AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

10639 stories from 30+ sources, refreshed continuously.

arXiv cs.CL (NLP)Research

Recipes for Steering and Scaling LLMs via Sampling

arXiv:2608.26120v1 Announce Type: new Abstract: Large Language Models (LLMs) are probabilistic models, typically defined by an autoregressive factorization. While recent work has begun to study richer target distributions beyond the base model, the sampling strategies remain highly inefficient. In this paper, we present a flexible and theoretically grounded framework for steering and scaling autoregressive LLMs with sampling. Within this framework, we describe two algorithms -- one based on Sequ

Read source article
arXiv cs.CL (NLP)Research

Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention

arXiv:2608.26121v1 Announce Type: new Abstract: Large language models state false facts as fluently as true ones, yet a model often "knows" internally when it is on shaky ground: the probability it assigns to its own answer tends to dip on the facts it gets wrong. The usual way to act on this, teaching a model to abstain rather than guess, requires a labelled dataset of right and wrong answers. We ask whether the model's own confidence, which is free and needs no labels, can do that job instead.

Read source article
arXiv cs.CL (NLP)Research

Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs

arXiv:2608.26123v1 Announce Type: new Abstract: Large language models (LLMs) are trained predominantly on English-language internet text that over-represents certain cultural narratives, raising concerns that models flatten the diversity of non-Western storytelling traditions into a single homogenized archetype. We present a pilot computational study examining this across three maximally distinct Indian regional oral and literary traditions: the Rajasthani Pabuji epic, classical Tamil Sangam poe

Read source article
arXiv cs.CL (NLP)Research

Natural-Language Policies to Executable Decisions: An Interpretable Large Language Model Framework

arXiv:2608.26124v1 Announce Type: new Abstract: Pricing automation in large-scale tourism is challenging because travel orders are highly unstructured, while pricing policies are complex, rapidly evolving, and inherently open-ended. Traditional rule engines are brittle and costly to maintain, whereas unconstrained LLM agents lack the reliability and auditability required for financial decisions. We present a production-grade LLM-powered pricing system with a strict decision boundary: LLMs perfor

Read source article
arXiv cs.CL (NLP)Research

Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales

arXiv:2608.26125v1 Announce Type: new Abstract: Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation, particularly when detached from sociocultural context. We propose a \emph{training-time} explainability framework that aligns model reasoning with human-annotated rationales, improving both classification performance and i

Read source article
arXiv cs.CL (NLP)Research

TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack

arXiv:2608.26126v1 Announce Type: new Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows require joint grounding in normative specifications, operational telemetry, vendor-specific fault evidence, and exact RF/network calculations. However, current LLM integration in telecom remains bottlenecked by a two-sided capability gap: generic reasoners often lack telecom-specific grounding, while domain-specific tele

Read source article
arXiv cs.CL (NLP)Research

FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes

arXiv:2608.26129v1 Announce Type: new Abstract: Scientific peer review datasets have trained AI systems exclusively on Computer Science and Machine Learning venues, producing models that critique ablation studies yet have never seen a biology reviewer demand contamination controls or a chemist question Nuclear Magnetic Resonance (NMR) spectral assignments. We introduce FIRSTPASS, the first large-scale peer review dataset built on complete multi-round editorial dialogues from a multidisciplinary

Read source article
arXiv cs.CL (NLP)Research

Agents Don't Paginate: First-Chunk Selection for LLM Tool Responses

arXiv:2608.26130v1 Announce Type: new Abstract: Coding agents built on large language models (LLMs), such as Claude Code, Cursor, OpenAI Codex, GitHub Copilot, and Aider, receive tool responses that routinely exceed the agent's per-turn token budget. The standard remedy, pagination, is available in every protocol that produced these responses; yet across the corpus of session logs from a public Model Context Protocol middleware we observed no agent-initiated requests for a second chunk. The firs

Read source article
arXiv cs.CL (NLP)Research

Evaluating Language Models in Realistic Conversational Contexts

arXiv:2608.26131v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed to serve open-ended, multi-turn interactions, evaluating conversational quality at human scale has become a central challenge. Existing evaluation frameworks built for summarization, translation, or short-form QA tasks fall short of adequately measuring the consistency of human-scale dialogue, especially when derivation and validation of these metrics themselves often rely on synthetic rathe

Read source article
arXiv cs.CL (NLP)Research

Agent Seer: Synthesizing Scenarios from Specification Understanding

arXiv:2608.26133v1 Announce Type: new Abstract: Evaluating AI agents that use external tools requires realistic test scenarios that capture how practitioners compose tools and iterate across conversation turns. Constructing such scenarios by hand demands deep domain expertise, does not scale across tool ecosystems, and produces static benchmarks that cannot track evolving APIs. We observe that tool specifications -- function names, natural-language descriptions, and typed parameter schemas -- al

Read source article
arXiv cs.CL (NLP)Research

Data Science Approaches to Evaluating Honours Candidates

arXiv:2608.26135v1 Announce Type: new Abstract: We present a modular data-science pipeline for estimating public sentiment towards individuals from fragmented, unstructured open-source intelligence (OSINT). The method chains web search, text extraction, relevance filtering, tokenisation, co-reference resolution, and sentiment analysis to convert heterogeneous web material into auditable person-level sentiment distributions. We compare AFINN and VADER with MINOS, a domain-informed sentiment algor

Read source article
arXiv cs.CL (NLP)Research

Interpretable, Fairly Evaluated Automated L2 Speaking Assessment that Beats the Single-Human Ceiling and Why Pause Encoding Does Not Change LLM Fluency Scores

arXiv:2608.26137v1 Announce Type: new Abstract: Second-language (L2) English learners can rarely rehearse speaking with a partner. Speaking is also the most anxiety-laden skill. These gaps drive a fast-growing market for automated speaking practice and scoring. But an automated score is trustworthy only if it is accurate, interpretable, fair, and benchmarked against the right human bar. We build an interpretable feature-plus-LLM hybrid for spontaneous L2 dialogue. We evaluate it without ever fit

Read source article
arXiv cs.CL (NLP)Research

Cross-Platform Generalisation Failure in Mental Health Natural Language Processing: A Five-Axis Fairness Audit of Transformer Models on Social Media

arXiv:2608.26138v1 Announce Type: new Abstract: We introduce the Cross-Platform Fairness Evaluation (CPFE) framework -- a five-axis audit protocol covering discriminative performance, calibration, statistical significance, prediction equity, and attribution stability -- and apply it to four transformer models (BERT, RoBERTa, Emotion-DistilRoBERTa, GoEmotions-RoBERTa) trained on a Kaggle mental health corpus (n=35,556) and evaluated on Reddit (n=6,257) and Twitter (n=2,883) test sets with emotion

Read source article
arXiv cs.CL (NLP)Research

Syntax vs. Semantics: How Transformers Learn Deep Dependencies

arXiv:2608.26139v1 Announce Type: new Abstract: Large Language Models demonstrate remarkable syntactic fluency, yet the optimization dynamics governing their acquisition of deep semantic dependencies remain poorly understood. We propose a mechanistic framework that models this learning process as a competition between Surface Statistics and Deep Semantics. Our theoretical analysis identifies a ``Gradient Starvation" phenomenon where the error signals for sparse semantic dependencies are actively

Read source article
arXiv cs.CL (NLP)Research

Affix Cache for Diffusion Large Language Models

arXiv:2608.26140v1 Announce Type: new Abstract: Diffusion Large Language Models (DLLMs) enable non-autoregressive decoding and bidirectional context modeling, but efficient inference remains challenging. Unlike autoregressive systems, whose key-value (KV) cache can be reused for shared prefixes, DLLMs couple the KV states of shared context tokens with evolving generated tokens through bidirectional attention, making naive cache reuse stale while full recomputation is expensive. We present ACache

Read source article
arXiv cs.CL (NLP)Research

AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking

arXiv:2608.26141v1 Announce Type: new Abstract: Multimodal large language models have demonstrated strong document reasoning capabilities by incorporating explicit thinking processes. While this capability significantly improves performance on challenging tasks, current models apply such deep reasoning uniformly to all questions, resulting in unnecessary computational overhead for simple task. This not only degrades user experience but also negatively impact accuracy on benchmark datasets. We id

Read source article
arXiv cs.CL (NLP)Research

Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation

arXiv:2608.26142v1 Announce Type: new Abstract: Referring Expression Segmentation (RES) aims to generate pixel-wise segmentation masks from complex and implicit textual queries. While recent advances in Multimodal Large Language Models (MLLMs) have substantially boosted RES performance, their prohibitive computational overhead remains a critical bottleneck, which, however, is rarely explored. To fill this gap, we first evaluate typical token compression methods on this task and observe a surpris

Read source article
arXiv cs.CL (NLP)Research

Why Current XAI Is Not Enough for Arabic NLP: A Critical Survey of the Explainability Gap

arXiv:2608.26144v1 Announce Type: new Abstract: Explainable AI (XAI) is now a major theme in NLP; however, Arabic NLP remains under-explained in three connected senses. First, there is a method gap: Arabic XAI relies heavily on a small set of post-hoc techniques such as LIME, SHAP, attention visualization, and saliency, while broader NLP XAI offers richer diagnostic, counterfactual, probing, rationale-based, and human-centered methods. Second, there is a task gap: existing Arabic XAI work is con

Read source article
Apple Machine LearningResearch

LLMs Are Not (Consistently) Bayesian: Quantifying Internal (In)consistencies of LLMs’ Probabilistic Beliefs

Modern AI systems are being deployed in complex domains such as medicine, science, and law, where there is often not a single correct answer given the observed evidence. Such systems must be able to represent and update uncertain beliefs about the world as new evidence arrives to make rational decisions. We introduce the novel technique of studying LLMs as information processing rules and utilize the information processing gap—the deviation from Bayes updates—to study the internal (in)consistenc

Read source article
Apple Machine LearningResearch

Agent Seer: Synthesizing Scenarios from Specification Understanding

Evaluating AI agents that use external tools requires realistic test scenarios that capture how practitioners compose tools and iterate across conversation turns. Constructing such scenarios by hand demands deep domain expertise, does not scale across tool ecosystems, and produces static benchmarks that cannot track evolving APIs. We observe that tool specifications—function names, natural-language descriptions, and typed parameter schemas—already encode sufficient semantic information to synthe

Read source article
r/MachineLearningResearch

py-evoFE: Automated Evolutionary Feature Engineering for Tabular ML in Python (Genetic Algorithms + Scikit-Learn + Polars) [P]

<!-- SC_OFF --><div class="md"><p>Hey everyone!</p> <p>I’m excited to announce the release of <strong><code>py-evoFE</code></strong> (v0.3.0) — an open-source Python library that uses genetic algorithms to automatically discover, combine, and optimize feature transformations for tabular datasets.</p> <ul> <li><strong>GitHub:</strong> <a href="https://github.com/tanopereira/py-evoFE">https://github.com/tanopereira/py-evoFE</a></li> <li><strong>PyPI:</strong> <code>pip install py-evoFE</code></li>

Read source article
r/MachineLearningResearch

Can AI Improve Itself? RSI Might Be the Answer [R]

<!-- SC_OFF --><div class="md"><p>Can an AI make other AIs better? And what stops it from just cheating? Last month, an OpenAI eval agent escaped its sandbox and broke into Hugging Face, apparently to grab test solutions from a benchmark. It's exactly what you'd expect from a system that rewrites agents and reads its own grades. We set out to measure recursive self-improvement anyway, with the exam locked outside its sandbox.</p> <p>We introduce HarnessOpt-Bench, which scores an LLM on how much

Read source article
NVIDIA BlogResearch

GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026

NVIDIA’s Gamescom announcements are revealing what’s next for GeForce NOW, with new ways to play, more supported devices and platforms, and even more big PC games headed to the cloud. New NVIDIA DLSS 4.5 technology controls give members more ways to fine-tune gameplay, while expanded support for new Steam devices, GOG single sign-on, Firefox browser […]

Read source article
MIT Tech ReviewResearch

The Download: inside OpenAI’s Hugging Face hack, and a new EV takes on the US

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. The inside story on why OpenAI agents hacked Hugging Face The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with…

Read source article
The Download: inside OpenAI’s Hugging Face hack, and a new EV takes on the US
Apple Machine LearningResearch

From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers

Designing effective reward signals for open-domain question answering is challenging because high-quality responses must simultaneously satisfy multiple aspects of answer quality that are difficult to capture with a holistic scalar objective. We introduce a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions, providing fine-grained supervision during post-training. Averaged across three evaluation axes

Read source article
r/MachineLearningResearch

A dataset with 52 Text to image model evaluation [P]

<!-- SC_OFF --><div class="md"><p>I created a simple text to image benchmark.</p> <p>I curated <strong>192 prompts that are difficult for T2I models</strong> in various ways: text rendering, spatial reasoning, human realism, negations, etc...</p> <p>I then asked a VLM to judge every output against a pre-specified binary question with the ground truth baked in.</p> <p>I'm publishing all the results including the images. (Most public T2I leaderboards don't publish the actual images and that's a sh

Read source article
NVIDIA BlogResearch

NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory

The next wave of AI is placing new demands on infrastructure. As AI agents and trillion-parameter workloads become mainstream, the performance of AI infrastructure depends not only on compute, but on how compute, memory, storage, networking and software are designed together as a unified system. To help hyperscalers and AI innovators build the next generation […]

Read source article
NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory
MIT Tech ReviewResearch

The inside story on why OpenAI agents hacked Hugging Face

The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released today. The hack, which a group of agents undertook to find solutions for a cybersecurity test that they were stuck on, has confirmed some experts’…

Read source article
Amazon ScienceResearch

When LLM judges agree, should we believe them?

Discounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.

Read source article
MIT Tech ReviewResearch

The Download: the Kids issue arrives, and Bill Gates reveals his AI fears

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Introducing: the Kids issue If the desire to limit kids’ use of technology was once a subcurrent, it has become a raging flood. Countries around the world are banning children from…

Read source article
The Download: the Kids issue arrives, and Bill Gates reveals his AI fears
IEEE Spectrum AIResearch

New Platform Peers Inside AI’s Black Box

Prompt Claude, ChatGPT, Gemini, or any other popular large language model with a question like “What is the best film ever made?” and the response will vary. And you (and most worryingly, the people who built the LLM) have little idea exactly how it came up with that specific answer. This mysterious behavior can be useful in some situations. But—as highlighted by a recent incident where OpenAI could not explain why its advanced prerelease model hacked AI company Hugging Face—it can have negative

Read source article
New Platform Peers Inside AI’s Black Box
MIT Tech ReviewResearch

Raised on AI

When my oldest child was born, I immediately set up Gmail and Twitter accounts in her name. I broadly announced her birth online and proceeded to plaster her photo across all sorts of platforms. In short, I began creating her digital footprint long before she could stand on her own two feet. Fast-forward a couple…

Read source article
Raised on AI
MIT Tech ReviewResearch

AI models flub these intelligence tests. Can you fare any better?

Puzzles and games have been central to AI development since the very beginning. Just as we humans like to test our smarts with crosswords or logic puzzles, developers can test how far models have advanced with a gaming gauntlet. The term “machine learning” was popularized in a 1959 article by the IBM computer scientist Arthur…

Read source article
r/MachineLearningResearch

Millwright — experimenting with an end-to-end machine learning framework in Rust [P]

<!-- SC_OFF --><div class="md"><p>I've been working on an open-source project called <strong>Millwright</strong>, an attempt to explore what an end-to-end machine learning workflow could look like in Rust.</p> <p><a href="https://millwright-rs.dev/">https://millwright-rs.dev/</a></p> <p>This started while I was learning and building ML tooling in Rust.</p> <p>I kept finding capable individual libraries, but also gaps between them. Training a model was rarely the problem. Building the workflow ar

Read source article
MIT Tech ReviewResearch

Bill Gates says we’ve passed AI’s danger thresholds. Now what?

It’s a glorious day in Kirkland, Washington, an affluent Seattle suburb on the eastern shore of Lake Washington. The temperature is in the mid-80s, and the sky is incapable of being any more blue. The view from the Gates Ventures conference room overlooks the Carillon Point Marina, where a flotilla of expensive boats bob in…

Read source article
Apple Machine LearningResearch

PROOF-Gen: From Optimized Data to Better Distillation

Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-calling agents re-run this stage on a daily or weekly cadence, paying the frontier-teacher cost each cycle, yet the mechanism is generate-and-filter (keep the teacher’s passing trajectories, discard the rest) and each cycle leaves behind the same hard scenarios because failures supply no signal. On τ

Read source article
Apple Machine LearningResearch

IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining

Recent advancements in large language models have intensified the need for efficient and deployable models within limited inference budgets. Structured pruning pipelines have shown promise in token efficiency compared to training target-size models from scratch. In this paper, we advocate incorporating enlarged model pretraining, which is often ignored in previous works, into pruning. We study the enlarge-and-prune pipeline as an integrated system to address two critical questions: whether it is

Read source article