AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

34182 stories from 30+ sources, refreshed continuously.

arXiv cs.CL (NLP)Research

GRAG: Generic Response-Augmented Generation Framework for Personalized Conversational Systems

arXiv:2606.21097v1 Announce Type: new Abstract: Deploying highly capable personalized conversational agents in resource-constrained or privacy-sensitive environments remains a significant challenge. We identify a fundamental bottleneck in the existing approaches: current training paradigms treat personalization and grounding as a single monolithic learning problem. Under these paradigms, language models are forced to simultaneously address what to say (content grounding) and how to say it in a u

Read source article
arXiv cs.CL (NLP)Research

LLM-Based Multi-Reference Evaluation for Efficient and Robust Assessment of Phrase Break Annotations

arXiv:2606.21098v1 Announce Type: new Abstract: Reliable evaluation of phrase break annotations is crucial, as subtle variations in prosodic boundaries directly affect the clarity and naturalness of speech. However, existing approaches exhibit major limitations: single-reference evaluation assumes a unique gold phrasing for an utterance despite multiple valid phrasings, while human judgment, though flexible, is labor-intensive and unscalable. To address these, we propose LLM-based Multi-Referenc

Read source article
arXiv cs.CL (NLP)Research

A Multi-Agent Audit Framework for High-Stakes Reasoning: Evaluation and Interpretability in Clinical Mental Health Screening

arXiv:2606.21123v1 Announce Type: new Abstract: High-stakes reasoning tasks necessitate transparent and verifiable workflows, yet conventional single-model large language models (LLMs) often struggle with hallucination and low interpretability under zero-shot paradigms. To address this general AI challenge, we propose a Multi-Agent Audit Framework that simulates a collaborative, multi-step verification process. We empirically validate this architecture in the sensitive domain of clinical mental

Read source article
arXiv cs.CL (NLP)Research

AdaMem: Learning What to Remember for Personalized Long-Horizon LLM Agents

arXiv:2606.21144v1 Announce Type: new Abstract: Long-term memory systems for Large Language Model (LLM) agents typically try to \emph{remember everything}, extracting memories uniformly to retain as many facts as possible. In production, however, inference cost and finite context budgets make this untenable: beyond consolidating raw dialogue into memory, an agent must exert \emph{write control}, efficiently keeping only the information each user actually cares about. Otherwise, long-horizon pers

Read source article
arXiv cs.CL (NLP)Research

Who Checks the Citations? Benchmarking Legal Hallucination Detection

arXiv:2606.21155v1 Announce Type: new Abstract: Attorneys, judges, and pro se filers increasingly use AI to draft legal documents, yet these tools frequently fabricate citations. Despite predictions that newer models would hallucinate less or that court sanctions would deter negligent filers, we found over 1,000 filings containing fabricated citations -- with this number growing year-over-year. This study evaluates whether AI-based systems can mitigate these errors by automatically detecting hal

Read source article
arXiv cs.CL (NLP)Research

Dementia-Agents: A Multi-Modal Multi-Agent System for Dementia Staging and Phenotyping

arXiv:2606.21168v1 Announce Type: new Abstract: Dementia diagnosis requires integrating multi-modal clinical assessments from diverse informants and clinicians under incomplete and heterogeneous data conditions. Yet most AI-driven approaches remain Alzheimer's disease (AD)-centric, framing the problem as binary AD detection or three-stage AD progression modeling within well-curated research settings. This pathology-driven paradigm overlooks the broader, syndrome-level nature of dementia, which s

Read source article
arXiv cs.CL (NLP)Research

OpenWER: Improving Cross-Lingual ASR Evaluation and Enabling Token-Based Accuracy Metrics

arXiv:2606.21237v1 Announce Type: new Abstract: Advances in deep learning and end-to-end Automatic Speech Recognition (ASR) have enabled robust multilingual models, but evaluation metrics remain limited in assessing accuracy. Efforts to improve or replace the common metric Word Error Rate (WER) often focus on English, leaving evaluations for low-resource languages under-explored and hindering fair cross-lingual comparisons. We present OpenWER, an open-source implementation that improves WER robu

Read source article
arXiv cs.CL (NLP)Research

SCOPE: Sequential Conformal Probing for Reliable OOD Rejection in LLM Services

arXiv:2606.21255v1 Announce Type: new Abstract: Rejecting inputs outside the defined in-distribution (IND) service scope is critical for large language model (LLM) services, where unsupported requests should be filtered before full generation. Existing out-of-distribution (OOD) detectors often rely on final outputs or final-layer representations, leaving unclear where service-boundary signals are most clearly encoded inside the model; they also lack a theoretical guarantee for held-out inputs. I

Read source article
arXiv cs.CL (NLP)Research

Factual Retrieval in LLMs Is a Redundant, Distributed and Non-Contiguous Process

arXiv:2606.21345v1 Announce Type: new Abstract: Large language models (LLMs) store and recall factual knowledge, yet the precise mechanism of how entity representations are transformed to enable specific attribute retrieval remains underexplored. In this work, we investigate this mechanism through the lens of an "attribute-computation path"-a sequence of computational steps over the entity representation required to elicit a target attribute. We then propose an iterative patching protocol to ide

Read source article
arXiv cs.CL (NLP)Research

Finetuning with Scientific Data Increases Hallucinations: A Multi-domain Factuality Evaluation of LLMs

arXiv:2606.21359v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to communicate and explain scientific concepts, yet their tendency to hallucinate poses significant risks in this high stakes use-case. Prior hallucination evaluation work remains largely restricted to the biomedical domain, treats hallucination as a binary task, and has not examined the growing family of scientifically fine-tuned LLMs. We address these gaps with SciFactCheck, a benchmark of 2,500

Read source article
arXiv cs.CL (NLP)Research

Evaluation of Small Language Models for Arabic Language Processing

arXiv:2606.21460v1 Announce Type: new Abstract: This paper evaluates the performance of twelve Small Language Models (SLMs) on Arabic natural language processing tasks. The study introduces a benchmark of 240 Arabic test items distributed across eight domains and ten language skills, covering both comprehension-oriented and generation-oriented tasks. All models were evaluated under a controlled zero-shot setting using a standardized Arabic-only prompt template. Model responses were assessed thro

Read source article
arXiv cs.CL (NLP)Research

Towards Pedagogically Aligned LLM Tutors for Math Mistake Remediation

arXiv:2606.21502v1 Announce Type: new Abstract: Large language models have strong potential for use in intelligent tutoring systems, but they often fail to follow effective pedagogical strategies, such as guiding students without revealing final answers. We study the application of a two-stage alignment pipeline for math mistake remediation, combining supervised fine-tuning on tutoring dialogs with Direct Preference Optimization on synthetic preference pairs. We construct a dataset that integrat

Read source article
arXiv cs.CL (NLP)Research

Dissecting Agentic RAG: A Component Ablation for Multi-Hop QA with a Local 7B Model

arXiv:2606.21553v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) systems combine iterative reasoning loops, query decomposition, and adaptive retrieval to tackle multi-hop question answering. However, the contribution of each component remains poorly understood, particularly under resource-constrained settings using only local language models. Many agentic designs add adaptive retrieval routing and deeper retrieval loops on the assumption that the added complexity hel

Read source article
arXiv cs.CL (NLP)Research

Rubric-as-Experts: Case-Specific MQM Rubrics for Translation Quality Evaluation

arXiv:2606.21559v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong potential in fine-grained translation quality evaluation (QE), yet existing MQM-based approaches typically rely on fixed rubric configurations shared across all translation samples. However, translation instances often differ substantially in error complexity, ambiguity, and required evaluation granularity, making static rubric allocation suboptimal for span-level error detection. We find that larger M

Read source article
arXiv cs.CL (NLP)Research

Per-Entity Bias Mapping for AI Visibility: Why Brand Mentions Require Entity-Specific Calibration

arXiv:2606.21595v1 Announce Type: new Abstract: AI-mediated answer systems increasingly determine how brands and organizations are represented to users. Existing approaches reduce visibility to mention rate or citation frequency. This paper argues that aggregate metrics are insufficient because entities exhibit systematically different AI visibility error profiles. We introduce Per-Entity Bias Mapping (PEBM): a ten-dimensional framework distinguishing raw from verified mentions. Three failure mo

Read source article
arXiv cs.CL (NLP)Research

LLM and Human Modes of Representation

arXiv:2606.21616v1 Announce Type: new Abstract: Much work on the cognitive foundations of AI has focussed on comparisons between the ways in which Large Language Models (LLMs) and humans process information and represent it. One aspect of this comparison involves determining the extent to which LLMs can achieve or surpass human performance on a variety of cognitively interesting tasks. A second explores points of convergence and divergence between LLM and human systems for processing information

Read source article
arXiv cs.CL (NLP)Research

CulMind: Benchmarking Multimodal Understanding and Reasoning in Chinese Cultural Heritage

arXiv:2606.21618v1 Announce Type: new Abstract: Evaluating Multimodal Large Language Models (MLLMs) in Chinese Cultural Heritage (CCH) requires fine-grained reasoning over visual, textual, stylistic, and historical clues. However, existing CCH benchmarks mainly emphasize final-answer accuracy, while the accuracy and completeness of reasoning processes remain underexplored. To address this gap, we introduce CulMind and CulMind-R: a high-quality benchmark for multimodal CCH covering 50 tasks from

Read source article
arXiv cs.CL (NLP)Research

Evaluating Document-Tuned Transformer Representations for Person-level Mental Health Assessment

arXiv:2606.21622v1 Announce Type: new Abstract: Person-level psychological assessment requires aggregating meaning across many messages from the same individual, a task that document-level training objectives were not explicitly designed for. We present a systematic, empirical comparison between architecturally matched traditional (a) base-transformers and (b) document-tuned-transformers (further contrastively fine-tuned at the document-level, sometimes referred to as "sentence transformers") un

Read source article
arXiv cs.CL (NLP)Research

CuratorKIT : Data Curation and Synthetic Data Generation for LLM Post-Training

arXiv:2606.21631v1 Announce Type: new Abstract: Data curation is a critical part of post-training pipelines for large language models, yet existing tools often treat ingestion, deduplication, synthetic generation, and quality filtering as separate stages. This fragmentation makes it difficult to audit pipeline decisions or understand why individual samples are rejected. CuratorKIT is an open-source Python library that covers this full lifecycle in a single configurable pipeline. The framework is

Read source article
arXiv cs.CL (NLP)Research

Behavioral and Representational Evidence of Binomial Ordering Preferences in Large Language Models

arXiv:2606.21645v1 Announce Type: new Abstract: Large language models (LLMs) can readily reproduce conventional expressions, yet their ability to model gradient frequency distributions remains underexplored. We investigate this using linguistic binomials, such as men and women, where both word permutations are grammatically valid but exhibit distinct, cross-linguistic variations in conventionality. We formalize binomial ordering as a distributional alignment problem, and construct a multilingual

Read source article
arXiv cs.LGResearch

Robustness Cannot be Reduced to Regularization: Studying Adversarial Training Beyond the Linear Case

arXiv:2606.21488v1 Announce Type: new Abstract: The vulnerability of ML models to adversarial examples has recently emerged as a major concern. While adversarial training is one of the most effective countermeasures to this issue, its high computational cost remains an obstacle to practical deployment. Recent progress in reducing this cost has relied, in the case of linear models, on a formal equivalence between the adversarial risk and a simpler form of regularized risk. This enabled significan

Read source article
arXiv cs.LGResearch

LIG: Layer-wise Integrated Gradients for Within-Layer Flow Analysis in Transformers

arXiv:2606.21564v1 Announce Type: new Abstract: Transformers achieve strong performance, but their internal computations remain opaque. We view each Transformer layer as a dynamic graph whose nodes are token representations and per-head attention outputs, with Multi-Head Attention (ATT) and MLP as module boundaries. On this graph we use LIG (Layer-wise Integrated Gradients), which applies set-to-set Integrated Gradients (IG) at nonlinear module boundaries. Set-to-set IG applies IG to a map from

Read source article
arXiv cs.LGResearch

The Cost Geometry of Belief: finite-resource inference under noisy observation

arXiv:2606.21585v1 Announce Type: new Abstract: We equip the space of beliefs with a cost geometry (what it costs to pass from one belief to another): optimal transport in Wasserstein space, reweighted conformally by Fisher information (the price of the precision at stake), distinct from the Fisher-Rao metric. In the setting we consider, a finite machine maintains a digital twin of a system; observing the territory through finite, noisy sensors, we model its coherent output as a belief: a probab

Read source article
arXiv cs.LGResearch

Geometric and Information Compression of Representations in Deep Learning

arXiv:2606.21593v1 Announce Type: new Abstract: Deep neural networks transform input data into latent representations that support a wide range of downstream tasks. These representations can be characterized along information-theoretic and geometric dimensions, but their relationship remains poorly understood. A central open question is whether low mutual information (MI) between inputs and representations necessarily implies geometrically compressed latent spaces and vice versa. We investigate

Read source article
arXiv cs.LGResearch

Learning to Place Guards by Reinforcement: A Geo-Free Neural Policy for the Vertex-Guard Art Gallery Problem

arXiv:2606.21604v1 Announce Type: new Abstract: Neural combinatorial optimization (NCO) has shown that policies trained by reinforcement can construct strong solutions to NP-hard problems directly from raw instances. What such a policy actually learns, as opposed to what its decoder expresses, remains much less clear. We study this distinction on the vertex-guard Art Gallery Problem, the NP-hard task of choosing polygon vertices from which to observe an entire region. A pointer-network policy is

Read source article
arXiv cs.LGResearch

HERALD: High-Throughput Block Diffusion LLM Serving via CPU-GPU Cooperative KV Cache Retrieval

arXiv:2606.21633v1 Announce Type: new Abstract: Diffusion LLMs (dLLMs) improve GPU utilization over autoregressive decoding by generating multiple tokens per forward pass, but their KV cache still grows linearly with context, limiting throughput at long contexts. KV cache offloading to host DRAM alleviates this memory pressure, but the limited PCIe bandwidth necessitates recalling only a sparse subset of KV entries. In block dLLMs, the relevant KV entries remain consistent across denoising steps

Read source article
arXiv cs.LGResearch

A new classification method based on Minimum Spanning Trees

arXiv:2606.21639v1 Announce Type: new Abstract: Minimum Spanning Trees have been used in unsupervised learning, particularly in clustering tasks, due to their ability to recognize clusters by removing edges that are considered inconsistent in defining those clusters. This paper aims to study the use of Minimum Spanning Trees in supervised learning. Specifically, we propose a classification algorithm based on Minimum Spanning Trees. To improve its performance, we introduce a robust version of the

Read source article
arXiv cs.LGResearch

Expressivity Saturation: Reduced Affine Region Usage Under Increasing Task Complexity

arXiv:2606.21687v1 Announce Type: new Abstract: Piecewise-affine neural networks (e.g., with ReLU or LeakyReLU activations) implement continuous piecewise-affine maps, and the number of affine regions provides a natural proxy for expressive capacity. However, the gap between theoretical region capacity and the affine regions realized after training remains insufficiently understood. We study this gap from two complementary perspectives. First, we give a rigorous, architecture-dependent theorem f

Read source article
arXiv cs.CL (NLP)Research

EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory

arXiv:2606.21649v1 Announce Type: new Abstract: Existing embedding models are inherently static: they encode text segments in isolation, ignoring their surrounding context and temporal order. This paper introduces EvoEmbedding, a novel embedding model that generates evolvable representations for retrieval. It is tailored for long-context scenarios, where information is dynamic, sequential, and requires continuous state tracking. Our design is simple: EvoEmbedding maintains a continuously updated

Read source article
arXiv cs.CL (NLP)Research

TACO: Task-Aware Column Description Generation Using LLMs

arXiv:2606.21685v1 Announce Type: new Abstract: Generating accurate and informative column descriptions (e.g. "membership status of customers" for the column name "cust_mem") is essential for a wide range of downstream NLP tasks on tabular data, including NL2SQL, table question answering, and entity linking. This problem arises in enterprises, domain sciences, government data portals, and so on. Despite its importance, most real-world datasets suffer from missing or cryptic documentation, often

Read source article
arXiv cs.CL (NLP)Research

Clinical Term Extraction using Open-Source Small Language Models

arXiv:2606.21689v1 Announce Type: new Abstract: Clinical information for amyotrophic lateral sclerosis (ALS) care documented in unstructured clinical notes limits downstream analysis without extraction into structured formats. Open-source small language models with few-shot prompting for detecting the presence of ALS-relevant clinical terms in patient documentation were evaluated without task-specific training data. The detection task targeted 17 categories spanning functional scores, respirator

Read source article
arXiv cs.CL (NLP)Research

When Compression Helps and When It Hurts: Condition-Aware Analysis of Chain-of-Thought Distillation

arXiv:2606.21704v1 Announce Type: new Abstract: Chain-of-Thought (CoT) distillation transfers multi-step reasoning from large reasoning models to smaller students, but verbose teacher traces inflate both training and inference cost. Existing CoT compression methods fall into two families, selective pruning and generative rewriting, yet prior studies have left key factors entangled: granularity is confounded with importance criteria in pruning, restructuring level is rarely isolated in rewriting,

Read source article
arXiv cs.CL (NLP)Research

PrivacyAlign: Contextual Privacy Alignment for LLM Agents

arXiv:2606.21710v1 Announce Type: new Abstract: AI agents acting on behalf of users are constantly making decisions, and for users to trust their agents, those decisions must align with what they actually want. Privacy is an important alignment problem for agents: every message, post, or tool call an agent makes is a contextual judgment about what is appropriate to share, with whom, and under which conditions. Because such judgments depend on social expectations and norms, human judgment does no

Read source article
arXiv cs.CL (NLP)Research

Denoising Iterative Self-Correction: Structured Verification Loops for Reliable LLM Reasoning

arXiv:2606.21724v1 Announce Type: new Abstract: Large language models produce fluent but often incorrect multi-step reasoning, and naive correction methods risk degrading already-correct answers. We introduce Denoising Iterative Self-Correction (DISC), a test-time procedure that treats verification question outputs as noisy measurements of where a solution may be corrupted. Using these signals, DISC progressively reduces errors across multiple verify-judge-correct passes, analogous to traditiona

Read source article
arXiv cs.CL (NLP)Research

CalVerT: Augmenting Agents with Calibrated Verifier Telemetry Improves Action and Learning in Knowledge-Intensive Tasks

arXiv:2606.21777v1 Announce Type: new Abstract: LLM agents in knowledge intensive question answering take retrieval and reasoning actions with incomplete knowledge about whether their current answer is uncertain, unsupported, or already complete. This produces two failure modes: committing to confident but unsupported answers, which hurts accuracy, and over-retrieving when the evidence in hand already suffices, resulting in wasted compute. To give agents a more complete picture of the state spac

Read source article
arXiv cs.CL (NLP)Research

When to Plan, When to Polish: Noise Level as a Granularity Axis for Diffusion Language Models

arXiv:2606.21802v1 Announce Type: new Abstract: Standard tokenwise diffusion LMs keep training corruption and inference commitment at token granularity throughout denoising. At high noise, this leaves scattered local fragments rather than coherent evidence, making it hard to form early coarse structure, exactly what planning-sensitive generation requires. Hierarchical planning methods add coarse stages to separate planning from wording, but they need extra planners, block latents, or two stage d

Read source article
arXiv cs.CL (NLP)Research

Test-Time Training with Next-Token Prediction

arXiv:2606.21803v1 Announce Type: new Abstract: Next-token prediction is the self-supervised signal that trains language models, and every observed prompt token provides the same signal at test time. We study whether this signal can define the inner-loop objective for test-time training (TTT) in pretrained long-context language models. Many TTT architectures require models to be trained with test-time adaptation in mind, limiting their direct applicability to released LLM checkpoints. While rece

Read source article
arXiv cs.CL (NLP)Research

Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI Dialogue

arXiv:2606.21844v1 Announce Type: new Abstract: As AI systems integrate into online spaces, differentiating them from humans in conversations is increasingly important. We present Inverse Turing Bench, a benchmark that evaluates LLMs and other models on their ability to differentiate humans and AI in multi-turn text. The benchmark provides a collection of paired dialogue transcripts, wherein one dialogue is between two humans and the other is between a human and an AI. The task is to correctly i

Read source article
arXiv cs.CL (NLP)Research

Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers

arXiv:2606.21848v1 Announce Type: new Abstract: We propose Keyless Attention, an attention mechanism that eliminates the key projection entirely, operating over queries and values only. This yields a Value-Only Cache that reduces KV cache memory and access overhead by exactly 50% over standard attention, while matching or exceeding standard attention's decode throughput. Beyond efficiency, we introduce Depth-$m$ Attention Factorization: standard attention computes a depth-2 factorization of the

Read source article
arXiv cs.CL (NLP)Research

The Language-Energy Divide: Measuring Energy Costs of Multilingual LLM Inference

arXiv:2606.21869v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in multilingual settings, yet the energy costs of serving these models across different languages remain poorly understood. We present a systematic study of inference energy consumption across languages with ML.Energy framework (Chung et al., 2026). We find striking disparities: energy consumption per output token varies by up to 8.3 times across languages, while total energy for a fixed set of

Read source article