AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

34895 stories from 30+ sources, refreshed continuously.

arXiv cs.CL (NLP)Research

SCOPE: Sequential Conformal Probing for Reliable OOD Rejection in LLM Services

arXiv:2606.21255v1 Announce Type: new Abstract: Rejecting inputs outside the defined in-distribution (IND) service scope is critical for large language model (LLM) services, where unsupported requests should be filtered before full generation. Existing out-of-distribution (OOD) detectors often rely on final outputs or final-layer representations, leaving unclear where service-boundary signals are most clearly encoded inside the model; they also lack a theoretical guarantee for held-out inputs. I

Read source article
arXiv cs.CL (NLP)Research

Factual Retrieval in LLMs Is a Redundant, Distributed and Non-Contiguous Process

arXiv:2606.21345v1 Announce Type: new Abstract: Large language models (LLMs) store and recall factual knowledge, yet the precise mechanism of how entity representations are transformed to enable specific attribute retrieval remains underexplored. In this work, we investigate this mechanism through the lens of an "attribute-computation path"-a sequence of computational steps over the entity representation required to elicit a target attribute. We then propose an iterative patching protocol to ide

Read source article
arXiv cs.CL (NLP)Research

Finetuning with Scientific Data Increases Hallucinations: A Multi-domain Factuality Evaluation of LLMs

arXiv:2606.21359v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to communicate and explain scientific concepts, yet their tendency to hallucinate poses significant risks in this high stakes use-case. Prior hallucination evaluation work remains largely restricted to the biomedical domain, treats hallucination as a binary task, and has not examined the growing family of scientifically fine-tuned LLMs. We address these gaps with SciFactCheck, a benchmark of 2,500

Read source article
arXiv cs.CL (NLP)Research

Evaluation of Small Language Models for Arabic Language Processing

arXiv:2606.21460v1 Announce Type: new Abstract: This paper evaluates the performance of twelve Small Language Models (SLMs) on Arabic natural language processing tasks. The study introduces a benchmark of 240 Arabic test items distributed across eight domains and ten language skills, covering both comprehension-oriented and generation-oriented tasks. All models were evaluated under a controlled zero-shot setting using a standardized Arabic-only prompt template. Model responses were assessed thro

Read source article
arXiv cs.CL (NLP)Research

Towards Pedagogically Aligned LLM Tutors for Math Mistake Remediation

arXiv:2606.21502v1 Announce Type: new Abstract: Large language models have strong potential for use in intelligent tutoring systems, but they often fail to follow effective pedagogical strategies, such as guiding students without revealing final answers. We study the application of a two-stage alignment pipeline for math mistake remediation, combining supervised fine-tuning on tutoring dialogs with Direct Preference Optimization on synthetic preference pairs. We construct a dataset that integrat

Read source article
arXiv cs.CL (NLP)Research

Dissecting Agentic RAG: A Component Ablation for Multi-Hop QA with a Local 7B Model

arXiv:2606.21553v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) systems combine iterative reasoning loops, query decomposition, and adaptive retrieval to tackle multi-hop question answering. However, the contribution of each component remains poorly understood, particularly under resource-constrained settings using only local language models. Many agentic designs add adaptive retrieval routing and deeper retrieval loops on the assumption that the added complexity hel

Read source article
arXiv cs.CL (NLP)Research

Rubric-as-Experts: Case-Specific MQM Rubrics for Translation Quality Evaluation

arXiv:2606.21559v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong potential in fine-grained translation quality evaluation (QE), yet existing MQM-based approaches typically rely on fixed rubric configurations shared across all translation samples. However, translation instances often differ substantially in error complexity, ambiguity, and required evaluation granularity, making static rubric allocation suboptimal for span-level error detection. We find that larger M

Read source article
arXiv cs.CL (NLP)Research

Per-Entity Bias Mapping for AI Visibility: Why Brand Mentions Require Entity-Specific Calibration

arXiv:2606.21595v1 Announce Type: new Abstract: AI-mediated answer systems increasingly determine how brands and organizations are represented to users. Existing approaches reduce visibility to mention rate or citation frequency. This paper argues that aggregate metrics are insufficient because entities exhibit systematically different AI visibility error profiles. We introduce Per-Entity Bias Mapping (PEBM): a ten-dimensional framework distinguishing raw from verified mentions. Three failure mo

Read source article
arXiv cs.CL (NLP)Research

LLM and Human Modes of Representation

arXiv:2606.21616v1 Announce Type: new Abstract: Much work on the cognitive foundations of AI has focussed on comparisons between the ways in which Large Language Models (LLMs) and humans process information and represent it. One aspect of this comparison involves determining the extent to which LLMs can achieve or surpass human performance on a variety of cognitively interesting tasks. A second explores points of convergence and divergence between LLM and human systems for processing information

Read source article
arXiv cs.CL (NLP)Research

CulMind: Benchmarking Multimodal Understanding and Reasoning in Chinese Cultural Heritage

arXiv:2606.21618v1 Announce Type: new Abstract: Evaluating Multimodal Large Language Models (MLLMs) in Chinese Cultural Heritage (CCH) requires fine-grained reasoning over visual, textual, stylistic, and historical clues. However, existing CCH benchmarks mainly emphasize final-answer accuracy, while the accuracy and completeness of reasoning processes remain underexplored. To address this gap, we introduce CulMind and CulMind-R: a high-quality benchmark for multimodal CCH covering 50 tasks from

Read source article
arXiv cs.CL (NLP)Research

Evaluating Document-Tuned Transformer Representations for Person-level Mental Health Assessment

arXiv:2606.21622v1 Announce Type: new Abstract: Person-level psychological assessment requires aggregating meaning across many messages from the same individual, a task that document-level training objectives were not explicitly designed for. We present a systematic, empirical comparison between architecturally matched traditional (a) base-transformers and (b) document-tuned-transformers (further contrastively fine-tuned at the document-level, sometimes referred to as "sentence transformers") un

Read source article
arXiv cs.CL (NLP)Research

CuratorKIT : Data Curation and Synthetic Data Generation for LLM Post-Training

arXiv:2606.21631v1 Announce Type: new Abstract: Data curation is a critical part of post-training pipelines for large language models, yet existing tools often treat ingestion, deduplication, synthetic generation, and quality filtering as separate stages. This fragmentation makes it difficult to audit pipeline decisions or understand why individual samples are rejected. CuratorKIT is an open-source Python library that covers this full lifecycle in a single configurable pipeline. The framework is

Read source article
arXiv cs.CL (NLP)Research

Behavioral and Representational Evidence of Binomial Ordering Preferences in Large Language Models

arXiv:2606.21645v1 Announce Type: new Abstract: Large language models (LLMs) can readily reproduce conventional expressions, yet their ability to model gradient frequency distributions remains underexplored. We investigate this using linguistic binomials, such as men and women, where both word permutations are grammatically valid but exhibit distinct, cross-linguistic variations in conventionality. We formalize binomial ordering as a distributional alignment problem, and construct a multilingual

Read source article
arXiv cs.LGResearch

Robustness Cannot be Reduced to Regularization: Studying Adversarial Training Beyond the Linear Case

arXiv:2606.21488v1 Announce Type: new Abstract: The vulnerability of ML models to adversarial examples has recently emerged as a major concern. While adversarial training is one of the most effective countermeasures to this issue, its high computational cost remains an obstacle to practical deployment. Recent progress in reducing this cost has relied, in the case of linear models, on a formal equivalence between the adversarial risk and a simpler form of regularized risk. This enabled significan

Read source article
arXiv cs.LGResearch

LIG: Layer-wise Integrated Gradients for Within-Layer Flow Analysis in Transformers

arXiv:2606.21564v1 Announce Type: new Abstract: Transformers achieve strong performance, but their internal computations remain opaque. We view each Transformer layer as a dynamic graph whose nodes are token representations and per-head attention outputs, with Multi-Head Attention (ATT) and MLP as module boundaries. On this graph we use LIG (Layer-wise Integrated Gradients), which applies set-to-set Integrated Gradients (IG) at nonlinear module boundaries. Set-to-set IG applies IG to a map from

Read source article
arXiv cs.LGResearch

The Cost Geometry of Belief: finite-resource inference under noisy observation

arXiv:2606.21585v1 Announce Type: new Abstract: We equip the space of beliefs with a cost geometry (what it costs to pass from one belief to another): optimal transport in Wasserstein space, reweighted conformally by Fisher information (the price of the precision at stake), distinct from the Fisher-Rao metric. In the setting we consider, a finite machine maintains a digital twin of a system; observing the territory through finite, noisy sensors, we model its coherent output as a belief: a probab

Read source article
arXiv cs.LGResearch

Geometric and Information Compression of Representations in Deep Learning

arXiv:2606.21593v1 Announce Type: new Abstract: Deep neural networks transform input data into latent representations that support a wide range of downstream tasks. These representations can be characterized along information-theoretic and geometric dimensions, but their relationship remains poorly understood. A central open question is whether low mutual information (MI) between inputs and representations necessarily implies geometrically compressed latent spaces and vice versa. We investigate

Read source article
arXiv cs.LGResearch

Learning to Place Guards by Reinforcement: A Geo-Free Neural Policy for the Vertex-Guard Art Gallery Problem

arXiv:2606.21604v1 Announce Type: new Abstract: Neural combinatorial optimization (NCO) has shown that policies trained by reinforcement can construct strong solutions to NP-hard problems directly from raw instances. What such a policy actually learns, as opposed to what its decoder expresses, remains much less clear. We study this distinction on the vertex-guard Art Gallery Problem, the NP-hard task of choosing polygon vertices from which to observe an entire region. A pointer-network policy is

Read source article
arXiv cs.LGResearch

HERALD: High-Throughput Block Diffusion LLM Serving via CPU-GPU Cooperative KV Cache Retrieval

arXiv:2606.21633v1 Announce Type: new Abstract: Diffusion LLMs (dLLMs) improve GPU utilization over autoregressive decoding by generating multiple tokens per forward pass, but their KV cache still grows linearly with context, limiting throughput at long contexts. KV cache offloading to host DRAM alleviates this memory pressure, but the limited PCIe bandwidth necessitates recalling only a sparse subset of KV entries. In block dLLMs, the relevant KV entries remain consistent across denoising steps

Read source article
arXiv cs.LGResearch

A new classification method based on Minimum Spanning Trees

arXiv:2606.21639v1 Announce Type: new Abstract: Minimum Spanning Trees have been used in unsupervised learning, particularly in clustering tasks, due to their ability to recognize clusters by removing edges that are considered inconsistent in defining those clusters. This paper aims to study the use of Minimum Spanning Trees in supervised learning. Specifically, we propose a classification algorithm based on Minimum Spanning Trees. To improve its performance, we introduce a robust version of the

Read source article
arXiv cs.LGResearch

Expressivity Saturation: Reduced Affine Region Usage Under Increasing Task Complexity

arXiv:2606.21687v1 Announce Type: new Abstract: Piecewise-affine neural networks (e.g., with ReLU or LeakyReLU activations) implement continuous piecewise-affine maps, and the number of affine regions provides a natural proxy for expressive capacity. However, the gap between theoretical region capacity and the affine regions realized after training remains insufficiently understood. We study this gap from two complementary perspectives. First, we give a rigorous, architecture-dependent theorem f

Read source article
arXiv cs.CL (NLP)Research

EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory

arXiv:2606.21649v1 Announce Type: new Abstract: Existing embedding models are inherently static: they encode text segments in isolation, ignoring their surrounding context and temporal order. This paper introduces EvoEmbedding, a novel embedding model that generates evolvable representations for retrieval. It is tailored for long-context scenarios, where information is dynamic, sequential, and requires continuous state tracking. Our design is simple: EvoEmbedding maintains a continuously updated

Read source article
arXiv cs.CL (NLP)Research

TACO: Task-Aware Column Description Generation Using LLMs

arXiv:2606.21685v1 Announce Type: new Abstract: Generating accurate and informative column descriptions (e.g. "membership status of customers" for the column name "cust_mem") is essential for a wide range of downstream NLP tasks on tabular data, including NL2SQL, table question answering, and entity linking. This problem arises in enterprises, domain sciences, government data portals, and so on. Despite its importance, most real-world datasets suffer from missing or cryptic documentation, often

Read source article
arXiv cs.CL (NLP)Research

Clinical Term Extraction using Open-Source Small Language Models

arXiv:2606.21689v1 Announce Type: new Abstract: Clinical information for amyotrophic lateral sclerosis (ALS) care documented in unstructured clinical notes limits downstream analysis without extraction into structured formats. Open-source small language models with few-shot prompting for detecting the presence of ALS-relevant clinical terms in patient documentation were evaluated without task-specific training data. The detection task targeted 17 categories spanning functional scores, respirator

Read source article
arXiv cs.CL (NLP)Research

When Compression Helps and When It Hurts: Condition-Aware Analysis of Chain-of-Thought Distillation

arXiv:2606.21704v1 Announce Type: new Abstract: Chain-of-Thought (CoT) distillation transfers multi-step reasoning from large reasoning models to smaller students, but verbose teacher traces inflate both training and inference cost. Existing CoT compression methods fall into two families, selective pruning and generative rewriting, yet prior studies have left key factors entangled: granularity is confounded with importance criteria in pruning, restructuring level is rarely isolated in rewriting,

Read source article
arXiv cs.CL (NLP)Research

PrivacyAlign: Contextual Privacy Alignment for LLM Agents

arXiv:2606.21710v1 Announce Type: new Abstract: AI agents acting on behalf of users are constantly making decisions, and for users to trust their agents, those decisions must align with what they actually want. Privacy is an important alignment problem for agents: every message, post, or tool call an agent makes is a contextual judgment about what is appropriate to share, with whom, and under which conditions. Because such judgments depend on social expectations and norms, human judgment does no

Read source article
arXiv cs.CL (NLP)Research

Denoising Iterative Self-Correction: Structured Verification Loops for Reliable LLM Reasoning

arXiv:2606.21724v1 Announce Type: new Abstract: Large language models produce fluent but often incorrect multi-step reasoning, and naive correction methods risk degrading already-correct answers. We introduce Denoising Iterative Self-Correction (DISC), a test-time procedure that treats verification question outputs as noisy measurements of where a solution may be corrupted. Using these signals, DISC progressively reduces errors across multiple verify-judge-correct passes, analogous to traditiona

Read source article
arXiv cs.CL (NLP)Research

CalVerT: Augmenting Agents with Calibrated Verifier Telemetry Improves Action and Learning in Knowledge-Intensive Tasks

arXiv:2606.21777v1 Announce Type: new Abstract: LLM agents in knowledge intensive question answering take retrieval and reasoning actions with incomplete knowledge about whether their current answer is uncertain, unsupported, or already complete. This produces two failure modes: committing to confident but unsupported answers, which hurts accuracy, and over-retrieving when the evidence in hand already suffices, resulting in wasted compute. To give agents a more complete picture of the state spac

Read source article
arXiv cs.CL (NLP)Research

When to Plan, When to Polish: Noise Level as a Granularity Axis for Diffusion Language Models

arXiv:2606.21802v1 Announce Type: new Abstract: Standard tokenwise diffusion LMs keep training corruption and inference commitment at token granularity throughout denoising. At high noise, this leaves scattered local fragments rather than coherent evidence, making it hard to form early coarse structure, exactly what planning-sensitive generation requires. Hierarchical planning methods add coarse stages to separate planning from wording, but they need extra planners, block latents, or two stage d

Read source article
arXiv cs.CL (NLP)Research

Test-Time Training with Next-Token Prediction

arXiv:2606.21803v1 Announce Type: new Abstract: Next-token prediction is the self-supervised signal that trains language models, and every observed prompt token provides the same signal at test time. We study whether this signal can define the inner-loop objective for test-time training (TTT) in pretrained long-context language models. Many TTT architectures require models to be trained with test-time adaptation in mind, limiting their direct applicability to released LLM checkpoints. While rece

Read source article
arXiv cs.CL (NLP)Research

Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI Dialogue

arXiv:2606.21844v1 Announce Type: new Abstract: As AI systems integrate into online spaces, differentiating them from humans in conversations is increasingly important. We present Inverse Turing Bench, a benchmark that evaluates LLMs and other models on their ability to differentiate humans and AI in multi-turn text. The benchmark provides a collection of paired dialogue transcripts, wherein one dialogue is between two humans and the other is between a human and an AI. The task is to correctly i

Read source article
arXiv cs.CL (NLP)Research

Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers

arXiv:2606.21848v1 Announce Type: new Abstract: We propose Keyless Attention, an attention mechanism that eliminates the key projection entirely, operating over queries and values only. This yields a Value-Only Cache that reduces KV cache memory and access overhead by exactly 50% over standard attention, while matching or exceeding standard attention's decode throughput. Beyond efficiency, we introduce Depth-$m$ Attention Factorization: standard attention computes a depth-2 factorization of the

Read source article
arXiv cs.CL (NLP)Research

The Language-Energy Divide: Measuring Energy Costs of Multilingual LLM Inference

arXiv:2606.21869v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in multilingual settings, yet the energy costs of serving these models across different languages remain poorly understood. We present a systematic study of inference energy consumption across languages with ML.Energy framework (Chung et al., 2026). We find striking disparities: energy consumption per output token varies by up to 8.3 times across languages, while total energy for a fixed set of

Read source article
arXiv cs.CL (NLP)Research

Scaling Performance and Low-Resource Annotation with Many-Shot In-Context Learning for Named Entity Recognition

arXiv:2606.21890v1 Announce Type: new Abstract: In-context learning (ICL) with large language models (LLMs) has emerged as a powerful alternative to fine-tuning for Named Entity Recognition (NER), achieving strong performance with minimal annotation and no additional training. However, prior work has shown that despite their adaptability, LLMs still lag behind fully supervised models such as fine-tuned BERT in structured tasks like NER. While existing studies on ICL for NER have mainly explored

Read source article
arXiv cs.CL (NLP)Research

Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding

arXiv:2606.21906v1 Announce Type: new Abstract: Autoregressive generation in large language models (LLMs) conventionally decodes from the final layer, assuming that deeper representations yield more reliable next-token predictions. We revisit this assumption by revealing a recurring Guess-Refine-Perturb dynamic: early layers form coarse guesses, intermediate layers refine reasoning-relevant semantics, and final layers can perturb these refined predictions toward generic or alignment-preferred to

Read source article
arXiv cs.CL (NLP)Research

Pre-Generation Hallucination Detection in Large Language Models via Soft-Target Attention Probing

arXiv:2606.21917v1 Announce Type: new Abstract: Detecting hallucination risk before generation enables abstention, retrieval augmentation, and routing decisions without incurring the cost of decoding. While prior work has shown that such risk can be estimated from a model's internal representations, existing approaches treat this as binary classification over a single decoded output. We instead formulate it as a risk-estimation problem. Under this formulation, we introduce soft-target supervisio

Read source article
arXiv cs.CL (NLP)Research

Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts

arXiv:2606.21939v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in contexts requiring complex moral reasoning and value trade-offs. However, existing evaluations typically rely on item-level behavioral metrics, which fail to capture how models structurally prioritize competing values as a cohesive system. To address this, we propose a symmetric human-LLM evaluation framework, grounded in Q methodology, to measure value-structure alignment. Under our protoco

Read source article
arXiv cs.CL (NLP)Research

OpenBioRQ: Unsolved Biomedical Research Questions for Agents

arXiv:2606.21959v1 Announce Type: new Abstract: A working citation looks like proof -- but the fact that a link resolves does not mean the cited paper supports the claim. I find that current agentic models rarely fabricate citations (over $99\%$ resolve), yet roughly $15.9\%$ link to the wrong paper. Existing benchmarks miss this failure mode: when a question has a fixed answer key, a model can reproduce the expected source from that key rather than independently verifying that the source suppor

Read source article
arXiv cs.CL (NLP)Research

Can LLMs Control Readability? A Multi-Dimensional Evaluation Framework for CEFR-Controlled Arabic Generation

arXiv:2606.21981v1 Announce Type: new Abstract: While Large Language Models (LLMs) can generate fluent Arabic text, their ability to reliably control readability levels remains unclear. We propose a multi-dimensional evaluation framework for Common European Framework of Reference for Language (CEFR)-controlled Arabic text generation, assessing whether instruction-following LLMs can serve as reliable generators for adaptive language learning. Our framework integrates controlled prompting, automat

Read source article
arXiv cs.CL (NLP)Research

Adding Robust Code-Switching Capabilities to High Performance Multilingual ASR

arXiv:2606.21990v1 Announce Type: new Abstract: Code-switching (CSW) remains challenging for large multi-lingual ASR systems in real-world deployment. While fine-tuning on synthetic CSW data is possible, it generally degrades strong monolingual baselines. Our goal is to preserve these capabilities while extending models to handle complex code-switching, including morphological variations across languages. We propose Bayesian factorized adaptation, which learns to efficiently integrate switching-

Read source article