AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

34182 stories from 30+ sources, refreshed continuously.

arXiv cs.CL (NLP)Research

Scaling Performance and Low-Resource Annotation with Many-Shot In-Context Learning for Named Entity Recognition

arXiv:2606.21890v1 Announce Type: new Abstract: In-context learning (ICL) with large language models (LLMs) has emerged as a powerful alternative to fine-tuning for Named Entity Recognition (NER), achieving strong performance with minimal annotation and no additional training. However, prior work has shown that despite their adaptability, LLMs still lag behind fully supervised models such as fine-tuned BERT in structured tasks like NER. While existing studies on ICL for NER have mainly explored

Read source article
arXiv cs.CL (NLP)Research

Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding

arXiv:2606.21906v1 Announce Type: new Abstract: Autoregressive generation in large language models (LLMs) conventionally decodes from the final layer, assuming that deeper representations yield more reliable next-token predictions. We revisit this assumption by revealing a recurring Guess-Refine-Perturb dynamic: early layers form coarse guesses, intermediate layers refine reasoning-relevant semantics, and final layers can perturb these refined predictions toward generic or alignment-preferred to

Read source article
arXiv cs.CL (NLP)Research

Pre-Generation Hallucination Detection in Large Language Models via Soft-Target Attention Probing

arXiv:2606.21917v1 Announce Type: new Abstract: Detecting hallucination risk before generation enables abstention, retrieval augmentation, and routing decisions without incurring the cost of decoding. While prior work has shown that such risk can be estimated from a model's internal representations, existing approaches treat this as binary classification over a single decoded output. We instead formulate it as a risk-estimation problem. Under this formulation, we introduce soft-target supervisio

Read source article
arXiv cs.CL (NLP)Research

Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts

arXiv:2606.21939v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in contexts requiring complex moral reasoning and value trade-offs. However, existing evaluations typically rely on item-level behavioral metrics, which fail to capture how models structurally prioritize competing values as a cohesive system. To address this, we propose a symmetric human-LLM evaluation framework, grounded in Q methodology, to measure value-structure alignment. Under our protoco

Read source article
arXiv cs.CL (NLP)Research

OpenBioRQ: Unsolved Biomedical Research Questions for Agents

arXiv:2606.21959v1 Announce Type: new Abstract: A working citation looks like proof -- but the fact that a link resolves does not mean the cited paper supports the claim. I find that current agentic models rarely fabricate citations (over $99\%$ resolve), yet roughly $15.9\%$ link to the wrong paper. Existing benchmarks miss this failure mode: when a question has a fixed answer key, a model can reproduce the expected source from that key rather than independently verifying that the source suppor

Read source article
arXiv cs.CL (NLP)Research

Can LLMs Control Readability? A Multi-Dimensional Evaluation Framework for CEFR-Controlled Arabic Generation

arXiv:2606.21981v1 Announce Type: new Abstract: While Large Language Models (LLMs) can generate fluent Arabic text, their ability to reliably control readability levels remains unclear. We propose a multi-dimensional evaluation framework for Common European Framework of Reference for Language (CEFR)-controlled Arabic text generation, assessing whether instruction-following LLMs can serve as reliable generators for adaptive language learning. Our framework integrates controlled prompting, automat

Read source article
arXiv cs.CL (NLP)Research

Adding Robust Code-Switching Capabilities to High Performance Multilingual ASR

arXiv:2606.21990v1 Announce Type: new Abstract: Code-switching (CSW) remains challenging for large multi-lingual ASR systems in real-world deployment. While fine-tuning on synthetic CSW data is possible, it generally degrades strong monolingual baselines. Our goal is to preserve these capabilities while extending models to handle complex code-switching, including morphological variations across languages. We propose Bayesian factorized adaptation, which learns to efficiently integrate switching-

Read source article
arXiv cs.CL (NLP)Research

Benchmarking Large Language Models for Grapheme-to-Phoneme Conversion: A Japanese Case Study

arXiv:2606.22009v1 Announce Type: new Abstract: Grapheme-to-phoneme (G2P) conversion is essential for controllable and robust text-to-speech, and large language models (LLMs), with broad linguistic knowledge, offer a promising approach. We benchmarked over 30 LLMs on Japanese G2P, comparing them with conventional morphological analyzers on 3000 manually annotated sentences. We evaluated two prompting strategies: a parse mode, where the LLM performs morphological analysis followed by rule-based k

Read source article
arXiv cs.AIResearch

The AI Evaluability Gap: The Missing Layer for Managing Risk and Sustaining Value

arXiv:2606.21015v1 Announce Type: new Abstract: Organizations deploying AI face two fundamental governance challenges: managing AI risk and sustaining AI value. Both depend on evidence whose sufficiency cannot be taken for granted. We call the shared underlying challenge the AI Evaluability Gap: the condition in which organizations lack sufficient evidence to support high-confidence governance decisions regarding either risk or value. We argue that this gap reflects a category error in current p

Read source article
arXiv cs.AIResearch

Mind the Noise: Sensitivity of Transformer-based Interaction-Aware Trajectory Prediction Models to Noisy Data

arXiv:2606.21344v1 Announce Type: new Abstract: Trajectory prediction allows autonomous vehicles to anticipate the future behavior of surrounding objects (or agents) and, accordingly, maximize the safety and efficiency of their driving. State-of-the-art Transformed-based interaction-aware trajectory prediction models, which rely on attention mechanisms to capture multi-agent interactions and maximize prediction accuracy, are commonly trained and evaluated on long-range high-quality datasets. The

Read source article
arXiv cs.AIResearch

Calibration Is Not Control: Why LLM-Agent Oversight Needs Intervention

arXiv:2606.21399v1 Announce Type: new Abstract: Runtime oversight for LLM agents is commonly framed as scalar risk prediction: estimate failure likelihood, confidence, or uncertainty, then intervene once the score crosses a threshold. We argue that this framing targets the wrong object for control. The relevant question is not how likely the agent is to fail if it continues, but whether an available intervention would improve the outcome. Two trajectory prefixes can have the same risk estimate w

Read source article
arXiv cs.AIResearch

Don't Blindly Trust It: How Unreliable Feedback Breaks Tool-Using LLM Agents

arXiv:2606.21409v1 Announce Type: new Abstract: Tool-augmented agents are typically evaluated by their gains under reliable external feedback. Yet these gains leave open a key counterfactual: when feedback is unreliable, would the agent be better off receiving no task evidence? We study this question with a controlled matched-loop comparison that fixes the agent loop, prompt, action space, and decoding, while varying only the returned observation: faithful, misleading, or absent. Across question

Read source article
arXiv cs.AIResearch

AutoRAS: Learning Robust Agentic Systems with Primitive Representations

arXiv:2606.21445v1 Announce Type: new Abstract: The automated design of agentic systems offers a promising pathway for scaling large language models (LLMs) beyond single-agent reasoning. While prior work has advanced task performance through handcrafted or automatically generated multi-agent workflows, robustness is often treated as an afterthought, leaving systems vulnerable to external adversaries and internal failures. We propose AutoRAS, a framework for the Automated design of Robust Agentic

Read source article
arXiv cs.AIResearch

Towards Transparent Mental Health Insights: An Explainable AI Model for Career-Related Depression and Anxiety Among University Students Using Structured Data

arXiv:2606.21474v1 Announce Type: new Abstract: Career anxiety and depression among university students present a growing challenge to mental health and academic achievement. This study proposes an Explainable AI (XAI) framework using multimodal data and Federated Learning (FL) to identify early indicators of career-related mental health problems in a privacy-preserving and culturally responsive manner. The framework combines structured behavioral data and facial emotion features from interview

Read source article
arXiv cs.AIResearch

Balancing Performance and Diversity in GRPO Autoregressive Text-to-Image Post-Training

arXiv:2606.21498v1 Announce Type: new Abstract: Autoregressive text-to-image (T2I) generation has recently advanced rapidly, yet aligning generated images with human preferences remains challenging. GRPO-style online reinforcement learning provides an effective framework; however, existing methods typically treat reference-policy divergence as fixed, despite its direct impact on policy optimization. We study this overlooked factor within a unified f-divergence framework, encompassing forward KL,

Read source article
arXiv cs.AIResearch

AI Alignment From Social Choice Perspectives

arXiv:2606.21550v1 Announce Type: new Abstract: Alignment from human feedback uses human judgments about model outputs to steer the behavior of language models after pretraining. When those judgments reflect conflicting views of desirable behavior, the learned objective becomes an aggregate determination of what the model should prefer. We survey recent work that has studied this aggregation problem through the lens of social choice theory. We illustrate how the social choice perspective helps i

Read source article
arXiv cs.AIResearch

Composing Verifiable Conceptual Models via Building Blocks: Towards Design-Time Verification of Agentic AI Workflows

arXiv:2606.21565v1 Announce Type: new Abstract: Agentic AI systems orchestrate multiple LLM-based agents through workflow architectures that coordinate decisions, tools, and external actions. While current platforms emphasize runtime safeguards, little support exists for verifying workflows during system design. From a Modeling \& Simulation perspective, this gap is analogous to composing conceptual models without verifying whether their building blocks interact coherently. We propose a design-t

Read source article
arXiv cs.AIResearch

Counsel: A Meta-Evaluation Dataset for Agentic Tasks

arXiv:2606.21627v1 Announce Type: new Abstract: As agentic systems tackle increasingly complex multi-step tasks, evaluating their trajectories presents a major bottleneck - human annotation of a single trajectory on popular agentic benchmarks can take hours, making it difficult to scale evaluations for measuring performance or curating training data. This has driven widespread reliance on automated approaches such as LLM-as-a-judge (LLMJ) to critique agents at the process and outcome-levels at s

Read source article
arXiv cs.AIResearch

ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks

arXiv:2606.21654v1 Announce Type: new Abstract: Computer use agents are evaluated almost exclusively on atomic desktop tasks, but realistic desktop work requires sustaining state across multiple objectives. We study this gap with ChainWorld, which composes atomic OSWorld tasks into long horizon desktop workloads through directional compatibility search while preserving the source evaluators. The resulting workload contains 347 chains of length two to four and compares two renderings of the same

Read source article
arXiv cs.AIResearch

Hallucination as Context Drift: Synchronization Protocols for Multi-Agent LLM Systems

arXiv:2606.21666v1 Announce Type: new Abstract: Multi-agent LLM systems routinely produce hallucinated outputs that cannot be explained by model deficiencies alone. A significant class of these failures arises not from model incapacity but from context drift: the divergence of internal knowledge states between concurrent agents. When agents enter a collaborative task with mismatched or stale representations of shared world state, their joint reasoning produces contradictions that manifest as hal

Read source article
arXiv cs.AIResearch

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents

arXiv:2606.21740v1 Announce Type: new Abstract: Translating natural-language planning intent into verified plans is a longstanding challenge: people communicate goals in language, while classical planners require formal PDDL specifications. Recent agentic frameworks bridge this gap by orchestrating a pool of specialized repair agents inside a verifier-checked refinement loop, but the orchestrator at the centre is itself a prompted frontier LLM, paying a frontier-LLM API call at every refinement

Read source article
arXiv cs.AIResearch

AgentCAT: Simulating Computerized Adaptive Testing via Multi-Agent Large Language Models

arXiv:2606.21832v1 Announce Type: new Abstract: Computerized Adaptive Testing (CAT), as a key technology for personalized education, aims to accurately assess examinee proficiency by retrieving exercises dynamically matching current ability estimates. However, existing CAT research is constrained by limitations of static offline data and isolated component optimization. Restricted by partial labels in offline logs, researchers degrade the dynamic assessment process into static sequence predictio

Read source article
arXiv cs.AIResearch

Measuring What Persists: Conditioning Mechanisms and a Geometric Framework for AI Agent Identity

arXiv:2606.21843v1 Announce Type: new Abstract: AI agents in long-context applications drift from their specified identity. Current methods detect this only after qualitative degradation is visible. We present a geometric framework for measuring identity structure using $\sqrt{\mathrm{JSD}}$ metric spaces and magnitude homology from enriched category theory, where identity is non-geodesic structure and drift is its relaxation toward the geodesic. Validated on a persistent AI agent, the framework

Read source article
arXiv cs.AIResearch

ForEx: A Formal Verification Framework for Explainable Reasoning in Logical Fallacy Detection and Annotation

arXiv:2606.21867v1 Announce Type: new Abstract: Current evaluations of Large Language Models (LLMs) on logical fallacy detection focus on predicted labels, but do not establish whether those labels are supported by the reasoning the models provide. We propose ForEx (Formal Verification for Explainable Reasoning), a framework that translates LLM-generated explanations into Lean4 and verifies whether the translated rationale is derivable under encoded premises, not the logical validity of the orig

Read source article
arXiv cs.AIResearch

AgentRiskBOM: A Risk-Scoping Security Bill of Materials for Agentic AI Systems

arXiv:2606.21877v1 Announce Type: new Abstract: Agentic AI systems retrieve private context, invoke tools, write files, call external services, coordinate with other agents, and may act without human approval. Existing bill of materials artifacts improve transparency for dependencies, model metadata, and training provenance, but leave an agentic transparency gap: capability opacity, the absence of a structured account of what a deployed agent can access, remember, change, delegate, and prove aft

Read source article
arXiv cs.AIResearch

Learning the ARTS of Search for Automated Discovery

arXiv:2606.21891v1 Announce Type: new Abstract: Scientific discovery can be formulated as an iterative search process over the space of hypotheses and experiments. Contemporary methods navigate this space using heuristics such as MCTS. These algorithms conflate the merit of a hypothesis with the quality of its experimental execution. A promising hypothesis with preliminary execution is therefore ranked below a modest hypothesis whose execution is refined. Moreover, prior methods prune the search

Read source article
arXiv cs.AIResearch

Holmes: Multimodal Agentic Diagnosis for Mixed-Language Mobile Crashes at Industrial Scale

arXiv:2606.21963v1 Announce Type: new Abstract: Diagnosing mobile crashes in ultra-large-scale industrial applications is a formidable challenge due to the sheer volume of code, the complexity of mixed-language environments, and the inability to reproduce failures locally. Traditional static analysis struggles with scalability, while existing LLM-based agents often rely on reproducible environments unavailable in post-mortem scenarios. We present Holmes, a multi-agent system that automates root

Read source article
arXiv cs.AIResearch

Human vs Machine Mathematical Difficulty on Project Euler: An Experimental Analysis

arXiv:2606.21972v1 Announce Type: new Abstract: We study how the effort and success probability of frontier AI systems scale with human difficulty on problems from Project Euler, an online platform of computational mathematics problems. Our dataset, from the MathArena benchmark, consists of 3840 attempts across 50 problems and 26 model configurations, with problem difficulty measured by the site's public human solve times. Motivated by a proposal of Timothy Gowers, we test a power-law relation $

Read source article
arXiv cs.CVResearch

Efficient Document Tampering Localization with Multi-Level Discrepancy Features and Unified DCT-Quantization Embedding

arXiv:2606.22285v1 Announce Type: new Abstract: Localizing document tampering is extremely challenging, as manipulations are crafted to appear visually consistent and often leave only subtle traces that are nearly invisible to the human eye. In prior work, evaluation has been largely dominated by synthetic benchmarks that closely match the training distribution, and methods have shown steady progress under this setting. However, these gains often translate poorly to human-made forgeries and to c

Read source article
arXiv cs.CVResearch

T-IMPACT: A Severity-Aware Benchmark for Contextual Image-Text Manipulation

arXiv:2606.22339v1 Announce Type: new Abstract: Recent advances in vision-language models and generative editing systems have made it increasingly easy to produce persuasive multimodal misinformation by altering images, text, or both jointly. However, existing datasets focus mainly on authenticity, out-of-context mismatch, or manipulation type, and rarely capture how strongly an edit changes the likely interpretation of a post. We introduce T-IMPACT, a first-release severity-aware benchmark for

Read source article
arXiv cs.CVResearch

Customizing Video Portraits via Identity-ActionDecoupling

arXiv:2606.22347v1 Announce Type: new Abstract: Identity-Preserving Text-to-Video Generation (IPT2V) seeks to synthesize a temporally coherent video from a reference image and a textual description, while simultaneously preserving the subject's identity and allowing fine-grained control over facial dynamics. Although recent methods such as ID-Animator and ConsisID inject identity features only at inference time, they ignored the ID-irrelevant information contained in Facial embedding, leading to

Read source article
arXiv cs.CVResearch

Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning

arXiv:2606.22394v1 Announce Type: new Abstract: Consistency distillation has significantly accelerated the inference of diffusion models. In this work, we reveal an intriguing asymmetry: while Logit-Normal sampling priors are highly efficacious for standard iterative generation, consistency distillation exhibits a distinctly different difficulty profile (e.g., U-shaped). We identify that the primary optimization bottlenecks reside at the boundary stages (initialization or final refinement) rathe

Read source article
arXiv cs.CVResearch

Multi-cancer detection using a computationally efficient CNN with transfer learning

arXiv:2606.22400v1 Announce Type: new Abstract: This study introduces a computationally efficient convolutional neural network (CNN) architecture enhanced with transfer learning for multi-cancer detection using biomedical images. The proposed lightweight CNN model is designed to reduce computational complexity while maintaining high classification performance, making it suitable for deployment in resource-constrained environments. We evaluate this approach on three distinct tumor datasets compri

Read source article
arXiv cs.CVResearch

Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition

arXiv:2606.22416v1 Announce Type: new Abstract: We address the problem of training on long-tailed data for video action recognition. We propose to augment the training set using a text-to-video generative model, conditioned on diverse text prompts grounded in action profiles and training exemplars. Our approach, called Gen2Balance, converts an imbalanced training set into a balanced combination of real and generated video clips. To effectively learn from such data, we employ a two-stage training

Read source article
arXiv cs.CVResearch

FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation

arXiv:2606.22424v1 Announce Type: new Abstract: Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural-language instructions in unseen scenes. While Large Models (LMs) have advanced VLN-CE, their performance remains severely degraded by real-world visual corruptions, a critical yet underexplored domain constraint. We introduce Temporal Conditional Flow Decorruptor (FlowDec), a novel image restoration framework tailored for LM-based VLN-CE. FlowDec in

Read source article
arXiv cs.CVResearch

Physically-guided Image Generation for Multi-Projection Mapping

arXiv:2606.22477v1 Announce Type: new Abstract: Projection Mapping (PM) enables seamless superimposition of digital content onto real-world 3D objects, serving as a fundamental technique for immersive visualization, digital twins, and interactive art. Although text-to-image diffusion models have greatly facilitated customized content creation, directly integrating them into practical PM pipelines remains challenging due to the mismatch between idealized 2D generation and physical constraints. To

Read source article
arXiv cs.CVResearch

Biological Sex Determination in Cadavers Using Deep Learning Algorithms from Computed Tomography Images of Pelvis and Skull

arXiv:2606.22515v1 Announce Type: new Abstract: Sexual identification of decomposed cadavers challenges traditional methods dependent on visual anthropological analysis. This study evaluates state-of-the-art deep learning (including YOLO26, YOLO11, ConvNeXt-Tiny, EfficientNetV2, ViT-B16, VGG16, and ResNet50) with transfer learning to automatically determine biological sex from forensic computed tomography (CT) scans. We analyzed 141 autopsied cadavers from the Forensic Medical Institute of Goi\^

Read source article
arXiv cs.CVResearch

Trajectory Forcing: Structure-First Generation with Controllable Semantic Trajectories

arXiv:2606.22527v1 Announce Type: new Abstract: Diffusion and flow-based generative models produce strong images, yet their controllability remains largely endpoint-centric: users specify conditions and receive final outputs, while the intermediate generative dynamics remain hidden. Recent methods have begun to exploit generation order and process decomposition to improve sample quality, but still treat intermediate states as internal computation rather than objects for interaction. We propose T

Read source article
arXiv cs.CVResearch

PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models

arXiv:2606.22540v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models provide a unified paradigm for robotic manipulation, yet their real-world deployment is often bottlenecked by execution efficiency. While existing efforts predominantly focus on compute-centric efficiency to reduce per-step inference latency, the intrinsic \textbf{policy efficiency} of these models remains largely unexplored. Policy efficiency is fundamentally affected by two factors, namely the effective executa

Read source article
arXiv cs.CVResearch

HiMatch-AD: DINOv3-driven Hierarchical Matching for Training-free Medical Anomaly Detection

arXiv:2606.22556v1 Announce Type: new Abstract: Anomaly detection is essential for medical image analysis, where pathological regions often appear as rare deviations from normal anatomical structures. While training-based methods have achieved promising performance, they require task-specific optimization and extensive normal data, which limits scalability across modalities and institutions. Training-free approaches offer greater flexibility by leveraging pretrained visual representations, yet e

Read source article