AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

34182 stories from 30+ sources, refreshed continuously.

arXiv cs.CL (NLP)Research

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories

arXiv:2606.22329v1 Announce Type: new Abstract: LLM-as-a-judge has become the dominant approach to scalable evaluation in NLP pipelines, yet judges themselves carry systematic biases that raw accuracy hides: they favor responses placed in slot A (position bias), they prefer longer responses regardless of quality (verbosity bias), and their reliability degrades sharply in lower-resource languages. We introduce BabelJudge, an open-source benchmark and reliability audit framework that measures all

Read source article
arXiv cs.CL (NLP)Research

How Does Research Evolve? Tracing Cross-Domain Trajectories in NLP, ML, and CV with Claim-Grounded Typed Citations

arXiv:2606.22342v1 Announce Type: new Abstract: How does research evolve, and what substrate would let us forecast where it goes next? Scientific progress is not simply a uniform accumulation of facts: ideas extend prior methods, address known limitations, realize proposed future directions, and sometimes dispute earlier claims. Existing citation graphs usually collapse these roles into a single homogeneous edge type, limiting how we can analyze scientific progress. We address this gap by propos

Read source article
arXiv cs.CL (NLP)Research

Curiosity as Linguistic Intervention: Using LLM Tutoring Dialogues to Influence Exploratory Learning Behavior

arXiv:2606.22349v1 Announce Type: new Abstract: Large Language Models (LLMs) provide a new opportunity to study how language shapes exploratory cognition because conversational strategies can be systematically manipulated at inference time. We introduce CURIOBOT, a framework that operationalizes Berlyne's collative variables, novelty, complexity, conflict, and uncertainty, as adaptive linguistic interventions for conversational tutoring. Across 270 tutoring conversations spanning multiple model

Read source article
arXiv cs.CL (NLP)Research

ORBIT: Training-Free Multi-Attribute Behavioral Steering via Orthogonal Subspace Rotation

arXiv:2606.22357v1 Announce Type: new Abstract: Language models are widely used in assistant settings, where controlling behavioral attributes is often essential. Activation steering modifies hidden-state representations at inference time, providing a lightweight, training-free mechanism that can be toggled at runtime. Existing methods, however, have focused primarily on steering a single attribute at a time. When multiple attributes must be controlled simultaneously, naive summation of per-attr

Read source article
arXiv cs.CL (NLP)Research

First-Token Broadcasters: Mechanistic Origins of Language Identity and Distributed Robustness in Transformers

arXiv:2606.22361v1 Announce Type: new Abstract: Why do multilingual language models sometimes generate in the wrong language, and why is this so hard to fix? We introduce Language Identity Head Ablation (LIHA), a causal intervention that zeros each attention head individually and measures the resulting language switch rate across a parallel dataset of 2,700 prompt-language pairs spanning seven languages. Applied to GPT-2, LIHA identifies a small set of first-token broadcaster heads - led by L6H1

Read source article
arXiv cs.CL (NLP)Research

Knowledge-Graph Grounding Helps LLMs Only for Out-of-Training Knowledge: A Controlled Study on Clinical Question Answering

arXiv:2606.22419v1 Announce Type: new Abstract: A recent Nature Medicine study reports that general-purpose frontier LLMs outperform specialized retrieval-augmented clinical tools on medical benchmarks, and that retrieval can hurt strong models. We ask the natural follow-up: does structured knowledge-graph (KG) grounding change this, and when does grounding help at all? We contribute two results. First, a reproduction: the study's headline HealthBench score (~88) is the Consensus variant, not fu

Read source article
arXiv cs.CL (NLP)Research

Words as Difference Makers: How Large Language Models Determine Causal Structure in Text

arXiv:2606.22430v1 Announce Type: new Abstract: Because large language models (LLMs) are impressively successful in predicting text, it appears that they must have access to a 'world model' representing causal and definitional structure. However, the dominant formalisms of modern causal inference -- Judea Pearl's interventionist approach and the Neyman-Rubin potential outcomes framework -- struggle to illuminate how LLMs learn causal structure. I resolve this puzzle by arguing that LLMs employ a

Read source article
arXiv cs.CL (NLP)Research

CASPER in the Machine: Insights into Character Variety in LLM-Generated Stories

arXiv:2606.22454v1 Announce Type: new Abstract: As LLM-generated text is increasingly used, especially in fictional domains, we explore how much LLM-generated stories differ from human-written stories. In this work, we focus on characters. We borrow definitions from narratology to analyze eight intricate dimensions of character, such as stylization and wholeness. These dimensions consider more than just basic characteristics. They assess how characters are portrayed within their stories. After a

Read source article
arXiv cs.AIResearch

IRumAI: Reinforcement Learning for Indian Rummy

arXiv:2606.21975v1 Announce Type: new Abstract: Despite its massive player base and complex hidden-information dynamics, Indian Rummy has received no reinforcement learning attention. Existing agents rely on combinatorial search, which is tactically strong but slow at inference. We present IRumAI, the first RL agent for the domain. IRumAI integrates Proximal Policy Optimization (PPO), meld-aware observation encoding, deadwood-driven reward shaping, and a dual-branch convolutional architecture. I

Read source article
arXiv cs.AIResearch

CFAgentBench: A Reproducible Environment and Benchmark for Autonomous Construction-Finance Agents

arXiv:2606.22000v1 Announce Type: new Abstract: We introduce CFAgentBench, a reproducible, self-hostable environment and benchmark for autonomous construction-finance agents: a CFO/controller-class agent operating across the real software stack a US construction finance team runs - ERP, project management, email, documents, pay applications, payroll, certified payroll, lien waivers, and bank/treasury portals. It contains 1,014 machine-gradeable task specifications across 8 domains and 77 familie

Read source article
arXiv cs.AIResearch

Nous: A Predictive World Model for Long-Term Agent Memory

arXiv:2606.22030v1 Announce Type: new Abstract: We present Nous, a novel agent memory architecture grounded in the principle that knowledge is prediction, not storage. Rather than persisting facts as database records, vector embeddings, or knowledge-graph triples, Nous maintains a predictive world model: a collection of categorical probability distributions, called dimensions, one per entity-attribute pair observed in conversation. Each incoming observation is scored by its information-theoretic

Read source article
arXiv cs.AIResearch

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice

arXiv:2606.22327v1 Announce Type: new Abstract: The explosive demand for interactive Large Language Model serving has highlighted the management of the Key-Value cache's dynamic memory footprint as a critical area for performance optimization in inference engines. Modern inference systems overwhelmingly rely on time-centric scheduling heuristics, such as Shortest Job First. However, their theoretical optimality is rooted in traditional schedule modeling, failing to capture the highly dynamic, 2D

Read source article
arXiv cs.AIResearch

Hypothesis-Driven Skill Optimization for LLM Agents

arXiv:2606.22330v1 Announce Type: new Abstract: External skills can improve action-oriented LLM agents without changing model weights, but persistent skill updates are risky when they are distilled from sparse or noisy trajectories. A plausible reflection may encode a useful procedure, a spurious shortcut, or a rule that the target executor cannot reliably follow. We propose Hypothesis-Driven Skill Optimization (HDSO), a train-free framework in which both the skill curator and the agent executor

Read source article
arXiv cs.AIResearch

Reference-Free Assessment of Physical Consistency in World Model-based Video Generation

arXiv:2606.22363v1 Announce Type: new Abstract: We introduce reference-free measures for evaluating the physical consistency of generated videos, combining relative and absolute approaches to assess fidelity. Although tools like WorldGym or WorldEval enable robotic simulation via video generation, physical fidelity gaps often prevent these environments from accurately reproducing real-world task success rates of VLA models. Unlike existing evaluation methods, which require costly human voting (E

Read source article
arXiv cs.AIResearch

ARIA: A Causal-Aware Framework for Rescuing LLM Reasoning in Trustworthy Materials Discovery

arXiv:2606.22375v1 Announce Type: new Abstract: Generative models have revolutionized the process of materials discovery, yet they often fail to satisfy underlying physical causality. Through an analysis of Large Language Models (LLMs) augmented with knowledge graphs derived from current literature, we uncover a phenomenon termed contextual tunneling, where models "over-anchor" on narrow, retrieved evidence while suppressing global physical reasoning. To address this problem, we introduce ARIA,

Read source article
arXiv cs.AIResearch

MetaPS: Adaptive Programmatic Strategy Selection for Market Agents

arXiv:2606.22385v1 Announce Type: new Abstract: No single market strategy always wins: momentum, mean reversion, risk control,and event-driven rules can each succeed or fail as market conditions change.Rather than asking large language models to directly generate market actions,we study an executable decision paradigm where an agent selects from a library of programmatic strategies, each implemented as a code module mapping market observations to actions.We propose \textbf{MetaPS}, a simulation-

Read source article
arXiv cs.AIResearch

PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems

arXiv:2606.22388v1 Announce Type: new Abstract: LLM agents increasingly operate in large tool ecosystems, where real-world tasks require discovering relevant tools, inferring implicit sub-goals, and adapting to dynamic environments over long horizons. However, existing benchmarks rarely evaluate planning under retrieval-limited tool visibility. To address this gap, we introduce PlanBench-XL, an interactive benchmark of 327 retail tasks over 1,665 tools that tests whether agents can iteratively r

Read source article
arXiv cs.AIResearch

Code Isn't Memory: A Structural Codebase Index Inside a Coding Agent

arXiv:2606.22417v1 Announce Type: new Abstract: Coding agents now interleave LLMs with retrieval over the working repository, and retrieval implementations vary widely across deployed harnesses. Inside a fixed coding-agent harness on a fixed model, does adding a structural codebase index actually change cost or resolve? We ran three arms (the harness with the index, the same harness without it, and an agentic-grep comparator) on SWE-PolyBench Verified and SWE-bench Pro with Claude Opus 4.7 held

Read source article
arXiv cs.AIResearch

SVGym (SciVerseGym): An Environment for Reinforcement Learning and Bayesian Optimization in Crystal Discovery

arXiv:2606.22425v1 Announce Type: new Abstract: Machine-learned interatomic potentials now enable efficient atomistic evaluation for interactive materials discovery, yet closed-loop crystal search methods remain fragmented across bespoke pipelines for editing, relaxation, scoring, constraints, and bookkeeping. We introduce SciVerseGym, a Gymnasium-compatible environment for sequential crystal discovery that frames crystal design as a Markov decision process. Agents observe an atomistic structure

Read source article
arXiv cs.AIResearch

A Differentiable Atari VCS:A Complex, Fully Known Ground Truth for Explainable AI

arXiv:2606.22447v1 Announce Type: new Abstract: Explanation requires ground truth: to verify an account of a system we must know its inner functioning-just what is missing where explainable AI (XAI) is most needed. Systems we can study fall into two camps. Simple, procedural one-decision trees, rule lists, sparse linear models-have a known but trivial mechanism, so explaining them tests nothing; genuinely complex ones-deep networks, real-world tasks-need XAI but have no ground-truth inner functi

Read source article
arXiv cs.AIResearch

PRIME: Evaluating Prompt Resolution Under Incompatible Instructions in LLMs

arXiv:2606.22470v1 Announce Type: new Abstract: Large language models (LLMs) often encounter conflicting prompts, although current instruction following benchmarks assess those meta-instructions in isolation, limiting the insights about how models process conflicting instructions. We introduce a framework \textit{PRIME}(\textit{Prompt Resolution under Incompatible Meta-Instructions Evaluation}) to analyze behavior of LLMs when provided with conflicting instructions. \textit{PRIME} purposefully p

Read source article
arXiv cs.AIResearch

VADAOrchestra: Neurosymbolic Orchestration of Adaptive Reasoning Workflows

arXiv:2606.22485v1 Announce Type: new Abstract: Decision-making in real-world settings rarely follows a fixed script. Instead, it unfolds as a dynamic reasoning process in which the appropriate course of action evolves as new context and data become available. Traditional Business Process Management systems provide rigor, determinism, and auditability, yet they generally struggle to adapt their execution at runtime. Conversely, agentic systems based on Large Language Models (LLMs) bring flexibil

Read source article
arXiv cs.AIResearch

Deep Learning-Based Sign Language Recognition from Videos and Cross-Lingual Translation to Indian Vernaculars

arXiv:2606.22494v1 Announce Type: new Abstract: Sign language is a primary mode of communication for the global deaf and hard-of-hearing community, yet automated tools that recognize sign gestures from video and translate them into natural language text remain limited, particularly for low-resource Indian languages. We present a two-stage deep learning pipeline that (i) classifies short sign language video clips into English word labels using a fine-tuned VideoMAE video transformer, and (ii) tra

Read source article
arXiv cs.AIResearch

Grounded Scaling: Why Agentic AI Needs Deterministic Environments

arXiv:2606.22495v1 Announce Type: new Abstract: Long-chain agent execution fails exponentially in environments designed for human tolerance: with per-step determinism $\delta < 1$, $k$-step chain success degrades as $\delta^k$. The AGI-to-ASI scaling debate (Genewein et al., 2026) has so far framed progress as a race between compute growth and a list of frictions (data wall, abstraction barrier, embodied bottleneck, multi-agent trust); we argue that environment determinism is a complementary bin

Read source article
arXiv cs.AIResearch

Imagine to Ensure Safety in Hierarchical Reinforcement Learning

arXiv:2606.22509v1 Announce Type: new Abstract: This work investigates the safe exploration problem in reinforcement learning, where an agent must maximize cumulative performance while simultaneously satisfying safety constraints. This challenge becomes even more pronounced in long-horizon tasks, where existing safe methods face fundamental limitations due to compounding estimation errors and restricted exploration capabilities. To address this problem, we propose a method that combines a learna

Read source article
arXiv cs.AIResearch

Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents

arXiv:2606.22528v1 Announce Type: new Abstract: Modern LLM agents increasingly rely on context compaction, summarization, or eviction to keep long-running sessions within a token budget. We show that this context-management layer is a safety-critical failure surface: in-context governance constraints that agents reliably obey while visible can be silently removed by compaction, causing the same agent to perform prohibited tool actions later in the session. We call this failure mode Governance De

Read source article
arXiv cs.AIResearch

MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop

arXiv:2606.22557v1 Announce Type: new Abstract: Computer use agents (CUAs) have advanced rapidly in desktop automation, and a growing number of users deploy CUAs such as OpenClaw on Mac Mini for always-on automation. However, existing benchmarks, including those for macOS, evaluate agents without framework augmentation and rely on binary evaluation. As a result, they fail to capture both the framework capabilities leveraged by modern CUAs and the partial progress on long-horizon, multi-applicati

Read source article
arXiv cs.AIResearch

Text2DSL: LLM-Based Code Generation for Domain-Specific Languages

arXiv:2606.22586v1 Announce Type: new Abstract: Domain-specific languages (DSLs) are widely used for managing operating system security policies, yet manually authoring rules in such languages demands high expertise and is error-prone. This paper formalises the task of automatic DSL code generation from natural language descriptions - Text2DSL - as a distinct problem class, separate from Text-to-SQL and general-purpose code generation. We introduce the PolkitBench dataset comprising 4,204 verifi

Read source article
arXiv cs.CVResearch

Generative Relightable Avatars

arXiv:2606.22718v1 Announce Type: new Abstract: We present Generative Relightable Avatars (GRA), a person-specific method for photorealistic free-view rendering and environment-map relighting of full-body humans. We postulate that modeling fine-grained appearance details is inherently a one-to-many problem that can benefit from a generative formulation. In contrast to fully regressive relightable avatar methods, GRA follows a hybrid approach that combines controllable, physics-grounded relightin

Read source article
arXiv cs.CVResearch

RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation

arXiv:2606.22749v1 Announce Type: new Abstract: Pre-trained Vision Foundation Models (VFMs) have become central to modern computer vision due to their powerful semantic representations and strong generalization ability. However, their patchified or pooled outputs are inherently low-resolution, limiting their effectiveness in tasks requiring fine-grained, pixel-level reasoning. Existing feature upsampling approaches either degrade semantic fidelity or rely on VFM-specific retraining and heavy arc

Read source article
arXiv cs.CVResearch

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations

arXiv:2606.22766v1 Announce Type: new Abstract: Audio Description aims to generate concise narrations of essential visual content in audio-visual media for blind and low-vision audiences. Existing methods either rely on prompting off-the-shelf multimodal models, which often mismatch AD style, or partially optimize training-based systems with next-token prediction, which under-explores model capacity and biases generation toward generic expressions. We present READ, the first reinforcement-learni

Read source article
arXiv cs.CVResearch

Visual Geometry Transformer in the Wild: Distractor-Free 3D Reconstruction

arXiv:2606.22787v1 Announce Type: new Abstract: Current end-to-end multi-view 3D reconstruction methods achieve impressive results, but rely on a restrictive static assumption: the scenes is entire distractor-free with perfect cross-view geometry. This reliance on idealized inputs causes even the most advanced methods to fail in real-world settings, where transient distractors and occlusions present. To address this, we propose Visual Geometry Transformer in the Wild (VGTW), an end-to-end framew

Read source article
arXiv cs.CVResearch

Policy-as-Data: Learning Generalizable HOI Diffusion Models from Simulated Physics

arXiv:2606.22806v1 Announce Type: new Abstract: Synthesizing realistic Human-Object Interactions (HOI) is critical for creating embodied avatars and functional virtual environments. However, current data-driven approaches primarily rely on motion capture datasets, which are expensive to scale and limited in functional diversity. Models trained with these datasets fail to generalize to unseen objects and maintain physical consistency over long horizons. In this paper, we propose a novel framework

Read source article
arXiv cs.CVResearch

Homographic Navigation: Geometry-Driven Camera Guidance for Deterministic Planar Capture

arXiv:2606.22834v1 Announce Type: new Abstract: We present homographic navigation, a geometry-centric framework for guiding camera acquisition toward precise capture of planar regions. Rather than treating homography as an output, we use it as an organizing variable that unifies learning, alignment, and evaluation. From a single annotated reference image, we generate unlimited synthetic training data via homographic augmentation and train a single-shot model for joint recognition and localizatio

Read source article
arXiv cs.CVResearch

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning

arXiv:2606.22873v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed in consumer, medical, financial, and enterprise applications. This broad deployment expands the safety surface: risks can arise from multimodal question answering, assistant responses, and cross-modal composition, while moderation policies may vary across products, regions, and deployment stages. Most existing guardrails either rely on fixed taxonomies or target only a narrow set of interactio

Read source article
arXiv cs.CVResearch

FedOT: Ownership Verification and Leakage Tracing via Watermarks for Federated LDMs

arXiv:2606.22875v1 Announce Type: new Abstract: Training Latent Diffusion Models (LDMs) within Federated Learning (FL) has attracted increasing attention due to its ability to combine the powerful generative capacity of LDMs with the privacy-preserving properties of FL. However, FL requires sharing the global model with multiple participants, which risks unauthorized model distribution or resale by malicious clients. While an intuitive approach is to adopt existing VAE-based watermarking techniq

Read source article
arXiv cs.CVResearch

PHOEBI: An Open-World Benchmark for Bacterial Identification in Phase-Contrast Microscopy

arXiv:2606.22890v1 Announce Type: new Abstract: Optical microscopy enables rapid, label-free imaging of live bacteria and is the standard instrument for species identification across clinical, environmental, and industrial microbiology. Yet field samples are routinely polymicrobial and may contain organisms that were never seen during system training, and no computer-vision benchmark tests multi-label species identification from phase-contrast microscopy (PCM) of such mixtures. We introduce Phas

Read source article
arXiv cs.CVResearch

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars

arXiv:2606.22905v1 Announce Type: new Abstract: Recent diffusion-based models have enabled realistic audio-driven avatar generation in real-time streaming. However, existing approaches struggle to maintain visual temporal consistency and fail to explicitly perceive user intent in complex interactive streaming scenarios. To address these challenges, we propose InteractiveAvatar, a real-time infinite-streaming video generation framework that supports visually consistent avatar video generation and

Read source article
arXiv cs.CVResearch

Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation

arXiv:2606.22918v1 Announce Type: new Abstract: Maintaining physical consistency in video generators and world models increasingly relies on vision-language models (VLMs) as automated judges that provide reward signals, ranking decisions, and data-filtering criteria. Yet VLMs differ substantially in training data and architecture, encoding physical phenomena through distinct internal representations. A single global evaluation schema therefore gives every VLM the same axes of competence, regardl

Read source article
arXiv cs.CVResearch

MythraGen: Two-Stage Retrieval Augmented Art Generation Framework

arXiv:2606.22924v1 Announce Type: new Abstract: Text-to-image generation has seen rapid advancements, especially with the development of generative models. However, challenges remain in achieving high-quality, contextually accurate image outputs that faithfully match the provided textual descriptions, especially in artistic generation. In this paper, we present a simple yet efficient retrieval augmented generation framework, namely MythraGen, for text-to-artistic image generation by integrating

Read source article