AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

34890 stories from 30+ sources, refreshed continuously.

arXiv cs.AIResearch

How Should Agents Read Demonstrations? Hierarchical Structure Beats Flat Action Logs

arXiv:2606.20978v1 Announce Type: new Abstract: Programming by Demonstration (PbD) offers a human-centered way to author procedural knowledge for LLM agents: users communicate what they want by showing rather than by writing prompts or code, making agent authoring accessible to non-programmers. The natural output of a PbD recording is a flat action log, but how this log is organized before being passed to the agent is an open design question with significant consequences for plan quality. We pro

Read source article
arXiv cs.AIResearch

BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge Discovery

arXiv:2606.20997v1 Announce Type: new Abstract: Biomedical researchers increasingly use AI-generated analyses and reports to interpret protein-level signals, but static outputs are often insufficient for research decision-making, where users need to inspect evidence, assess uncertainty, compare mechanisms, and refine hypotheses. We present \textsc{BioInsight}, a multi-agent system that moves from static biomedical report generation to interactive evidence-centered interactive interface generatio

Read source article
arXiv cs.AIResearch

Building Agent Harnesses for Scientific Curation from Multimodal Sources

arXiv:2606.21005v1 Announce Type: new Abstract: Scientific discovery workflows often depend on structured curation from the literature. This is difficult for current agents because the key evidence is scattered across long text, dense tables, and figures, and the final records often require reasoning across multiple evidence fragments rather than copying a single span. We study scientific curation from multimodal sources and introduce Beaver, an agent harness that extracts structured information

Read source article
arXiv cs.AIResearch

Agentic Time Machine as an Infrastructure for Future-Event Forecasting

arXiv:2606.21013v1 Announce Type: new Abstract: Forecasting future events is a critical challenge for large language model (LLM) agents, spanning domains from elections and monetary policy to financial markets. However, evaluating progress on this task presents a fundamental trade-off between efficiency and environment fidelity. While live evaluation benchmarks suffer from an inherently slow feedback loop, existing retrospective replays typically restrict agents to static, pre-frozen databases t

Read source article
arXiv cs.AIResearch

Negative Knowledge as Failure-aware Shared Memory for AutoResearch

arXiv:2606.21024v1 Announce Type: new Abstract: AI-assisted research systems generate many failed attempts, but those failures rarely become a durable, shared knowledge asset. We propose a negative knowledge memory layer: a curator agent converts each failed attempt into a bounded, typed record in a shared bank, and a downstream research agent explicitly adopts or rejects those records before proposing its next experiment. We evaluate this layer in two settings: same-task retry on ScienceAgentBe

Read source article
arXiv cs.AIResearch

Coherence Under Commitment: Probing Generalization and Vacuous Memorization in LLM Logical Reasoning

arXiv:2606.21083v1 Announce Type: new Abstract: Large language models (LLMs) deployed for logical reasoning in knowledge-intensive domains exhibit a subtle but critical failure: coherence can be vacuously achieved through systematic abstention. A model that withholds commitment to either entailment or refutation satisfies negation consistency while providing no utility. We introduce Coherence Under Commitment (CUC), a dual-query evaluation paradigm that jointly measures consistency and decisiven

Read source article
arXiv cs.AIResearch

Repeated post-training is not Self-improving: Diagnosing Scientific Amnesia in Continual DPO Pipelines

arXiv:2606.21089v1 Announce Type: new Abstract: Industrial LLM teams often ship behavior updates by repeatedly DPO-training a base model on sequences of related preference-data campaigns. The dominant failure mode in this regime is not always classical catastrophic forgetting: a pipeline may preserve previously learned behaviors while still failing to accumulate reusable methodological knowledge about how to train the next campaign. We call this failure mode scientific amnesia. This paper turns

Read source article
arXiv cs.AIResearch

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training

arXiv:2606.21090v1 Announce Type: new Abstract: Self-improvement can self-regress. In REINFORCE post-training for code, a model can quickly improve on its optimized metric and then collapse within the same training campaign. We study this in a controlled multi-seed testbed using Qwen-2.5-3B and Qwen-2.5-7B, trained on competitive-programming tasks with binary CodeGrader reward across 10 sequential 20-step campaigns. Across campaigns, pass@1 shows a robust rise-then-collapse pattern: it peaks wit

Read source article
arXiv cs.AIResearch

Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models

arXiv:2606.21121v1 Announce Type: new Abstract: Large language models can produce confident but protocol-invalid answers in domains where procedural compliance is critical. This paper presents Answer Engineering, a deterministic runtime and authoring layer that applies localized rule-guided interventions to the visible reasoning trajectory during standard autoregressive generation, without retraining, modifying model weights, or performing global search. The method is evaluated on a controlled c

Read source article
arXiv cs.AIResearch

PulseCX: Breaking the Closed-World Assumption in Real-Time CX

arXiv:2606.21124v1 Announce Type: new Abstract: Conversational AI agents in Customer Experience (CX) typically suffer from a Closed-World Constraint, ignoring high-velocity external shifts like viral trends or outages. Ad-hoc web search attempts to bridge this gap but often introduce prohibitive latency and context poisoning. We introduce PulseCX, a framework that decouples knowledge acquisition from consumption. Adopting a structure-first paradigm, PulseCX employs an asynchronous agent to linea

Read source article
arXiv cs.AIResearch

Learning Burst-Aware Early Warning Models for Capacity Stress under AI Workload Surges in Hyperscale Data Centers

arXiv:2606.21130v1 Announce Type: new Abstract: The rapid growth of large-scale AI workloads, particularly Large Language Model (LLM) training and inference, is fundamentally reshaping the operational dynamics of hyperscale data centers. Unlike traditional cloud workloads, AI-driven jobs exhibit bursty, high-intensity, and rapidly shifting resource demands, often leading to sudden capacity stress that cannot be effectively handled by reactive threshold-based mechanisms. In this paper, we propose

Read source article
arXiv cs.AIResearch

Trip+: Benchmarking Agents in Personalized Interactive Travel Planning

arXiv:2606.21169v1 Announce Type: new Abstract: Interactive travel planning has become a popular use case for language models. Agents are deployed to manage evolving preferences and unexpected disruptions over multiple turns. Such settings require models to make complex, profile-conditioned planning decisions. However, existing benchmarks often evaluate feasibility, personalization, or interaction in relatively isolated settings. We therefore introduce Trip+ to measure the ability of agents to p

Read source article
arXiv cs.AIResearch

Whistleblowing and the machine -- towards a considered position

arXiv:2606.21201v1 Announce Type: new Abstract: Artificial intelligent agents and autonomous systems are embedded in our environments. They are both a commercial product and a personal tool that generates a lot of data and can draw conclusions from it: machines generate and keep secrets. But should machines protect all secrets? It has been shown that artificial agents are able to whistleblow and it has been argued that digital multi-agent environments should allow for agents in them to whistlebl

Read source article
arXiv cs.AIResearch

ARCO: Adaptive Rubric with Co-Evolution for Multi-Step LLM-Based Agents

arXiv:2606.21262v1 Announce Type: new Abstract: Reinforcement learning for multi-step LLM agents often relies on scalar rewards that indicate success but cannot explain why a trajectory is good or bad. Rubric-based rewards improve interpretability through natural-language criteria, but existing methods score at the trajectory level and freeze the scorer behind a closed-source judge, leaving step-level credit assignment unresolved and the judge itself static. We propose ARCO (Adaptive Rubric CO-e

Read source article
arXiv cs.AIResearch

Towards Dys-XAI: Influence-Based Explanations for Dysarthria Severity Assessment

arXiv:2606.21306v1 Announce Type: new Abstract: Dysarthria severity assessment is essential for therapy planning and longitudinal monitoring, yet manual perceptual rating is time-consuming and variable across clinicians. Although deep learning models achieve strong performance, their black-box nature limits clinical adoption. Existing speech explainability methods typically provide acoustic feature importance scores that are difficult for end-users to interpret. We propose an influence-based, in

Read source article
arXiv cs.AIResearch

Social World Model for Lifelong Social Intelligence

arXiv:2606.21315v1 Announce Type: new Abstract: Social intelligence is a core competency for language agents, yet current research primarily focuses on static capability evaluation rather than how these skills are continuously shaped and accumulated. This gap calls for a shift toward sustainable learning paradigms. Currently, two methodological pain points exist: social interaction trajectories lack unified structured representations to form iterable learning signals, and capability improvement

Read source article
arXiv cs.LGResearch

MedTS-TTT: Test-Time Training for Medical Time Series Classification

arXiv:2606.21329v1 Announce Type: new Abstract: Medical time series (MedTS) signals such as electroencephalography (EEG) and electrocardiography (ECG) support many clinical applications. However, substantial subject-level heterogeneity often induces subject-level distribution shift, causing a fixed parameter set to generalize poorly to unseen individuals. Compared with domain adaptation methods that often depend on extra adaptation components or target-batch statistics, Test-Time Training (TTT)

Read source article
arXiv cs.LGResearch

A Reward-Petri-Net Interpretation of Temporal Behavior Trees

arXiv:2606.21350v1 Announce Type: new Abstract: This paper introduces an interpretation of Temporal Behavior Trees (TBTs) as Reward-Petri-Nets (RPNs) for reinforcement learning (RL). Designing reward functions for complex, long-horizon robotic tasks is notoriously difficult, especially when tasks have hierarchical structure and temporal constraints. TBTs extend conventional behavior trees (BTs) used in robotic applications by incorporating temporal properties into their leaf nodes. This allows T

Read source article
arXiv cs.LGResearch

Predictive Repair Management Using a Multi-Head Attention Transformer and Online Learning

arXiv:2606.21364v1 Announce Type: new Abstract: Accurate prediction of repair duration is an important challenge in product maintenance due to its implications for resource allocation, customer satisfaction, and operational performance. This study aims to develop a deep learning framework to help fleet repair shops accurately categorize repair time given product historical data. The study uses an automobile repair and maintenance dataset and creates an end-to-end predictive framework by employin

Read source article
arXiv cs.CVResearch

Structural Assessment for Understanding and Guiding Dataset Distillation in Discrete Token Space

arXiv:2606.21705v1 Announce Type: new Abstract: Dataset distillation (DD) has proven to reduce training cost while preserving accuracy. While promising, the factors that make one distilled dataset more effective than another remain poorly understood. In this work, we investigate this question through the lens of discrete visual tokenizers. Whereas many prior DD efforts emphasize matching global data distributions, we suggest that the effectiveness depends on which semantic concepts are captured

Read source article
arXiv cs.CVResearch

Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization

arXiv:2606.21736v1 Announce Type: new Abstract: Single domain generalization (SDG) aims to learn a robust model, which could perform well on many unseen domains while there is only one single domain available for training. One of the promising directions for achieving single-domain generalization is to generate out-of-domain (OOD) training data through data augmentation or image generation. Given the rapid advancements in AI-generated content (AIGC), this paper is the first to propose leveraging

Read source article
arXiv cs.CVResearch

Quantile Adaptive Temperature Scaling for Confidence Calibration

arXiv:2606.21749v1 Announce Type: new Abstract: Deep neural networks often produce poorly calibrated confidence estimates, overstating their certainty even when predictions are incorrect. Temperature Scaling remains the most widely used posthoc calibration method due to its simplicity and effectiveness, yet its global, uniform rescaling of logits fails to correct the highly heterogeneous structure of miscalibration observed across the confidence spectrum. In particular, the largest correctness c

Read source article
arXiv cs.CVResearch

Motion-Aware Reinforcement Learning For Object Localization

arXiv:2606.21764v1 Announce Type: new Abstract: We present MARLNet (Motion-Aware Reinforcement Learning Network), a PPO-based bounding-box refinement agent that incorporates a constant-velocity motion prior into the observation state and an action smoothness penalty into the reward function. The agent operates on 268-dimensional observations encoding the current proposal, a kinematic prediction, the previous action, and a 256-dimensional EfficientNet-B0 crop feature, and learns a five-dimensiona

Read source article
arXiv cs.CVResearch

RAPID: A Reproducible Multi-Agent Pipeline for Interpretable Disaster Damage Assessment from Satellite and Street-View Imagery

arXiv:2606.21819v1 Announce Type: new Abstract: Due to the increasing frequency and intensity of extreme climate events, there is a clear demand for intelligent, scalable, and autonomous approaches to disaster damage assessment. Existing methods, largely based on supervised learning and task-specific fine-tuning, struggle to generalize under domain shifts, long-tailed data distributions, and heterogeneous geospatial data sources, especially in disaster scenarios. They also often lack the ability

Read source article
arXiv cs.CVResearch

Zero-Shot Vision-Language Models for Classroom Engagement Recognition: A Benchmark Study of Prompt Sensitivity and Cross-Dataset Generalization

arXiv:2606.21861v1 Announce Type: new Abstract: Automated classroom engagement recognition holds substantial promise for scalable learning analytics, yet the suitability of modern Vision-Language Models (VLMs) for this task under zero-shot conditions remains largely unexplored. We present a systematic benchmark that evaluates five widely-used VLMs: CLIP, BLIP-VQA, GPT-4o, LLaVA-1.5-7B, and Qwen2.5VL-7B-Instruct across two complementary educational datasets: DAiSEE, an individual-student video da

Read source article
arXiv cs.CVResearch

Prompt-Calibrated SAM 3 for Open-Vocabulary Remote Sensing Semantic Segmentation

arXiv:2606.21863v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation (OVSS) in remote sensing images aims to segment categories beyond a fixed label space. Recent SAM 3-based methods provide a promising training-free foundation, yet three key issues remain: (1) a single class-name prompt lacks sufficient semantic coverage for complex remote sensing categories; (2) expanding each category into multiple prompts introduces redundant online text encoding; and (3) directly aggregatin

Read source article
arXiv cs.CVResearch

Fidelity- and Perception-Aware Local Implicit Attention for Arbitrary-Scale Image Super-Resolution

arXiv:2606.21910v1 Announce Type: new Abstract: Arbitrary-scale image super-resolution (ASISR) aims to reconstruct high-resolution images from low-resolution inputs over a continuous range of upscaling factors. While traditional pixel-regression approaches often produce overly smooth results that lack realistic details, recent diffusion methods can produce sharper and more realistic textures. However, these diffusion techniques frequently introduce the risk of structural hallucinations. To addre

Read source article
arXiv cs.CVResearch

Artic-O: End-to-End Articulated Object Reconstruction via Latent Geometry Learning

arXiv:2606.21938v1 Announce Type: new Abstract: Reconstructing articulated objects from sparse images requires recovering complete geometry, movable parts, and motion parameters. Recent methods typically separate geometry reconstruction, part reasoning, and articulation estimation into different stages. This separation can weaken consistency between shape, active parts, and motion, while also incurring substantial inference cost. We introduce Artic-O, an end-to-end, feed-forward framework for ar

Read source article
arXiv cs.CVResearch

Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG

arXiv:2606.21968v1 Announce Type: new Abstract: Vision-Language Models (VLMs) struggle as query-relevant objects become smaller. To address this, recent training-free approaches dynamically retrieve and zoom into local image regions. However, we show that indiscriminately applying retrieval ignores a critical vulnerability: the resolution-context trade-off. Patch-based zooming recovers details for small targets, but can split large objects and destroy global spatial context; attention-based retr

Read source article
arXiv cs.CVResearch

CoDMD: Copula-aware Distribution Matching Distillation for Fast Video Generation

arXiv:2606.21982v1 Announce Type: new Abstract: Few-step distillation for video diffusion models has attracted significant attention, driven by the urgent demand for efficient deployment in real-world scenarios. However, Distribution Matching Distillation (DMD), a leading paradigm, tends to degrade under limited NFE budgets, manifesting in video generation as layout instability, oversaturation, and broken motion dynamics. We trace this failure to a structural limitation: standard DMD is an intra

Read source article
arXiv cs.CVResearch

One-Shot Data Selection for Medical Image Classification via Graph Coverage

arXiv:2606.22002v1 Announce Type: new Abstract: Training medical image classifiers on entire datasets is wasteful when annotation budgets are limited: not all samples contribute equally, yet acquiring expert labels is expensive. Active learning reduces annotation cost through iterative querying, but assumes repeated access to an oracle and requires multiple rounds of model training. One-shot geometry-based methods such as facility location avoid retraining but operate on pairwise distances that

Read source article
arXiv cs.CVResearch

IDAG-Edit: Multi-Object Video Editing via Instance-Decoupled Attention and Guidance

arXiv:2606.22042v1 Announce Type: new Abstract: Diffusion-based video editing has made significant progress; however, achieving precise and temporally consistent object-level control, especially in multi-object scenarios, remains challenging due to attention leakage, identity drift, and unstable temporal dynamics. In this work, we propose IDAGEdit, a training-free framework for fine-grained multi-object video editing with strong temporal consistency. The framework adopts Layout-guided Attention

Read source article
arXiv cs.CVResearch

Learning Cross-View Semantic Priors for Single-Reference Unseen Object Pose Estimation

arXiv:2606.22076v1 Announce Type: new Abstract: Single-reference unseen object 6D pose estimation reduces object onboarding by estimating poses of arbitrary novel objects from only one reference view. Recent correspondence-based pipelines have achieved robust performance with vision foundation model (VFM) features. However, they typically treat these features as intra-view descriptors, leaving dense visual-semantic cues, including appearance, structure, and context, insufficiently exchanged acro

Read source article
arXiv cs.CVResearch

Morphology-Aware Multimodal Representation Learning for Insect Phylogenetic Reconstruction

arXiv:2606.22077v1 Announce Type: new Abstract: Morphological traits provide important evidence for phylogenetic reconstruction and evolutionary relationship analysis. Recent image-based approaches have introduced deep learning, particularly convolutional models, to derive morphological features from specimen images, but these methods generally rely on single-modality visual representations and do not explicitly incorporate morphological semantics. This study proposes a morphology-aware multimod

Read source article
arXiv cs.CVResearch

Accurate identification and measurement of the precipitate area by two-stage deep neural networks in novel chromium-based alloys

arXiv:2606.22112v1 Announce Type: new Abstract: The performance of advanced materials for extreme environments is underpinned by their microstructure, including the size and distribution of reinforcing phases. Chromium-based superalloys are a recently proposed alternative to conventional face-centred-cubic superalloys for high-temperature applications, such as Concentrated Solar Power, and their development requires efficient measurement of precipitate volume fraction and size distribution from

Read source article
arXiv cs.CVResearch

Surgical Anatomy Recognition with Context Learning using Foundation Representations

arXiv:2606.22124v1 Announce Type: new Abstract: Accurate recognition of anatomical structures is essential for safe and effective minimally invasive surgery (MIS), yet it remains underexplored in surgical computer vision due to limited annotated data and methods tailored primarily to natural scenes. In this work, we present a combined dataset and model framework to advance anatomy-aware perception in MIS. First, we introduce ATLAS-120k, a large-scale clip-level semantic segmentation dataset comp

Read source article
arXiv cs.CVResearch

Feed-forward Motion In-betweening for Any 4D

arXiv:2606.22131v1 Announce Type: new Abstract: 4D dynamics (3D geometry evolving over time) is a fundamental representation of the physical world and plays a crucial role in world modeling (e.g., animation and games). Owing to the scarcity of large-scale, long-horizon 4D mesh data with arbitrary shapes, early text-to-4D methods rely on distillation or test-time optimization from video diffusion priors, making inference prohibitively slow. Recent feed-forward generators greatly reduce inference

Read source article
arXiv cs.CVResearch

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training

arXiv:2606.22158v1 Announce Type: new Abstract: Achieving human-like reasoning in Vision-Language Models (VLMs) remains a long-standing challenge. Recent approaches leverage Chain-of-Thought (CoT) rationales generated by human annotators or proprietary models to improve reasoning, which is costly and difficult to scale. Self-training offers a promising alternative by using models own outputs as supervision. However, existing methods often suffer from visual hallucinations -- where rationales des

Read source article
arXiv cs.AIResearch

On the Identifiability of User Adaptation in Co-Adaptive Neural Interfaces

arXiv:2606.20569v1 Announce Type: new Abstract: We analyze identifiability in co-adaptive human-machine systems. We show that closed-loop encoder estimates do not uniquely identify user adaptation, but instead reflect properties of the joint system. We discuss implications for interpreting behavioral adaptation and propose conditions for identification.

Read source article
arXiv cs.AIResearch

Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies

arXiv:2606.20599v1 Announce Type: new Abstract: Tree of Thought (ToT) search has become a promising direction for improving the reasoning capabilities of large language models, but deploying these methods in practice raises a question that has received little systematic attention: how do different search strategies behave under varying compute budgets, model sizes, and problem difficulties? In this work, we evaluate two representative ToT methods; DPTS, a Monte Carlo tree search based approach,

Read source article