AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

34680 stories from 30+ sources, refreshed continuously.

arXiv cs.AIResearch

Agent Behavior Mining: Generative AI Agent Governance in Business Processes

arXiv:2606.20669v1 Announce Type: new Abstract: As organizations increasingly deploy generative AI agents to automate business processes, they face a governance dilemma: although these agents can increase operational flexibility, their non-deterministic nature challenges the control and standardization that Business Process Management seeks to enforce. This paper addresses this \emph{invisible autonomy risk} by introducing \emph{Agent Behavior Mining}, a governance capability that enables the ap

Read source article
arXiv cs.AIResearch

Democratizing and accelerating AI-driven pathology research through agentic intelligence

arXiv:2606.20677v1 Announce Type: new Abstract: Computational pathology has advanced rapidly with the emergence of foundation models, yet widespread adoption remains limited by substantial technical complexity and programming requirements. Here we present PathLab, an autonomous agentic framework that translates natural-language research objectives into executable and validated computational pathology workflows through the structured composition of domain-specific skills and tools. By organizing

Read source article
arXiv cs.AIResearch

From Question Answering to Task Completion: A Survey on Agent System and Harness Design

arXiv:2606.20683v1 Announce Type: new Abstract: LLM-based agents mark a shift from passive question answering to active task completion: they perceive environments, invoke tools, maintain state, and act over extended horizons. As agent systems have evolved from prompt engineering to workflows and context engineering, harness engineering, and agent-native training with co-evolution, a central question has become increasingly important: where does the bottleneck in agent performance reside, in the

Read source article
arXiv cs.AIResearch

Artificial Intelligence as Monism: Ontological, Organisational, and Methodological Implications

arXiv:2606.20704v1 Announce Type: new Abstract: This paper argues that Artificial Intelligence should be understood as a form of monism: a unified substance that cannot be decomposed into separate elements such as data, algorithms, or technical architectures. Drawing from philosophical traditions of monism, dualism, and holism, the paper contends that AI is not merely a collection of components but a single, indivisible essence reflecting the phenomena it replicates. Treating AI as monism has de

Read source article
arXiv cs.AIResearch

Simulated Customers Never Walk Away: Decision Fidelity of LLM User Simulators Measured Against Real Purchase Outcomes

arXiv:2606.20708v1 Announce Type: new Abstract: LLM-as-user-simulation has become core infrastructure for conversational AI: agent benchmarks (tau-bench), training pipelines, and a growing body of fidelity studies all rely on LLMs role-playing the human side of dialogue. Existing frameworks measure communicative fidelity -- whether simulators talk like humans -- against ground truth from paid participants role-playing assigned goals. We argue this has a structural blind spot: when the goal is as

Read source article
arXiv cs.AIResearch

FairTutor: Equity-Aware Pedagogical LLM Routing for Budget-Constrained AI Tutoring

arXiv:2606.20713v1 Announce Type: new Abstract: Generative AI tutors provide real-time, personalized learning support, but also create a new education inequity: students with access to premium AI services may receive clearer explanations, more personalized guidance, and better scaffolding than students limited to free or low-cost services. To address this challenge, we propose FairTutor, an equity-aware model-routing framework that achieves cost-effective AI tutoring via pedagogically motivated

Read source article
arXiv cs.AIResearch

Escape from Delusional Echo Trap: Symmetry Breaking, Stochastic Dynamics and Mathematical Mitigation Strategies for Algorithmic Sycophancy

arXiv:2606.20718v1 Announce Type: new Abstract: We propose a rigorous and systematic mathematical framework for tracking the cognitive trajectories of a user, in the context of algorithmic sycophancy and AI-driven delusional spiraling. Using tools from dynamical systems theory and stochastic differential equations, we explore how individuals perceive, interpret, and update their beliefs as they interact with AI chatbots that possess hidden traits of sycophancy. We treat the evolving conviction a

Read source article
arXiv cs.AIResearch

When Web Agents Finish but Still Fail: Reproducible Triggers and Trace Diagnostics for Parallel Web Exploration

arXiv:2606.20724v1 Announce Type: new Abstract: Long-horizon web agents often fail in ways hidden by final-answer evaluation: they may visit useful pages, produce a well-formed answer, and terminate confidently while still missing fields, over-including unsupported items, or relying on stale evidence. We study these failures with Parallel WebBench, a parallel web-exploration benchmark containing 1,679 verified records: 350 manually curated parallel tasks and 1,329 reconstructed records with veri

Read source article
arXiv cs.AIResearch

Fara-1.5: Scalable Learning Environments for Computer Use Agents

arXiv:2606.20785v1 Announce Type: new Abstract: Collecting computer use data from human demonstrations is expensive and slow, motivating the need for scalable generation strategies. This requires two key ingredients: environments in which agents can act and verifiers that can judge whether their demonstrations succeeded. We introduce FaraGen1.5, a scalable data pipeline for computer use agents composed of three modular components: environments, solvers, and verifiers. FaraGen1.5 uses both live w

Read source article
arXiv cs.AIResearch

What Shapes Emergent Misalignment? Insights from Training Dynamics, Model Priors, and Data

arXiv:2606.20814v1 Announce Type: new Abstract: Emergent misalignment (EM) is a phenomenon in which models generalize with narrow fine-tuning, leading to broad (yet uneven) misalignment across evaluation questions. We study EM and its variability directly through the components of fine-tuning: training dynamics, model priors, and data. (1) We first explored how in-domain training loss relates to out-of-domain alignment scores across datasets and model families. Then, we tried to induce potential

Read source article
arXiv cs.AIResearch

Study on Quantitative Dynamic Epistemic Logic for Belief Revision

arXiv:2606.20837v1 Announce Type: new Abstract: Belief revision is a process in which an agent begins to believe in something she previously did not. I begin the paper by presenting, based on (G\"ardenfors, 1998; Hansson, 1999), postulates for belief revision that constitute the basis of the AGM theory. I will then briefly show the semantics of a modal logic introduced in (van Ditmarsch, 2005), which I call `$P$'. This logic formalizes static epistemic states and has greater expressive power tha

Read source article
arXiv cs.AIResearch

Process-Reward Tactic Evolution for Long-Horizon Bioinformatics Workflows

arXiv:2606.20839v1 Announce Type: new Abstract: LLM agents can write code and call tools, but reliable bioinformatics work requires long-horizon interaction with workflow software, typed data objects, provenance, and biological checks. We study this setting through Galaxy workflow execution. The agent must explore task data, construct or adapt an executable workflow DAG, bind inputs and dataset collections, monitor execution, debug failures, and validate biological outputs. We propose Process-Re

Read source article
arXiv cs.AIResearch

SignVLA: Real-Time Sign Language-Guided Robotic Manipulation via Attention LSTM and Vision-Language-Action Models

arXiv:2606.20857v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models enable robots to execute manipulation tasks from natural-language instructions grounded in visual observations. However, existing VLA interfaces primarily rely on speech or text input, limiting accessibility for deaf, hard-of-hearing, and speech-impaired users. We present SignVLA, a real-time sign-language-guided VLA framework for accessible human-robot interaction. The system introduces a modular sign-to-text in

Read source article
arXiv cs.AIResearch

When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study

arXiv:2606.20881v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in large language model reasoning, but relies on ground-truth supervision that is costly or infeasible, especially in coding tasks. Recent work addresses this by deriving rewards from a model's own signals, such as majority voting or confidence-based scores, achieving notable success on mathematical reasoning benchmarks. However, code generation poses distinct cha

Read source article
arXiv cs.CVResearch

Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion

arXiv:2606.21135v1 Announce Type: new Abstract: Human motion generation has been widely studied across diverse input modalities, text, music, and video, and recent efforts have unified these into single multimodal frameworks. However, while morphological factors such as gender and body shape are known to produce distinct kinematic signatures, no existing unified framework incorporates this into generation, treating all subjects as morphologically equivalent. We present Odoriko, the first unified

Read source article
arXiv cs.CVResearch

SEED: Simple ViT and Evolving Harness for Explainable Text Forgery Detection

arXiv:2606.21138v1 Announce Type: new Abstract: AI-assisted image editing threatens trust in financial, legal, and identity records. The GenText-Forensics Challenge at ACM MM 2026 addresses this by requiring structured forensic reports, in which integrating detection, pixel-level localization, and natural language explanation for multilingual text-centric forgery images. We present SEED, a modular system with three components. First, a similarity-guided pipeline augments training with diverse sy

Read source article
arXiv cs.CVResearch

ChronoLock: Protecting Videos from Unauthorized Text-to-Video Personalization

arXiv:2606.21146v1 Announce Type: new Abstract: Text-to-video (T2V) diffusion models have made it increasingly easy to synthesize realistic and temporally coherent videos, while recent personalization techniques allow such models to imitate a specific subject, style, or motion pattern from only a few reference clips. This capability creates a new data-misuse risk: videos shared online can be collected and used for unauthorized T2V fine-tuning. Existing protective perturbations are mainly designe

Read source article
arXiv cs.CVResearch

BadDreamer: Transferable Backdoor Attacks against Video World Models for Autonomous Driving

arXiv:2606.21172v1 Announce Type: new Abstract: Video world models are increasingly used in autonomous driving to forecast future scene evolution and provide future-aware spatio-temporal representations for downstream action prediction. In perception-to-action pipelines, these representations can directly influence ego-vehicle waypoint planning, making the learned future dynamics a critical security-sensitive component. Despite their promise, the training-time security risks of autonomous-drivin

Read source article
arXiv cs.CVResearch

Context-Aware Autoregressive Diffusion for Gloss-Wise Sign Language Production

arXiv:2606.21234v1 Announce Type: new Abstract: To generate natural and accurate sentence-level sign language, synthesizing the "gloss", the fundamental semantic unit, is essential. However, most current sign-language production (SLP) methods generate entire sequences at once. While this end-to-end approach is often efficient, it is prone to temporal drift and hand motion blur as sentences get longer, and fails to accurately control individual glosses. In this paper, we propose the Context-aware

Read source article
arXiv cs.CVResearch

A Neurosymbolic Framework for Interpretable Skeleton-Based Seizure Detection via Concept-Driven Logical Reasoning

arXiv:2606.21252v1 Announce Type: new Abstract: Video-based seizure detection is essential for the management of epilepsy patients, offering a non-invasive complement to electroencephalography. While several deep learning approaches have been developed for video-based seizure detection, none are inherently interpretable, limiting their adoption and translation into clinical practice. We present, to our knowledge, the first exploration of a neurosymbolic framework for video-based seizure detectio

Read source article
arXiv cs.CVResearch

Lightweight 3D Feature Pretraining by Bayesian Inversion of 2D Foundation Models

arXiv:2606.21292v1 Announce Type: new Abstract: We present Casper3D, a lightweight probabilistic framework for converting noisy multi-view 2D foundation-model embeddings into a latent 3D semantic representation. We model view-level semantic features as noisy observations of an underlying 3D semantic state and infer this state with a set-based variational model that incorporates relative pose during multi-view reasoning. Casper3D is trained by predicting held-out semantic observations from novel

Read source article
arXiv cs.CVResearch

A Test-time Actor-Critic Approach to News Images Generation

arXiv:2606.21304v1 Announce Type: new Abstract: This paper introduces the CERTH-ITI solution for the MediaEval NewsImages 2026 challenge, which focuses on generating images related to news headlines. Inspired by the Actor-Critic paradigm in reinforcement learning, we present a test-time, model-agnostic Actor-Critic Image Generation approach (ACIG). ACIG generates prompts for image creation, produces the images, evaluates the generated results, and if needed refines the image generation prompts a

Read source article
arXiv cs.CVResearch

LEViL: Label-Efficient Video Learning via Zero-Shot Distillation over VLM-Generated Pseudo-Label Spaces

arXiv:2606.21358v1 Announce Type: new Abstract: Supervised video pretraining is a common transfer learning practice for improving downstream action recognition performance. However, it requires large-scale labeled source datasets, and the effectiveness of the learned initialization is influenced by the similarity between the source and target domains. Constructing such labeled pretraining datasets for different target domains is costly and difficult to scale. To address these limitations, this s

Read source article
arXiv cs.CVResearch

FLM-Occ: Feed-forward Likelihood Maximization for Efficient Indoor Occupancy Prediction

arXiv:2606.21373v1 Announce Type: new Abstract: Recent indoor occupancy prediction methods adopt Gaussian primitives as a sparse 3D representation for computational efficiency. However, their training relies on voxel classification, which imposes only local constraints and lacks global supervision on the distribution of the primitives. Therefore, they inevitably predict spurious primitives in empty regions, undermining both representational and computational efficiency. To address this, we propo

Read source article
arXiv cs.CVResearch

OSOG: A Differentiable, Physics-Informed Synthetic Data Engine for Micro-Optical Environments

arXiv:2606.21381v1 Announce Type: new Abstract: Deep learning in computational microscopy is severely constrained by the scarcity of densely annotated datasets. While synthetic data generation has bridged this gap in macroscopic computer vision, traditional graphics engines rely on geometric ray-tracing, failing to capture the micro-optical phenomena required for microscopy. Conversely, while wave-optics formulations exist, rendering them computationally tractable at the scale required for deep

Read source article
arXiv cs.CVResearch

Technical Report for ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge: Exploring Query-Based Segmentation and Increased Spatial Context for Outdoor Scene Understanding

arXiv:2606.21456v1 Announce Type: new Abstract: In this report, we present our submission to the GOOSE 2D Fine-Grained Semantic Segmentation Challenge, organized as part of the Workshop on Field Robotics at ICRA 2026. The challenge combines data from the GOOSE and GOOSE-Ex datasets, which comprise more than 13k images captured from 4 distinct camera setups, annotated using a hierarchical taxonomy of 56 fine-grained classes and 11 broader categories. Starting from SegFormer as a baseline, we prog

Read source article
arXiv cs.CVResearch

Semi-Supervised Vision-Language-Action Model

arXiv:2606.21493v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models enable robots to predict actions directly from visual observations and language instructions, but adapting them to new environments still depends on costly action-labeled demonstrations. To reduce this dependence, we study semi-supervised VLA adaptation under limited supervision signals, where only a small portion of trajectories contain robot actions and the remaining trajectories provide action-unlabeled vision

Read source article
arXiv cs.CVResearch

Compressing Observation History into Agent Memory: Distilling Transformers into Recurrent Transformers

arXiv:2606.21562v1 Announce Type: new Abstract: Transformers are AI's workhorse with strong performance in modeling sequential data, but their computational cost becomes prohibitive when processing long sequences. We target long-horizon streaming vision and robotics applications like map-free pose estimation, where it is particularly impractical to store and maintain a history of observations. Recurrent Transformers address this limitation by maintaining fixed-size memory but their performance l

Read source article
arXiv cs.CVResearch

Radial Basis Function Networks as Projection Heads in Self-Supervised Learning

arXiv:2606.21590v1 Announce Type: new Abstract: Self-supervised learning (SSL) typically relies on a backbone encoder followed by a small multilayer perceptron (MLP) projection head, which is conventionally discarded after training, while backbone quality is assessed via costly linear probing on labeled data. We argue that this approach including discarding the projector is rather computationally wasteful. Instead, we propose replacing the MLP head with a radial basis function network (RBFN), wh

Read source article
arXiv cs.CVResearch

$\mu$Match: Foundation Models for Semi-supervised Learning and Domain Adaptation in EM

arXiv:2606.21605v1 Announce Type: new Abstract: Vision foundation models have substantially advanced computer vision, enabling state-of-the-art performance in zero- and few-shot settings. They have been successfully applied to biomedical imaging tasks ranging from organ segmentation in computed tomography to cell segmentation in light microscopy. Electron microscopy (EM) is a central modality for analyzing cellular ultrastructure due to its nanometer-scale resolution. However, the application of

Read source article
arXiv cs.CVResearch

Chehre: An Emoji-Prompted Video Dataset for Perceptually Diverse Facial Expression Recognition

arXiv:2606.21657v1 Announce Type: new Abstract: Facial expressions are nonverbal social signals used in human interaction, but facial expression recognition datasets often focus on static images, basic emotion categories, or single deterministic annotations. We introduce Chehre, an emoji-prompted video dataset for analyzing dynamic facial expressions across a wide range of expressions for exploring inter-individual perceptual diversity. In Chehre, participants were prompted to express and record

Read source article
arXiv cs.CVResearch

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

arXiv:2606.21661v1 Announce Type: new Abstract: Generating a coherent multi-shot video requires structured cross-shot memory. Subject appearance, scene context, and speaker identity must persist across cuts. Existing approaches either train end-to-end over fixed-length sequences and cannot scale, generate shot-by-shot with memory banks that grow linearly, or orchestrate pretrained generators under an LLM planner without a multi-shot-aware backbone. We present UnityShots, a memory-driven multi-sh

Read source article
arXiv cs.CVResearch

Enlight: Fast Low-Light Image Enhancement via Multi-Objective Optimization and Shadow-Aware Refinement

arXiv:2606.21674v1 Announce Type: new Abstract: We present ENLIGHT, a fast and training free framework for low-light image enhancement based on direct optimization of a perceptual objective. Unlike deep learning approaches that require large scale training data and supervision, ENLIGHT operates in a zero-shot manner by optimizing image quality at inference time. The method employs a two stage global to local optimization strategy. In the first stage, ENLIGHT performs global illumination adjustme

Read source article
arXiv cs.CVResearch

VT-DUDA: Visual Token Conditioning for Diffusion-guided Unsupervised Domain Adaptation

arXiv:2606.21700v1 Announce Type: new Abstract: Unsupervised domain adaptation (UDA) aims to learn a target-domain classifier from labeled source data and unlabeled target data under distribution shift. Recent diffusion-based UDA methods approach this problem by synthesizing labeled target-style images and training on the resulting synthetic data. However, their performance depends heavily on the conditioning design: class prompts provide only coarse guidance, while domain adaptation modules mai

Read source article
arXiv cs.LGResearch

Towards CSI-Native Foundation Models: A Channel-Adaptive Roadmap for 6G

arXiv:2606.20670v1 Announce Type: new Abstract: Wireless foundation models offer a path toward reusable channel state information (CSI) intelligence for sixth-generation (6G) systems. However, existing generic-backbone adaptation and CSI pretraining methods often treat CSI as task tensors rather than propagation-conditioned channel responses, thereby failing to capture the intrinsic time-frequency-spatial geometry of wireless environments. This paper presents a channel-adaptive roadmap toward CS

Read source article
arXiv cs.LGResearch

NeuroShield: A Device-Agnostic Foundation Model for EEG Authentication

arXiv:2606.20673v1 Announce Type: new Abstract: A central challenge in EEG authentication is that models are typically tied to the acquisition settings in which they are trained. In particular, variations in headset hardware, channel layout, and signal duration create heterogeneous recordings that existing models are not designed to handle, causing each new headset or dataset to be treated as a separate model-development problem. This fragmentation limits multi-dataset learning, hinders knowledg

Read source article
arXiv cs.LGResearch

Massive Activations Are Architecturally Robust: A Controlled Scratch/Commitment Residual Stream Test

arXiv:2606.20743v1 Announce Type: new Abstract: Trained transformers reliably develop massive activations, a small number of hidden dimensions whose magnitude is far above the median and which concentrate on the sequence-start token. Whether these outliers are a removable artifact of the residual stream's overloaded read and write role, or instead a functional necessity, is actively debated. We test the artifact hypothesis directly, with an architectural intervention. Our architecture, Ledger Re

Read source article
arXiv cs.LGResearch

B[FM]$^2$: Brain Foundation Model via Flow Matching with SplitUNet

arXiv:2606.20812v1 Announce Type: new Abstract: EEG foundation models can learn generalizable representations from large-scale EEG corpora to enable single-backbone transfer across diverse clinical and brain-computer interface tasks. Existing models typically discretize the continuous multi-channel EEG waveform into patches or codebook tokens and train a transformer with masked self-supervision. Recognizing that this discretization fragments continuous brain rhythms and obscures fine-grained tem

Read source article
arXiv cs.LGResearch

Machine Learning Classification of Cryopathy Syndromes: A Comprehensive Comparative Study

arXiv:2606.20874v1 Announce Type: new Abstract: Cryopathy syndromes are difficult to classify because laboratory patterns often overlap across diagnostic categories, while some diagnoses are rare. This makes routine interpretation of cryoglobulin-related tests challenging and increases dependence on expert judgment. The aim of this study was to develop and compare machine learning approaches for automated classification of cryopathy syndromes from laboratory data and to identify a practical stra

Read source article
arXiv cs.LGResearch

MMGNN: Multi-level, multi-color graph neural networks for molecular property prediction

arXiv:2606.20906v1 Announce Type: new Abstract: Molecular message-passing neural networks commonly propagate chemically diverse interactions through a single graph, which may mix interaction-specific signals and require deep propagation to capture long-range effects. We introduce the Multi-level, Multi-color Graph Neural Network (MMGNN), a hierarchical framework that decomposes a molecular graph into overlapping atom-type-pair-specific subgraphs while preserving atom-level resolution. MMGNN-2D c

Read source article