AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

31239 stories from 30+ sources, refreshed continuously.

arXiv cs.AIResearch

LabAgent: Customize Any Research Hubs for Scientific Discoveries Using AI Agents

arXiv:2609.13437v1 Announce Type: new Abstract: Scientific research is a continuous process that emphasizes inheritance. Methods developed by predecessors are often expanded upon by new researchers to explore more novel and in-depth scientific questions. However, the change of lab staff, such as student graduation, leads to a lack of personnel capable of replicating methods. Methods that have been developed with significant effort and resources cannot be continued. To address these limitations,

Read source article
arXiv cs.AIResearch

TimeThink: Eliciting Compositional Reasoning in Timeseries Large Language Models

arXiv:2609.13457v1 Announce Type: new Abstract: Timeseries multimodal large language models (TS-MLLMs) have recently begun leveraging the reasoning capabilities of large language models (LLMs) for question-answering tasks. However, these models often fail to capture dynamic temporal patterns, providing only implicit reasoning that lacks the underlying explanations critical for high-stakes applications like healthcare. While reinforcement learning (RL)-based timeseries language models aim to addr

Read source article
arXiv cs.AIResearch

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

arXiv:2609.13463v1 Announce Type: new Abstract: The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data renders human review impractical, driving the need for automated root-cause attribution (RCA). However, automated RCA methods using LLMs suffer from low diagnostic accuracy, especially as execu

Read source article
arXiv cs.AIResearch

Governing at Machine Speed: An Adaptive Intelligence Architecture for Real-Time AI Policy Enforcement

arXiv:2609.13466v1 Announce Type: new Abstract: Enterprise AI adoption has reached 78% of organizations globally, yet the infrastructure to govern that adoption has not kept pace. This paper identifies and characterizes the attestation deficit, a structural condition in which organizations maintain governance policies but cannot produce auditable, tamper-evident evidence of enforcement within regulatory timelines. Drawing on empirical data from the Stanford 2026 AI Index Report (362 documented i

Read source article
arXiv cs.AIResearch

OrchSLM: Probing the Dynamics of Small Language Model Orchestration

arXiv:2609.13470v1 Announce Type: new Abstract: Although large language models (LLMs) have demonstrated remarkable capabilities, their reliance on cloud-scale infrastructure poses fundamental challenges for deployment in agentic pipelines, including latency, privacy, connectivity, and substantial computational cost. Small language models (SLMs) offer a compelling alternative: recent studies suggest that many repetitive and narrowly scoped subtasks in agentic workloads may be better served by spe

Read source article
arXiv cs.AIResearch

Grounded Adjudication of Variations across Extracted TimeLines (GAVEL): Comparing Clinical Timelines Against Their Case Reports

arXiv:2609.13475v1 Announce Type: new Abstract: Existing pipelines for clinical timeline extraction from case reports are evaluated using an expert reference and are limited by imperfect reference annotations and imprecise event alignment. We developed GAVEL, an LLM judge protocol that compares two timelines with the case report and returns a discrepancy type, verdict, and report passage for each difference. We evaluated the event matcher, reviewed 2,738 findings from GPT5.6sol and DeepSeek V3.2

Read source article
arXiv cs.AIResearch

Token Efficient Task Execution via Application Behavior Modeling for Web Agents

arXiv:2609.13491v1 Announce Type: new Abstract: The strong performance of AI Agents across an impressive variety of tasks is driving an unprecedented investment in agentic infrastructures, however the cost of processing tokens is fast increasing. Web agents automate the execution of web-application tasks described in natural language, by analyzing the web-application's user interface (UI) and interacting with it. This work introduces OdoBot, a novel web-agent architecture that completes tasks at

Read source article
arXiv cs.AIResearch

From Legal Text to AI-specific Risk Sources: A Systematic Analysis of the EU AI Act's High-Risk Requirements

arXiv:2609.13535v1 Announce Type: new Abstract: The EU AI Act introduces mandatory requirements for high-risk AI systems with the explicit goal of ensuring the development and operation of trustworthy AI. At the same time, AI risk management practices rely on structured risk taxonomies to systematically identify and treat AI-specific risk sources. As both the AI Act and established risk taxonomies aim to address AI-induced risks, a natural question is whether they align in the risk sources they

Read source article
arXiv cs.AIResearch

Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents

arXiv:2609.13543v1 Announce Type: new Abstract: LLM agents are predominantly benchmarked on short, single-task trajectories, yet real deployments run for hours under contention, surfacing a different class of failures. We use the Clinical Environment Simulator (CES), in which an agent manages an entire emergency-department shift under continuous time and resource pressure, as a testbed: long-horizon execution failures manifest measurably in a single rollout under structured, multi-dimensional gr

Read source article
arXiv cs.AIResearch

AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents

arXiv:2609.13548v1 Announce Type: new Abstract: Web agents can utilize reusable tools to reduce the cost and latency of low-level browser interaction, but automatically discovered tool collections can be large, redundant, and poorly aligned with user demand. We present AutoTailor, a meta-agentic framework for constructing and maintaining a compact set of trajectory-derived Model Context Protocol (MCP) APIs. Offline, AutoTailor converts web trajectories into parameterized browser-automation progr

Read source article
arXiv cs.AIResearch

Toward a Decision-Assurance Layer for AI-Assisted Flight Planning in Air Traffic Management

arXiv:2609.13552v1 Announce Type: new Abstract: Generative AI is increasingly being used informally in Air Traffic Management (ATM) for tasks such as flight plan generation, trajectory interpretation, and constraint checking. Although these tools can reduce workload and accelerate planning, their non-deterministic outputs create safety and operational risks in human-in-the-loop settings. This paper proposes the AI Trust and Assurance Layer (ATAL), a model-agnostic decision assurance architecture

Read source article
arXiv cs.AIResearch

Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems

arXiv:2609.13559v1 Announce Type: new Abstract: Large Language Models (LLMs) with function-calling capabilities are becoming critical for modern agentic AI systems. Nevertheless, current deployments typically route inferences to powerful cloud-based models, incurring significant energy use and carbon emissions. We address this sustainability challenge with a carbon-aware routing framework that distributes function-calling queries across a three-tier edge-cloud architecture, combining edge and cl

Read source article
arXiv cs.AIResearch

A Hybrid Agentic AI Framework for Intelligent Supply Chain Analytics

arXiv:2609.13561v1 Announce Type: new Abstract: Efficient utilization of supply chain analytics for decision making remains a significant challenge for planners, as critical tasks such as database querying, key performance indicator (KPI) analysis, demand forecasting, and performance diagnosis require heterogeneous expertise spanning data engineering, operations research, and domain knowledge. In this work, we propose an agentic system for supply chain analytics that bridges the gap between busi

Read source article
arXiv cs.AIResearch

Causal multi-modal AI for personalized chemosensitivity prediction

arXiv:2609.13567v1 Announce Type: new Abstract: Chemotherapy improves survival for some patients with breast cancer, but doctors cannot reliably predict who. Current guidelines rely on recurrence scores as a proxy for treatment benefit, which may contribute to the overprescription of chemotherapy. Here we present a causal multi-modal AI model that predicts personalized chemosensitivity using routinely collected pathology and clinical information. We developed our model on a multi-national datase

Read source article
arXiv cs.AIResearch

How User-AI Mistreatment Occurs and Matters in Conversational Systems?

arXiv:2609.13579v1 Announce Type: new Abstract: Safety research often focuses on model-generated harms, but users may also direct hostility, coercion, and adversarial pressure at models. Understanding how and when that occurs is essential for accurately interpreting model behaviour, alignment drift, and real-world deployment risks. In this paper, we audit 777K English LMSYS-Chat-1M conversations with two independent detectors: an eight-category lexicon for hostility directed at the model, and th

Read source article
arXiv cs.AIResearch

FLoKD: Adaptive Knowledge Distillation for Federated Low-Rank LLM over Wireless Networks

arXiv:2609.13580v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong capabilities across a wide range of natural language processing tasks. However, conventional fine-tuning typically relies on centralized data collection, bringing in privacy concerns. Federated learning (FL) enables collaborative LLM fine-tuning without sharing raw client data, but its deployment over bandwidth-constrained wireless networks is hindered by the communication overhead of model-para

Read source article
arXiv cs.CVResearch

ArtSociety: Multi-Agent Multimodal Collaboration for Art Emotion Understanding

arXiv:2609.13240v1 Announce Type: new Abstract: The AffectiveArt Multidimensional Art Emotion Understanding task asks to jointly predict an artwork's fine-grained emotion (12 classes, 1549:1 head-to-tail ratio), binary valence/arousal, and five attribute-grounded descriptions -- sub-tasks that exhibit strong empirical trade-offs, so the single-model solutions we tried do not jointly optimize all of them well. We present ArtSociety, a multi-agent framework that assembles heterogeneous multimodal

Read source article
arXiv cs.CVResearch

Personalized and Explainable Blood Pressure Estimation from PPG via Hybrid CNN--Morphological Features

arXiv:2609.13190v1 Announce Type: new Abstract: Continuous cuffless blood pressure (BP) monitoring using photoplethysmography (PPG) offers a promising solution for personalized healthcare. However, existing methods have two major limitations. Handcrafted feature-based approaches rely on precise fiducial point detection and are limited to short-term analysis, while deep learning models, despite their accuracy, often operate as black boxes with limited physiological interpretability. To address th

Read source article
arXiv cs.CVResearch

Part Grounding, Not Action Knowledge: Locating the Bottleneck in VLM Affordance Prediction

arXiv:2609.13225v1 Announce Type: new Abstract: Benchmarks agree that vision-language models reason poorly about low-level manipulation, but an aggregate accuracy score does not say which step fails. We separate two steps that affordance questions conflate: identifying which part of an object to act on, and knowing what action that part requires. Across 19 articulated objects we asked eight models, spanning three developers, what motion a robot should apply. Under an open prompt, push was produc

Read source article
arXiv cs.CVResearch

Synthetic Leprosy Image Generation Using Mask-Conditioned Latent Diffusion and Transfer Learning from Large Chronic Wound Datasets

arXiv:2609.13226v1 Announce Type: new Abstract: Machine learning for neglected tropical diseases is limited by data, not algorithms: public annotated image sets for leprosy (Hansen's disease) number in the hundreds, orders of magnitude below what generative models require. We ask whether a model trained on abundant chronic wound photography transfers to this low-data regime. We build a three-stage pipeline. First, a DeepLabV3-ResNet50 segmentation network (validation Dice 0.876, IoU 0.799) suppl

Read source article
arXiv cs.CVResearch

What Does the Encoder Actually Decide? A Controlled Comparison of Vision Backbones on Joint Tree Segmentation and Stereo Depth

arXiv:2609.13232v1 Announce Type: new Abstract: A robot pruning trees needs two facts per pixel: whether it belongs to a tree, and its distance. Both are usually obtained via task heads attached to a vision backbone chosen by reputation rather than measurement. Holding dataset, decoders, losses, schedule, and evaluation fixed, we ask: how much does the encoder choice change joint semantic segmentation and stereo depth on thin vegetation? We build a hard parameter-sharing network with one encoder

Read source article
arXiv cs.CVResearch

Abstract-LoRA: Unlocking Single-Image Style Transfer through Targeted U-Net Block Training

arXiv:2609.13239v1 Announce Type: new Abstract: Diffusion models represent one of the most advanced paradigms in generative modeling. Leveraging their development, a growing number of style transfer methods based on diffusion models have been proposed. However, among these methods, multi-image style transfer approaches that require at least five to ten style examples tend to achieve more satisfactory results. Single-image methods, by contrast, often struggle with either insufficient content pres

Read source article
arXiv cs.CVResearch

Preserving Subject-Clarity in Image Outpainting with Multiscale Wavelet Supervision

arXiv:2609.13251v1 Announce Type: new Abstract: Commercial and advertising images are frequently affected by poor framing, partially cropped subjects, truncated text or logos, and insufficient context, all of which can reduce subject clarity, i.e., the ability of an image to clearly communicate its primary subject. Image outpainting offers a scalable solution by extending image boundaries and recovering missing content and context. However, existing diffusion-based outpainting methods often prod

Read source article
arXiv cs.CVResearch

Filling the Unseen: Scene Extrapolation via 3D Gaussian Splatting

arXiv:2609.13262v1 Announce Type: new Abstract: 3D Gaussian Splatting achieves photorealistic reconstruction within training view distribution, yet it degrades on out-of-distribution novel views, exhibiting holes in unobserved regions and artifacts in observable areas. Recent works formulate this task as extrapolation and interpolation and try to address it with generative models, but remain limited in extrapolation scale and quality. They repeat a generate-reconstruct-shift cycle to progressive

Read source article
arXiv cs.CVResearch

Structure-Token Evidence-Anchored Reasoning for Scientific Chart Understanding

arXiv:2609.13267v1 Announce Type: new Abstract: Scientific charts encode quantities in axes, legends, and geometric marks, yet large vision-language models still treat them as natural photographs. Visual in-context examples do not expose the coordinate frame; unconstrained chain-of-thought can name a plausible number that was never read from a bar. We present STEER (Structure-Token Evidence-anchored Reasoning), which freezes a Llama-3.2-Vision encoder and inserts three modules: a chart structure

Read source article
arXiv cs.CVResearch

CANAL: Channel-Aware Noise Allocation for Differentially Private Feature Distillation in Medical Image Segmentation

arXiv:2609.13271v1 Announce Type: new Abstract: Medical image segmentation needs diverse training data, but hospitals hold complementary scans they cannot share for privacy and regulatory reasons. Knowledge distillation can bridge this gap by exporting learned feature representations instead of images, but those representations still encode patient-specific anatomy and remain vulnerable to membership-inference and feature-inversion attacks. Adding calibrated Gaussian noise restores a differentia

Read source article
arXiv cs.CVResearch

SomBench: Benchmark Dataset for Advancing Machine Learning in Lunar Science

arXiv:2609.13277v1 Announce Type: new Abstract: Lunar orbital missions, such as Lunar Reconnaissance Orbiter, Kaguya/SELENE, Gravity Recovery and Interior Laboratory, and Lunar Prospector, among others, provide rich multi-instrument observations, but their heterogeneity in sampling, projection, and conventions limits reproducible machine learning (ML). We introduce SomBench, a unified, spatially-aligned, ML-ready lunar dataset aggregating 30+ co-registered layers from ten instruments across four

Read source article
arXiv cs.CVResearch

GaugeDefect: Detecting Surface Anomalies by Curvature of Feature Transport

arXiv:2609.13282v1 Announce Type: new Abstract: Industrial anomaly localization has advanced rapidly with feature-based, reconstruction-based, and distillation-based methods. Most of these methods score a region by asking how unusual its local appearance or feature representation is with respect to normal training images. This is a strong and practical formulation. In this work, we study a complementary geometric cue for cases where an abnormal region may still contain locally plausible visual f

Read source article
arXiv cs.CVResearch

Target-Checked Reliability Score Refinement for Video Question Answering

arXiv:2609.13288v1 Announce Type: new Abstract: Video-language models can answer multiple-choice questions with high confidence yet be wrong. We study whether answer-level reliability scores can be improved under target shift without retraining the models or changing their answers. We collect option-probability lists from three fixed video-language models under four deterministic video samplings and represent cross-view changes and cross-model agreement as a response graph. Using a labeled targe

Read source article
arXiv cs.CVResearch

VectorHarness: Recovering Editable, Relation-Preserving Structure from Scientific Graphics

arXiv:2609.13294v1 Announce Type: new Abstract: Converting scientific graphics into editable representations remains a challenging problem for image-to-code generation because of their heterogeneous elements and complex layouts. Recent multi-agent reconstruction systems have advanced this line of work, but often follow a copy-paste paradigm: the reconstructed image closely resembles the original, while complex regions remain effectively uneditable. We instead formulate a different objective, ras

Read source article
arXiv cs.CVResearch

Harnessing Image Question Dependence for Better VLM Test-time Reinforcement Learning

arXiv:2609.13296v1 Announce Type: new Abstract: Test-time reinforcement learning can adapt vision-language models (VLMs) to unlabeled target data, but its effectiveness is fundamentally limited by the reliability of self-generated learning signals. To assess the reliability of consensus-based learning signals, we analyze VLM test-time reinforcement learning across diverse VQA datasets and model sizes, revealing two limitations. First, gains from consensus-based test-time training largely come fr

Read source article
arXiv cs.CVResearch

HGSQ: Heatmap-Guided Sparse Query Detector for Real-Time Aerial Small Object Detection

arXiv:2609.13306v1 Announce Type: new Abstract: Real-time aerial small object detection is an important visual signal and image processing problem, requiring a detector to preserve fine-grained localization while avoiding redundant computation on large background regions. This paper focuses on this deployment-oriented aerial/UAV setting rather than claiming a universal detector for all object detection scenarios. Existing Transformer-based detectors provide strong global modeling, but their dens

Read source article
arXiv cs.CVResearch

GroundBench: A Factorized, Counterfactual Benchmark for Locating VLM Affordance Failures

arXiv:2609.13308v1 Announce Type: new Abstract: A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language models, with no model outperforming a constant baseline until the part was named. However, naming the part supplies information that a real system must infer, confounding visual grounding, mechanical reasoning, and category-to-action association. We introduce GroundBench, a diagnostic benchmark that s

Read source article
arXiv cs.CVResearch

Towards Practical Precision Agriculture: Real-Time Fruit Detection and Video Analytics on Embedded Edge Hardware

arXiv:2609.13551v1 Announce Type: new Abstract: Static-image benchmarks do not capture the computational and temporal requirements of practical orchard video analytics. This study presents an end-to-end framework for real-time fruit detection, tracking, and counting on the NVIDIA Jetson Orin Nano Super. A lightweight YOLO26s detector is trained independently on four public datasets representing apples, mangoes, blueberries, and strawberries under a common protocol. The models are deployed on emb

Read source article
arXiv cs.CVResearch

From Advertised Improvements to Measured Capabilities: Evaluating ChatGPT Images 2.5 on Forgery Tasks

arXiv:2609.13617v1 Announce Type: new Abstract: We evaluate whether the improvements advertised for ChatGPT Images 2.5 translate into better performance on forgery tasks with predetermined answers. We compare its Flare and Sunburst API models with GPT-Image-2 re-run in the same week, using receipt-field edits, repeated editing, product placement and fine-print rendering. After image registration, Flare and Sunburst show fewer OCR-detected changes to surrounding receipt text (31.7% and 31.2% vers

Read source article
arXiv cs.CVResearch

Multimodal Foundation Models Adaptation based on Domain-Aware Relaxed Orthogonal Subspace for Remote Sensing

arXiv:2609.13654v1 Announce Type: new Abstract: Pretrained foundation models (FMs) have achieved remarkable success in computer vision, yet their high fine-tuning cost limits practical deployment. Parameter-efficient fine-tuning (PEFT) methods such as Low-Rank Adaptation (LoRA) improve efficiency by constraining updates to a predefined low-rank subspace. However, when applied to remote sensing tasks with substantial domain shifts, the fixed subspace is constructed without observing the downstrea

Read source article
arXiv cs.LGResearch

BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents

arXiv:2609.13149v1 Announce Type: new Abstract: For local large language model agents, active context is a scarce resource: memory capacity, prefill latency, cache growth, and service objectives all constrain how many input tokens each call can afford. We present BudgetBench, an active-budget protocol and reference harness that treats the per-call input-token budget as the independent variable when comparing memory strategies. Holding the model, task, sampler, and decoding fixed, it sweeps budge

Read source article
arXiv cs.LGResearch

A derivative-fidelity failure mode in physics-informed neural networks: strengthened benchmark evidence from function-value training

arXiv:2609.13171v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) use automatic differentiation to impose differential-equation residuals, but good agreement in function values does not necessarily imply accurate derivatives. This paper formulates derivative fidelity as a failure mode of PINNs and tests it with one-dimensional benchmarks. Multilayer perceptrons are trained only on function values for sin(x) and exp(x), while second derivatives obtained by automatic differe

Read source article
arXiv cs.LGResearch

Early Prediction of Satellite Collision Probability Using a Hybrid TCN-Transformer Model for a CDM-Based Conjunction Analysis Framework

arXiv:2609.13191v1 Announce Type: new Abstract: The rapid expansion of operational satellites and orbital debris has increased the frequency of close approach events in low Earth orbit (LEO), creating a higher operational burden for satellite operators. This problem is especially critical for satellites using electric propulsion, where low-thrust maneuver capability imposes additional time constraints on collision avoidance planning. In current practice, Conjunction Data Messages (CDMs) provide

Read source article
arXiv cs.LGResearch

Evaluating LLM-Generated Rules for Heart Disease Prediction

arXiv:2609.13192v1 Announce Type: new Abstract: This study compares traditional machine learning models and Large Language Model (LLM)-generated rule-based systems for heart disease prediction using the UCI Heart Disease dataset. Several classifiers, including Logistic Regression, K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Naive Bayes, Decision Tree, and Random Forest, were evaluated alongside rule-based systems generated using GPT-4o and Claude Sonnet 4.6. Model performance was as

Read source article