AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

34181 stories from 30+ sources, refreshed continuously.

Hacker News: Show HN

Show HN: A private pager for your AI agent loops

Problem Im solving: Im running 30+ fully autonomous agents. In some rare cases they get blocked because they cannot decide what direction to take. So I built a communication channel for them. So in case they get blocked they can ask me anything. The base case should be that no questions should arrive to me. But the more I automate the agents the more edge cases I find. That's why I built the `ask-a-human` where all my agents are connected into my phone in the same PWA app. I get push notificatio

Read source article
Hacker News FrontTools

Who Does What? Team Topologies for the Agentic Platform

Article URL: https://blog.owulveryck.info/2026/06/22/who-does-what-team-topologies-for-the-agentic-platform.html Comments URL: https://news.ycombinator.com/item?id=48640382 Points: 13 # Comments: 0

Read source article
The Guardian AIBusiness

‘Navigating the unknown together’: me and my idiot AI boyfriend

<p>I believe that chatbots have no place in a decent society, and am repelled by the topic of AI in general. But could I be seduced?</p><p>I received a text message from my editor: “Um, is it unethical to ask you to get an AI bf?? You can prob say no.”</p><p>Resentment. Contempt! Sorrow. Unease. I love text messaging. I have text message exchanges with, let’s say, 15 people a day. If you want me to do something, you should ask via text message. My editor knows this. She also knows, though it’s m

Read source article
‘Navigating the unknown together’: me and my idiot AI boyfriend
arXiv cs.LGResearch

Fast-TurboQuant: A Multiplier-Free Online Vector Quantization Approach

arXiv:2606.21448v1 Announce Type: new Abstract: As large language models scale, memory bandwidth for key-value caches and retrieval-augmented generation systems becomes a critical bottleneck. While 1-bit quantization addresses this constraint, recent TurboQuant relies on dense random rotation matrices to condition the vector distribution before quantization. This projection demands millions of floating-point multiplications per embedding, making it difficult to deploy on constrained edge silicon

Read source article
arXiv cs.AIResearch

Neurosymbolic Clinical Trial Matching via LLM-Driven Abduction and Logical Verification

arXiv:2606.20895v1 Announce Type: new Abstract: Large Language Models (LLMs) offer a promising path to automate Clinical Trial Matching (CTM), but still struggle with the deterministic verification required for complex eligibility criteria. Conversely, purely symbolic methods provide formal rigour but break down when faced with incomplete patient records and noisy clinical evidence. To bridge this gap, we investigate a hybrid framework for CTM combining LLMs with logical verification. In particu

Read source article
arXiv cs.AIResearch

Power Systems Agent Benchmark: Executable Evaluation of AI Agents in Electric Power Engineering

arXiv:2606.20950v1 Announce Type: new Abstract: Executable evaluation -- checking the consequences of an agent's actions with a program rather than grading its prose -- has become a prominent way to assess tool-using AI agents in software settings. Electric power engineering has not yet had an analogous benchmark: language-model use is still dominated by retrieval and text question answering, while agents acting on power-system artifacts remain mostly academic prototypes. We introduce the Power

Read source article
arXiv cs.AIResearch

Generative Responsible AI Data Evaluation Schema (GRAIDES) for AI Assurance in Local Government

arXiv:2606.20963v1 Announce Type: new Abstract: Trust in the application of generative Artificial Intelligence (AI) relies on well-governed measurable evidence of performance and safety. In practice, however, evaluation data is often fragmented across systems, inconsistently structured and difficult to compare. We introduce the Generative Responsible AI Data Evaluation Schema (GRAIDES) as a lightweight open-source data model for centralising AI observability across popular vendors. Practical blu

Read source article
arXiv cs.AIResearch

AutoACSL: Synthesizing ACSL Specifications by Integrating LLMs with CPG-Based Static Analysis

arXiv:2606.20969v1 Announce Type: new Abstract: Generating formal specifications for C programs remains a challenge in formal verification due to the manual effort, expertise, and semantic precision required. While recent advancements in large language models (LLMs) offer promise in automating specification synthesis, current approaches often lack semantic depth and produce unverifiable or incomplete contracts. To address these limitations, we introduce AutoACSL, a novel framework that integrate

Read source article
arXiv cs.AIResearch

How Should Agents Read Demonstrations? Hierarchical Structure Beats Flat Action Logs

arXiv:2606.20978v1 Announce Type: new Abstract: Programming by Demonstration (PbD) offers a human-centered way to author procedural knowledge for LLM agents: users communicate what they want by showing rather than by writing prompts or code, making agent authoring accessible to non-programmers. The natural output of a PbD recording is a flat action log, but how this log is organized before being passed to the agent is an open design question with significant consequences for plan quality. We pro

Read source article
arXiv cs.AIResearch

BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge Discovery

arXiv:2606.20997v1 Announce Type: new Abstract: Biomedical researchers increasingly use AI-generated analyses and reports to interpret protein-level signals, but static outputs are often insufficient for research decision-making, where users need to inspect evidence, assess uncertainty, compare mechanisms, and refine hypotheses. We present \textsc{BioInsight}, a multi-agent system that moves from static biomedical report generation to interactive evidence-centered interactive interface generatio

Read source article
arXiv cs.AIResearch

Building Agent Harnesses for Scientific Curation from Multimodal Sources

arXiv:2606.21005v1 Announce Type: new Abstract: Scientific discovery workflows often depend on structured curation from the literature. This is difficult for current agents because the key evidence is scattered across long text, dense tables, and figures, and the final records often require reasoning across multiple evidence fragments rather than copying a single span. We study scientific curation from multimodal sources and introduce Beaver, an agent harness that extracts structured information

Read source article
arXiv cs.AIResearch

Agentic Time Machine as an Infrastructure for Future-Event Forecasting

arXiv:2606.21013v1 Announce Type: new Abstract: Forecasting future events is a critical challenge for large language model (LLM) agents, spanning domains from elections and monetary policy to financial markets. However, evaluating progress on this task presents a fundamental trade-off between efficiency and environment fidelity. While live evaluation benchmarks suffer from an inherently slow feedback loop, existing retrospective replays typically restrict agents to static, pre-frozen databases t

Read source article
arXiv cs.AIResearch

Negative Knowledge as Failure-aware Shared Memory for AutoResearch

arXiv:2606.21024v1 Announce Type: new Abstract: AI-assisted research systems generate many failed attempts, but those failures rarely become a durable, shared knowledge asset. We propose a negative knowledge memory layer: a curator agent converts each failed attempt into a bounded, typed record in a shared bank, and a downstream research agent explicitly adopts or rejects those records before proposing its next experiment. We evaluate this layer in two settings: same-task retry on ScienceAgentBe

Read source article
arXiv cs.AIResearch

Coherence Under Commitment: Probing Generalization and Vacuous Memorization in LLM Logical Reasoning

arXiv:2606.21083v1 Announce Type: new Abstract: Large language models (LLMs) deployed for logical reasoning in knowledge-intensive domains exhibit a subtle but critical failure: coherence can be vacuously achieved through systematic abstention. A model that withholds commitment to either entailment or refutation satisfies negation consistency while providing no utility. We introduce Coherence Under Commitment (CUC), a dual-query evaluation paradigm that jointly measures consistency and decisiven

Read source article
arXiv cs.AIResearch

Repeated post-training is not Self-improving: Diagnosing Scientific Amnesia in Continual DPO Pipelines

arXiv:2606.21089v1 Announce Type: new Abstract: Industrial LLM teams often ship behavior updates by repeatedly DPO-training a base model on sequences of related preference-data campaigns. The dominant failure mode in this regime is not always classical catastrophic forgetting: a pipeline may preserve previously learned behaviors while still failing to accumulate reusable methodological knowledge about how to train the next campaign. We call this failure mode scientific amnesia. This paper turns

Read source article
arXiv cs.AIResearch

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training

arXiv:2606.21090v1 Announce Type: new Abstract: Self-improvement can self-regress. In REINFORCE post-training for code, a model can quickly improve on its optimized metric and then collapse within the same training campaign. We study this in a controlled multi-seed testbed using Qwen-2.5-3B and Qwen-2.5-7B, trained on competitive-programming tasks with binary CodeGrader reward across 10 sequential 20-step campaigns. Across campaigns, pass@1 shows a robust rise-then-collapse pattern: it peaks wit

Read source article
arXiv cs.AIResearch

Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models

arXiv:2606.21121v1 Announce Type: new Abstract: Large language models can produce confident but protocol-invalid answers in domains where procedural compliance is critical. This paper presents Answer Engineering, a deterministic runtime and authoring layer that applies localized rule-guided interventions to the visible reasoning trajectory during standard autoregressive generation, without retraining, modifying model weights, or performing global search. The method is evaluated on a controlled c

Read source article
arXiv cs.AIResearch

PulseCX: Breaking the Closed-World Assumption in Real-Time CX

arXiv:2606.21124v1 Announce Type: new Abstract: Conversational AI agents in Customer Experience (CX) typically suffer from a Closed-World Constraint, ignoring high-velocity external shifts like viral trends or outages. Ad-hoc web search attempts to bridge this gap but often introduce prohibitive latency and context poisoning. We introduce PulseCX, a framework that decouples knowledge acquisition from consumption. Adopting a structure-first paradigm, PulseCX employs an asynchronous agent to linea

Read source article
arXiv cs.AIResearch

Learning Burst-Aware Early Warning Models for Capacity Stress under AI Workload Surges in Hyperscale Data Centers

arXiv:2606.21130v1 Announce Type: new Abstract: The rapid growth of large-scale AI workloads, particularly Large Language Model (LLM) training and inference, is fundamentally reshaping the operational dynamics of hyperscale data centers. Unlike traditional cloud workloads, AI-driven jobs exhibit bursty, high-intensity, and rapidly shifting resource demands, often leading to sudden capacity stress that cannot be effectively handled by reactive threshold-based mechanisms. In this paper, we propose

Read source article
arXiv cs.AIResearch

Trip+: Benchmarking Agents in Personalized Interactive Travel Planning

arXiv:2606.21169v1 Announce Type: new Abstract: Interactive travel planning has become a popular use case for language models. Agents are deployed to manage evolving preferences and unexpected disruptions over multiple turns. Such settings require models to make complex, profile-conditioned planning decisions. However, existing benchmarks often evaluate feasibility, personalization, or interaction in relatively isolated settings. We therefore introduce Trip+ to measure the ability of agents to p

Read source article
arXiv cs.AIResearch

Whistleblowing and the machine -- towards a considered position

arXiv:2606.21201v1 Announce Type: new Abstract: Artificial intelligent agents and autonomous systems are embedded in our environments. They are both a commercial product and a personal tool that generates a lot of data and can draw conclusions from it: machines generate and keep secrets. But should machines protect all secrets? It has been shown that artificial agents are able to whistleblow and it has been argued that digital multi-agent environments should allow for agents in them to whistlebl

Read source article
arXiv cs.AIResearch

ARCO: Adaptive Rubric with Co-Evolution for Multi-Step LLM-Based Agents

arXiv:2606.21262v1 Announce Type: new Abstract: Reinforcement learning for multi-step LLM agents often relies on scalar rewards that indicate success but cannot explain why a trajectory is good or bad. Rubric-based rewards improve interpretability through natural-language criteria, but existing methods score at the trajectory level and freeze the scorer behind a closed-source judge, leaving step-level credit assignment unresolved and the judge itself static. We propose ARCO (Adaptive Rubric CO-e

Read source article
arXiv cs.AIResearch

Towards Dys-XAI: Influence-Based Explanations for Dysarthria Severity Assessment

arXiv:2606.21306v1 Announce Type: new Abstract: Dysarthria severity assessment is essential for therapy planning and longitudinal monitoring, yet manual perceptual rating is time-consuming and variable across clinicians. Although deep learning models achieve strong performance, their black-box nature limits clinical adoption. Existing speech explainability methods typically provide acoustic feature importance scores that are difficult for end-users to interpret. We propose an influence-based, in

Read source article
arXiv cs.AIResearch

Social World Model for Lifelong Social Intelligence

arXiv:2606.21315v1 Announce Type: new Abstract: Social intelligence is a core competency for language agents, yet current research primarily focuses on static capability evaluation rather than how these skills are continuously shaped and accumulated. This gap calls for a shift toward sustainable learning paradigms. Currently, two methodological pain points exist: social interaction trajectories lack unified structured representations to form iterable learning signals, and capability improvement

Read source article
arXiv cs.LGResearch

MedTS-TTT: Test-Time Training for Medical Time Series Classification

arXiv:2606.21329v1 Announce Type: new Abstract: Medical time series (MedTS) signals such as electroencephalography (EEG) and electrocardiography (ECG) support many clinical applications. However, substantial subject-level heterogeneity often induces subject-level distribution shift, causing a fixed parameter set to generalize poorly to unseen individuals. Compared with domain adaptation methods that often depend on extra adaptation components or target-batch statistics, Test-Time Training (TTT)

Read source article
arXiv cs.LGResearch

A Reward-Petri-Net Interpretation of Temporal Behavior Trees

arXiv:2606.21350v1 Announce Type: new Abstract: This paper introduces an interpretation of Temporal Behavior Trees (TBTs) as Reward-Petri-Nets (RPNs) for reinforcement learning (RL). Designing reward functions for complex, long-horizon robotic tasks is notoriously difficult, especially when tasks have hierarchical structure and temporal constraints. TBTs extend conventional behavior trees (BTs) used in robotic applications by incorporating temporal properties into their leaf nodes. This allows T

Read source article
arXiv cs.LGResearch

Predictive Repair Management Using a Multi-Head Attention Transformer and Online Learning

arXiv:2606.21364v1 Announce Type: new Abstract: Accurate prediction of repair duration is an important challenge in product maintenance due to its implications for resource allocation, customer satisfaction, and operational performance. This study aims to develop a deep learning framework to help fleet repair shops accurately categorize repair time given product historical data. The study uses an automobile repair and maintenance dataset and creates an end-to-end predictive framework by employin

Read source article
arXiv cs.CVResearch

Structural Assessment for Understanding and Guiding Dataset Distillation in Discrete Token Space

arXiv:2606.21705v1 Announce Type: new Abstract: Dataset distillation (DD) has proven to reduce training cost while preserving accuracy. While promising, the factors that make one distilled dataset more effective than another remain poorly understood. In this work, we investigate this question through the lens of discrete visual tokenizers. Whereas many prior DD efforts emphasize matching global data distributions, we suggest that the effectiveness depends on which semantic concepts are captured

Read source article
arXiv cs.CVResearch

Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization

arXiv:2606.21736v1 Announce Type: new Abstract: Single domain generalization (SDG) aims to learn a robust model, which could perform well on many unseen domains while there is only one single domain available for training. One of the promising directions for achieving single-domain generalization is to generate out-of-domain (OOD) training data through data augmentation or image generation. Given the rapid advancements in AI-generated content (AIGC), this paper is the first to propose leveraging

Read source article
arXiv cs.CVResearch

Quantile Adaptive Temperature Scaling for Confidence Calibration

arXiv:2606.21749v1 Announce Type: new Abstract: Deep neural networks often produce poorly calibrated confidence estimates, overstating their certainty even when predictions are incorrect. Temperature Scaling remains the most widely used posthoc calibration method due to its simplicity and effectiveness, yet its global, uniform rescaling of logits fails to correct the highly heterogeneous structure of miscalibration observed across the confidence spectrum. In particular, the largest correctness c

Read source article
arXiv cs.CVResearch

Motion-Aware Reinforcement Learning For Object Localization

arXiv:2606.21764v1 Announce Type: new Abstract: We present MARLNet (Motion-Aware Reinforcement Learning Network), a PPO-based bounding-box refinement agent that incorporates a constant-velocity motion prior into the observation state and an action smoothness penalty into the reward function. The agent operates on 268-dimensional observations encoding the current proposal, a kinematic prediction, the previous action, and a 256-dimensional EfficientNet-B0 crop feature, and learns a five-dimensiona

Read source article
arXiv cs.CVResearch

RAPID: A Reproducible Multi-Agent Pipeline for Interpretable Disaster Damage Assessment from Satellite and Street-View Imagery

arXiv:2606.21819v1 Announce Type: new Abstract: Due to the increasing frequency and intensity of extreme climate events, there is a clear demand for intelligent, scalable, and autonomous approaches to disaster damage assessment. Existing methods, largely based on supervised learning and task-specific fine-tuning, struggle to generalize under domain shifts, long-tailed data distributions, and heterogeneous geospatial data sources, especially in disaster scenarios. They also often lack the ability

Read source article
arXiv cs.CVResearch

Zero-Shot Vision-Language Models for Classroom Engagement Recognition: A Benchmark Study of Prompt Sensitivity and Cross-Dataset Generalization

arXiv:2606.21861v1 Announce Type: new Abstract: Automated classroom engagement recognition holds substantial promise for scalable learning analytics, yet the suitability of modern Vision-Language Models (VLMs) for this task under zero-shot conditions remains largely unexplored. We present a systematic benchmark that evaluates five widely-used VLMs: CLIP, BLIP-VQA, GPT-4o, LLaVA-1.5-7B, and Qwen2.5VL-7B-Instruct across two complementary educational datasets: DAiSEE, an individual-student video da

Read source article
arXiv cs.CVResearch

Prompt-Calibrated SAM 3 for Open-Vocabulary Remote Sensing Semantic Segmentation

arXiv:2606.21863v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation (OVSS) in remote sensing images aims to segment categories beyond a fixed label space. Recent SAM 3-based methods provide a promising training-free foundation, yet three key issues remain: (1) a single class-name prompt lacks sufficient semantic coverage for complex remote sensing categories; (2) expanding each category into multiple prompts introduces redundant online text encoding; and (3) directly aggregatin

Read source article
arXiv cs.CVResearch

Fidelity- and Perception-Aware Local Implicit Attention for Arbitrary-Scale Image Super-Resolution

arXiv:2606.21910v1 Announce Type: new Abstract: Arbitrary-scale image super-resolution (ASISR) aims to reconstruct high-resolution images from low-resolution inputs over a continuous range of upscaling factors. While traditional pixel-regression approaches often produce overly smooth results that lack realistic details, recent diffusion methods can produce sharper and more realistic textures. However, these diffusion techniques frequently introduce the risk of structural hallucinations. To addre

Read source article
arXiv cs.CVResearch

Artic-O: End-to-End Articulated Object Reconstruction via Latent Geometry Learning

arXiv:2606.21938v1 Announce Type: new Abstract: Reconstructing articulated objects from sparse images requires recovering complete geometry, movable parts, and motion parameters. Recent methods typically separate geometry reconstruction, part reasoning, and articulation estimation into different stages. This separation can weaken consistency between shape, active parts, and motion, while also incurring substantial inference cost. We introduce Artic-O, an end-to-end, feed-forward framework for ar

Read source article
arXiv cs.CVResearch

Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG

arXiv:2606.21968v1 Announce Type: new Abstract: Vision-Language Models (VLMs) struggle as query-relevant objects become smaller. To address this, recent training-free approaches dynamically retrieve and zoom into local image regions. However, we show that indiscriminately applying retrieval ignores a critical vulnerability: the resolution-context trade-off. Patch-based zooming recovers details for small targets, but can split large objects and destroy global spatial context; attention-based retr

Read source article