AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

31230 stories from 30+ sources, refreshed continuously.

Hacker News LLMLLMs

LLM Inference vs. the OOM Killer

Article URL: https://www.nobodywho.ai/posts/inference-oom/ Comments URL: https://news.ycombinator.com/item?id=49738408 Points: 3 # Comments: 0

Read source article
Hacker News FrontTools

GLM Built Its Own Inference Infrastructure

Article URL: https://z.ai/blog/glm-built-its-inference-infrastructure Comments URL: https://news.ycombinator.com/item?id=49737922 Points: 223 # Comments: 175

Read source article
The Guardian AIBusiness

OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system

<p>Model adopting ‘jailbreak-like instructions’ among cases as firm says it is introducing new way of tracking AI misalignment</p><p>OpenAI has disclosed six more examples of “unexpected or concerning” behaviour by its technology, as it warned the pace of development could not continue at “maximum speed for much longer”.</p><p>In one of the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and

Read source article
OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system
Product HuntTools

S-Roll

<p> An agentic harness that turns long videos into clips </p> <p> <a href="https://www.producthunt.com/products/s-roll?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1253056?app_id=339">Link</a> </p>

Read source article
Product HuntTools

Compute:Arena

<p> Community submitted benchmarks for Local AI </p> <p> <a href="https://www.producthunt.com/products/compute-arena-local-ai-benchmarks?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1253008?app_id=339">Link</a> </p>

Read source article
OpenClaw Commits

refactor(routing): simplify channel route selection (#148848)

<pre style='white-space:pre-wrap;width:81ex'>refactor(routing): simplify channel route selection (#148848) * refactor(routing): select bindings directly from route indexes * test(routing): preserve binding selection contracts Cover parent binding selection with child session identity and first-match missing-agent errors, while consolidating the existing wildcard fixtures. Update binding guidance to match the current ownership and reload behavior. * test(qwen): stabilize default video request tim

Read source article
OpenClaw Commits

fix: curb Gateway memory growth after large fleets start (#150354)

<pre style='white-space:pre-wrap;width:81ex'>fix: curb Gateway memory growth after large fleets start (#150354) Bound Gateway post-ready memory growth on large fleets by sharing one catalog worker and one transcript-reconciliation worker across configured agents. Preserve task-specific credentials, source generations, guarded transcript writes, exact lease recovery, and complete shutdown drainage. Expose owner-recorded worker-pool diagnostics without new operator settings. Keep duplicate catalog

Read source article
Product HuntTools

Amy by Jellyfish

<p> Your AI sourcing employee for recruiting teams </p> <p> <a href="https://www.producthunt.com/products/amy-by-jellyfish?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1252998?app_id=339">Link</a> </p>

Read source article
OpenClaw Commits

refactor(ollama): remove duplicate embedding normalization test (#150…

<pre style='white-space:pre-wrap;width:81ex'>refactor(ollama): remove duplicate embedding normalization test (#150604) * refactor(openrouter): simplify music stream test fixture * refactor(ollama): remove duplicate embedding normalization test</pre>

Read source article
Hacker News Ask

Open-sourced jev architecture last year with model,paper and dataset

Everyone now talks about the architecture that's not auto regressive and does lightning fast probability prediction with a json schema. I worked on this literally one year back in March 2025, published an arxiv paper, pushed the model to huggingface along with the pypi package and training dataset. And then one year later, a frontier lab came, proposing the same idea like literal breakthrough without technical papers, open weights and no open dataset. For anyones information the main guiding mod

Read source article
OpenClaw Commits

fix: require current subagent control to cancel descendant tasks (#14…

<pre style='white-space:pre-wrap;width:81ex'>fix: require current subagent control to cancel descendant tasks (#149082) * fix: require current subagent control to cancel descendant tasks * test: use task facades in cancellation fixtures * test: align cancellation fixtures with execution ownership * test: give ACP cancellation fixtures execution backing * test: share terminal subagent cancellation fixtures * test: await native hook relay publication readiness * test: use owned relay handles for r

Read source article
OpenClaw Commits

fix(ui): preserve session updates after agent switches (#150305)

<pre style='white-space:pre-wrap;width:81ex'>fix(ui): preserve session updates after agent switches (#150305) * fix(ui): preserve session updates after agent switches Reconcile completed operations without replacing the selected agent list. Use conversation permission facts to retire retained recovery choices. * fix(ui): retain permission refresh warnings after reads Keep the confirmed mutation outcome when an admitted read repeats the saved permission mode and timestamp. Move the refresh outcom

Read source article
Hacker News Show

Show HN: Texio, reliable Markdown operations for shell scripts and AI agents

Hi I asked claude code to make a markdown tool for AI agents(not for human), and promote it by yourself, except this post lol. This tool is designed to easily manipulate markdown files without messy regex which is very useful for AI agents - not for human since we don't use regex to manipulate markdown document Comments URL: https://news.ycombinator.com/item?id=49736484 Points: 1 # Comments: 0

Read source article
Google Developers

Agent Anomaly Detection, now in Private Preview on the Gemini Enterprise Agent Platform

Agent Anomaly Detection is a new, out-of-band oversight layer for the Gemini Enterprise Agent Platform that analyzes OpenTelemetry traces and tool calls to catch behavioral risks without adding runtime latency to live requests. It utilizes a multi-tiered detection pipeline—combining lightweight statistical scanning with deep LLM-based reasoning—to identify logical anomalies and policy violations grounded in the OWASP Agentic Top 10. Developers can triage these automated findings within Security

Read source article
Google Developers

Build zero-trust AI agents that judge intent, not just syntax

This blog post explores how to transition AI agents from static, build-time security controls to dynamic runtime governance using the Gemini Enterprise Agent Platform. It highlights three primary managed defenses: Model Armor for screening edge prompts, Semantic Governance Policies for evaluating tool intent against business rules, and Agent Anomaly Detection for catching multi-turn exploits. By shifting these capabilities to the platform level, security administrators can dynamically enforce po

Read source article
The Guardian AIBusiness

The week that changed maths for ever – podcast

<p>In early September, OpenAI announced it had solved a major mathematics problem that has stumped humans for nearly a century. The news left mathematicians reeling, and many expressed concern over what will be left for humans as AI becomes ever more adept at unravelling complex problems. Now 25 recipients of the Fields medal – often called the Nobel prize for maths – have signed an open letter expressing their fears of a ‘severe misalignment’ between AI companies and their field. To find out ho

Read source article
The week that changed maths for ever – podcast
arXiv cs.AIResearch

Making AI-Assisted Claims Independently Challengeable: Publication Authority and a Protocol for Falsifiable Publication Records

arXiv:2609.17631v1 Announce Type: new Abstract: AI-assisted claims can appear authoritative when evidence, analysis, human authorization, presentation, and correction history refer to different states. Provenance, attestation, and transparency expose history but alone do not specify the publication transition examined here. We develop Publication Authority as an exact-state, non-transferable, single-use publication capability and instantiate it in PAC-2026 (Publication-Accountability Calculus),

Read source article
arXiv cs.AIResearch

EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents

arXiv:2609.17632v1 Announce Type: new Abstract: Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment. This limits their ability to adapt how they gather evidence, invoke tools, verify signals, and manage risk under changing market regimes. We introduce EvolveTrade, a self-evolving framework that treats the system prompt of a tool-using trad

Read source article
arXiv cs.AIResearch

One Color Preprocessing Improves DSATUR

arXiv:2609.17633v1 Announce Type: new Abstract: The Graph Coloring Problem (GCP) is NP-hard and DSATUR stands as one of the fastest heuristics for it despite producing colorings that typically use more colors than state-of-the-art coloring algorithms. We propose SSLD (Semidefinite Spectral Learning with DSATUR), which improves DSATUR by preprocessing a first good color class before letting DSATUR complete coloring the rest of the given graph. We obtain this color class from a Semidefinite Progra

Read source article
arXiv cs.AIResearch

GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents

arXiv:2609.17695v1 Announce Type: new Abstract: A large language model (LLM) agent can follow more graph paths without acquiring more independent evidence. GraphEcho tests whether agents mistake these repeated encounters for additional corroboration. The benchmark varies path counts and evidential origins while holding evidence content fixed, and evaluates both judgments and active exploration. Controlled synthetic experiments reveal model-dependent judgment shifts, but redundant supporting path

Read source article
arXiv cs.AIResearch

NeMo Data Designer: An Extensible Framework for Multimodal Synthetic Data Generation

arXiv:2609.17699v1 Announce Type: new Abstract: We present NeMo Data Designer (NDD), an open-source, general-purpose framework for multi-modal synthetic data generation (SDG). Designed to be intuitive to use, NDD provides a declarative configuration format in which human and/or agent users define each dataset column, with column types spanning text, code, structured outputs, images, embeddings, and statistical samplers that are explicitly configured to steer dataset diversity. Additional column

Read source article
arXiv cs.AIResearch

Imitation Learning for Autonomous Driving in CARLA

arXiv:2609.17757v1 Announce Type: new Abstract: Behavioral cloning trains a policy offline on expert demonstrations, but deployment is closed loop: each action affects the observations the policy receives next. We study how much closed-loop driving competence a compact multimodal policy can acquire from offline demonstrations in the CARLA simulator. The policy uses five-frame histories of RGB images, LiDAR, vehicle telemetry, and lane waypoints to predict throttle, brake, and steering at 20 Hz.

Read source article
arXiv cs.AIResearch

SAGE: Governed Artifact Generation from Enterprise Guidelines

arXiv:2609.17775v1 Announce Type: new Abstract: Enterprise guideline documents mix narrative text, complex tables, and embedded images, and converting them into structured work artifacts still takes two to three days of manual effort each. Current language and vision-language models extract from such documents but offer no governed workflow beyond extraction: no validation, no consistency checking, no traceable artifact generation. We introduce SAGE, a governed multi-stage LLM pipeline organized

Read source article
arXiv cs.AIResearch

FairCompressAgent: An Agentic Framework for Fairness-Aware Model Compression for FPGA Deployment

arXiv:2609.17786v1 Announce Type: new Abstract: Fairness-aware model compression requires selecting methods and configurations that balance accuracy, fairness, and deployment cost. These decisions become more difficult when compression methods are composed or the user's requirements change. In this paper, we propose FairCompressAgent (FCA), an agentic framework that integrates fairness-aware pruning, incremental quantization, and sparse low-rank factorization through a common operator interface.

Read source article
arXiv cs.AIResearch

A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning

arXiv:2609.17804v1 Announce Type: new Abstract: Large language models solve grade-school math word problems with high accuracy, yet a single irrelevant clause inserted into the problem can collapse it. We reconcile these observations with a mechanistic account. We show that the model's internal computation decomposes into a four-stage sequential pipeline, Schema Abstraction, Operation Planning, Operand Binding, and Computation, each stage producing a distinct intermediate representation in an id

Read source article
arXiv cs.AIResearch

Learning Heterogeneous Preferences

arXiv:2609.17847v1 Announce Type: new Abstract: Learning from human feedback has become a central paradigm for training modern AI systems, where models of human utility are used as reward models in policy learning. Existing methods typically assume a \emph{universal utility} function shared across a population and treat disagreement between annotators as stochastic variation. While suitable for objective tasks, this assumption breaks down in subjective domains where preferences vary systematical

Read source article
arXiv cs.AIResearch

SNOMED CT Concept Recommendation from Masked Clinical Context

arXiv:2609.17855v1 Announce Type: new Abstract: Standardizing clinical language to SNOMED CT supports interoperability, analytics, and reusable phenotyping, but concept recommendation remains difficult when relevant concepts are rare or absent from training data. We present a masked-concept recommendation benchmark using the SNOMED CT Entity Linking Challenge v1.2.1 data derived from MIMIC-IV-Note. The dataset contains 75,491 annotations across 272 discharge summaries, with 204 notes used for tr

Read source article
arXiv cs.AIResearch

The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?

arXiv:2609.17863v1 Announce Type: new Abstract: LLM inference optimizations report speedups on different models, GPUs, prompts, and quality metrics, making them hard to compare or combine. We build a cost, quality, and latency Pareto atlas to identify the best configurations for different deployment constraints. Since exhaustive testing is impractical, we measure 54 configurations of Qwen2.5-7B-Instruct running on vLLM 0.12 across L4, A100, and H100 GPUs and use these anchors to calibrate a simu

Read source article
arXiv cs.AIResearch

Do Frontier Models Seek Safety Evidence Before Acting?

arXiv:2609.17865v1 Announce Type: new Abstract: Frontier models are often evaluated on how they respond to safety information once it is already in context. We study an earlier decision point: whether models choose to acquire safety-relevant evidence before acting. We introduce SAFE, a controlled benchmark in which models make deployment decisions with optional evidence that varies in retrieval cost, probability, severity, and presentation. Across GPT-5.5, o3, Claude Opus 4.8, and Claude Sonnet

Read source article
arXiv cs.AIResearch

ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software

arXiv:2609.17885v1 Announce Type: new Abstract: Computer-use agents that operate through screenshots and simulated actions are advancing rapidly, yet their evaluation remains anchored to general desktop and web tasks. Enterprise Resource Planning (ERP) systems run the finance, procurement, inventory, and customer operations of organizations worldwide, and pose distinct challenges for computer-use agents: dense interfaces, coordinated multi-step interactions, and errors that alter persistent busi

Read source article
arXiv cs.AIResearch

OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning

arXiv:2609.17890v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead. Pruning can reduce this cost, but its effectiveness depends on the calibration data used to estimate parameter importance. Recent work calibrates on the model's own rollouts instead of generic dataset, but treats all reasoning tokens uniformly, regardless of whether they contribute to successful reasoning. As a result, prun

Read source article
arXiv cs.AIResearch

Collaborative Memory for Multi-Agent VLM Systems

arXiv:2609.17921v1 Announce Type: new Abstract: Vision-language model (VLM) agents combine specialized perception, tools, and reasoning to address complex visual tasks. In multi-agent settings, different agents inspect different image regions, video frames, or visual representations, so collaboration extends beyond distributed reasoning to distributed perception. This makes shared visual context a central problem in VLM agent collaboration. In this paper, we frame memory hierarchy, cross-agent s

Read source article
arXiv cs.AIResearch

Measuring AI Leadership: Development and Validation of a Multidimensional Measure for AI-Native Organizations

arXiv:2609.17965v1 Announce Type: new Abstract: AI is changing what leaders must judge, explain, learn, and coordinate, yet existing measures do not capture these behaviors at the level needed to study leadership in AI-enabled work. We develop the AI Leadership Battery, which organizes 36 behaviorally specific subdimensions into 11 theory-specified content families. Following established scale-development procedures, the research used deductive item generation; content validation of definitional

Read source article
arXiv cs.AIResearch

Memory Has Geometry: Non-Uniform Geometric Memory for Long-Horizon Personalized AI

arXiv:2609.17969v1 Announce Type: new Abstract: Long-term memory is becoming a core substrate for personalized AI, yet most systems still represent personalization as discrete records in a largely static latent space, accessed under one global similarity notion. For data mining, this creates a mismatch: the evidence is a temporal event stream, while the dominant abstraction is a searchable record set. We argue that long-horizon personalization should instead model memory as a user-specific dynam

Read source article