AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

31239 stories from 30+ sources, refreshed continuously.

arXiv cs.CL (NLP)Research

Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding

arXiv:2609.09338v1 Announce Type: new Abstract: Speculative decoding is critical for accelerating LLM inference. However, the speedup is fragile: drafters are typically trained against a narrow distribution for a single target model, and their acceptance rate collapses under workload shifts. This is a striking inversion of modern LLM development, where target models are valued precisely for the broad generalization they acquire through large-scale pretraining. We argue that the natural remedy, p

Read source article
arXiv cs.CL (NLP)Research

SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection

arXiv:2609.09349v1 Announce Type: new Abstract: Modern LLMs demonstrate impressive multilingual performance, yet standard benchmarks primarily reward selecting correct answers rather than evaluating genuine factual understanding. We introduce Systematic Wikidata-based Object-Relation Distortion (SWORD), a benchmark that evaluates whether models consistently reject factual errors across languages. SWORD generates syntactically well-formed but factually incorrect statements in eight widely spoken

Read source article
arXiv cs.CL (NLP)Research

Auditable Emergency Triage for Maternal and Newborn Care in India

arXiv:2609.09356v1 Announce Type: new Abstract: At Noora Health, our nurses answer more than 50,000 medical queries per month on our WhatsApp-based service that provides caregivers with on-demand support. Their most time-critical task is emergency triage: deciding which queries need immediate in-person attention. To support them, we built a system that uses a large language model (LLM) to classify whether a message is an emergency and provide a rationale for interpretability. But the system was

Read source article
arXiv cs.CL (NLP)Research

Do LLMs Make More Mistakes If They Do Not Believe the Input Data?

arXiv:2609.09363v1 Announce Type: new Abstract: Large language models (LLMs) are prone to hallucinating or misinterpreting facts, which impairs their usability in retrieval-augmented generation or data-to-text systems. We analyse how faithfulness of LLMs to provided context depends on how plausible they perceive the context to be (context-memory conflict). To better identify error patterns, we make use of the increased difficulty of non-English and low-resource language text generation and input

Read source article
arXiv cs.CL (NLP)Research

Benchmarking Hybrid Deep Research Across Database Querying and Web Search

arXiv:2609.09410v1 Announce Type: new Abstract: While autonomous agents have made significant strides in "deep research" by iteratively navigating the open web to synthesize information, real-world problem-solving is rarely confined to a single environment. Complex analytical tasks inherently require agents to weave together evidence from both ambiguous unstructured text (e.g., the open web) and highly precise structured data (e.g., relational databases). However, existing benchmarks evaluate th

Read source article
arXiv cs.CL (NLP)Research

Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements

arXiv:2609.09425v1 Announce Type: new Abstract: Educational data filters have become a practical way to improve language-model pre-training, but most filters treat educational value as a single scalar property. This may be too broad for some applications, especially if the data set already features a high density of educational material. Useful learning material needs to be accurate, engaging, well structured, and appropriate for the intended audience and application (e.g. learner- vs teacher-fa

Read source article
arXiv cs.CL (NLP)Research

The Mutations of Machine Speech

arXiv:2609.09496v1 Announce Type: new Abstract: Algorithmic outputs now populate the digital environments through which contemporary life is organized. The role of law in facilitating and constituting (rather than merely responding to) these processes is gaining increasing traction across scholarly accounts. This inquiry traces the evolution of algorithmic outputs attending to their legal underpinnings and social implications, surfacing the mutations of machine speech. The first mutation redefin

Read source article
arXiv cs.CL (NLP)Research

BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models

arXiv:2609.09554v1 Announce Type: new Abstract: We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such as Whisper have revolutionized ASR, but most prominent models are highly multilingual. As a result, these models often perform poorly on languages less well-represented in their training set. While it has long been known that effective language adapta

Read source article
arXiv cs.CL (NLP)Research

Towards Automatic Evolution Tree Generation from Citation Graphs

arXiv:2609.09561v1 Announce Type: new Abstract: Surveys remain the primary way researchers grasp the lineage of methods within an AI subfield, but they scale poorly against the current rate of publication. Existing taxonomy-induction methods are largely leaf-bound and time-agnostic; they tend to force transitional papers into mature leaves and can create topological inversions between ancestors and descendants. We propose EvoTree, a staged framework that decouples conceptual backbone learning fr

Read source article
arXiv cs.CL (NLP)Research

Reproducing Omitted Temporal Expressions in Japanese News for Retrieval-Augmented Applications

arXiv:2609.09569v1 Announce Type: new Abstract: News articles often contain omitted temporal expressions, such as day-only or month-only mentions, which must be interpreted with reference to the publication date. When such articles are indexed or processed as standalone text in search and retrieval-augmented generation (RAG) systems, these omissions can cause temporal mismatches and unstable interpretation by large language models. We focus on reproducing omitted temporal expressions as concrete

Read source article
arXiv cs.CL (NLP)Research

Beyond Top Words: MonoTM for Topic Modeling with Interpretable Monosemantic Features

arXiv:2609.09575v1 Announce Type: new Abstract: Topic models summarize large text corpora, but top-ranked words often provide only a limited representation of topic semantics. Sparse autoencoders (SAEs) offer a way to move beyond word-level descriptors by extracting interpretable features from dense representations, yet how feature interpretability relates to topic-inference quality remains unclear. We introduce \textbf{MonoTM}, an interpretable topic modeling framework that decouples these role

Read source article
arXiv cs.CL (NLP)Research

SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia

arXiv:2609.09672v1 Announce Type: new Abstract: The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leaving Southeast Asian (SEA) languages critically underrepresented. We introduce SEA-SpeechBench, to the best of our knowledge, the first large-scale multitask benchmark that evaluates speech understanding in 11 SEA languages through 97,194 samples across

Read source article
arXiv cs.CL (NLP)Research

Looped GPT-BERT: Trading Parameters for Computation in Small Language Modeling

arXiv:2609.09691v1 Announce Type: new Abstract: When training data are limited, increasing parameter count is not the only way to improve language-model performance. A small parameter set, when repeatedly applied, can also deliver comparable performance. We study Looped GPT-BERT in the BabyLM 2026 Strict-small setting, combining GPT-BERT's masked next-token and causal language-modeling objectives with depth-wise parameter sharing. We train on a preprocessed 7.48M-word English corpus and compare

Read source article
arXiv cs.CL (NLP)Research

When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination

arXiv:2609.09696v1 Announce Type: new Abstract: Large language models are increasingly proposed as automated auditors of document quality, yet their reliability as detectors of planted errors is poorly characterised. We construct a contaminated corpus of 150 academic papers spanning supply chain management and medical research, injecting 450 known contaminants of three types: typographical corruption, semantic reversal, and absurd out-of-context insertion. We then evaluate Google Gemini 3.0 Pro'

Read source article
arXiv cs.CL (NLP)Research

StreamAlign: Streaming Text-Aligned Speech Tokenization

arXiv:2609.09719v1 Announce Type: new Abstract: Text-aligned speech tokenization methods have emerged to better align speech tokens with LLM token spaces, enabling more effective utilization of pretrained LLMs. However, they rely on offline automatic speech recognition (ASR), leading to two key limitations: (i) the need for complete utterances before tokenization, precluding real-time streaming, and (ii) vocabulary mismatch between ASR and LLMs, which reduces acoustic granularity from the subwor

Read source article
arXiv cs.CL (NLP)Research

SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design

arXiv:2609.09764v1 Announce Type: new Abstract: Social intelligence enables agents to read social context, infer intent, and adapt over sustained dialogue. As language models become autonomous collaborators, it is central to building effective and trustworthy human-AI interaction. Existing reinforcement learning methods optimize single-turn utterances and sparse outcome rewards, producing short-sighted policies that struggle to manage goal-relationship tensions across multi-turn interactions. We

Read source article
arXiv cs.CL (NLP)Research

CARRE: Counterfactual Action Retrieval and Reason Evaluation for Explainable Churn Prescription

arXiv:2609.09766v1 Announce Type: new Abstract: Churn models typically identify high-risk customers but do not specify which feasible retention action should be considered or why that action is appropriate. We present CARRE (Counterfactual Action Retrieval and Reason Evaluation), a three-stage framework that combines retrieval-augmented candidate generation, cost-aware counterfactual scoring, and large language model (LLM) reasoning. CARRE retrieves a predefined catalog of retention actions, est

Read source article
arXiv cs.CL (NLP)Research

SymbolicLight V2: Hybrid Neuromorphic Architecture and Sparse Execution for Low-Energy Language Inference

arXiv:2609.09772v1 Announce Type: new Abstract: SymbolicLight V2 combines sparse event computation with continuous-state processing in a hybrid neuromorphic language architecture. Extending V1's spike-gated dual paths, it adds graded signed events at further projections and softmax-free local attention. We implement the 194M-parameter model on an Alveo U50C FPGA using digital fixed-point arithmetic and on an ARM CPU using sparse integer execution. Across three same-checkpoint FPGA implementation

Read source article
Product Hunt — The best new products, every day

Modeinspect

<p> 99 Days Free AI Credits - Design product UI in your codebase </p> <p> <a href="https://www.producthunt.com/products/modeinspect-1-0?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1246291?app_id=339">Link</a> </p>

Read source article
Hacker News LLMLLMs

Training a 3.8B LLM to 0.384 CORE for $998

Article URL: https://hugovergnes.github.io/little-lm-3-8b/ Comments URL: https://news.ycombinator.com/item?id=49637435 Points: 74 # Comments: 13

Read source article
Simon WillisonLLMs

Quoting Calif Research

<blockquote cite="https://calif.io/research/weworm"><p>Today, we're releasing a demo of WeWorm, the first zero-click worm to spread through WeChat calls across iOS and Android. [...]</p> <p>The victim does not need to answer the call, or interact with their phone at all. Even if they do answer, they hear nothing, and the exploit still succeeds. [...]</p> <p>Working with AI, our team found the bug and wrote the first remote code execution (RCE) exploit in about two days. Building the worm took on

Read source article
The Guardian AIBusiness

Lawmakers blast AI companies after researcher warns of human extinction by 2030

<p>Former Anthropic employee Jacob Coxon said AI will become ‘superhuman systems’ that can cause human extinction by the end of the decade</p><p>Just a day after three Anthropic researchers warned that artificial intelligence could kill off humanity within the decade, lawmakers have begun lashing out about the risks of the burgeoning technology.</p><p>Ted Cruz, a republican senator from Texas, said in an <a href="https://www.facebook.com/reel/1712597293170453">interview</a> on ABC’s The View, th

Read source article
Lawmakers blast AI companies after researcher warns of human extinction by 2030
PyTorch

PyTorch Conference China 2026: Advancing the Open Source AI Stack

PyTorch Conference China 2026 brought the PyTorch community together in Shanghai on September 8–9 alongside KubeCon + CloudNativeCon and OpenInfra Summit, following sponsor-hosted co-located events on September 7. Technical discussions...

Read source article
PyTorch Conference China 2026: Advancing the Open Source AI Stack
Vercel Blog

Build with OpenAI Agents API on Vercel

You can now build and deploy long-running, tool-using agents with the OpenAI Agents API on Vercel. OpenAI manages the agent loop and session state, while Vercel hosts the application and connects each session to Vercel Sandbox for code execution and file access. With this integration, you get: An OpenAI-managed agent loop and session state Reliable Sandbox creation and reconnection through signed OpenAI webhooks and Vercel Queues An isolated execution environment for every agent session A persis

Read source article
Vercel Blog

Tako Search is free on AI Gateway through September 30

Tako Search is free exclusively on AI Gateway through September 30. It lets AI models search Tako's curated data and the live web, filter web results by domain or publication date, and use the results to answer questions with current information, citations, and visualizations. After September 30, searches are billed at standard rates. The same integration works with any model on AI Gateway, so you can switch models without changing your search setup. You also don't need a separate Tako account o

Read source article
OpenAI BlogLLMs

Introducing the Agents API

Build and launch cloud agents with the Agents API, a managed service powered by the Codex harness for orchestration, long-running sessions, and tool use.

Read source article
Eric Jang

Smooth Exponentials for Robotics

I recently reread Dario Amodei’s essays "Machines of Loving Grace" and "Adolescence of Technology." In these essays and his interviews, Dario speaks often of a "smooth exponential" of AI capabilities, as if machine intelligence "emerge[s] spontaneously from the right combination of data and raw computation."

Read source article
Simon WillisonLLMs

.blend URL Viewer

<p><strong>Tool:</strong> <a href="https://tools.simonwillison.net/blender-viewer">.blend URL Viewer</a></p> <p>I'm continuing to have a lot of fun with GPT-6 Astra and Blender (see <a href="https://til.simonwillison.net/llms/blender-coding-agents-macos">my TIL</a>).</p> <p>As a big fan of the <a href="https://en.wikipedia.org/wiki/Faberg%C3%A9_egg">Imperial Fabergé Easter eggs</a>, I've always thought it would be fun to make some new ones that celebrate popular culture.</p> <p>Yesterday I decid

Read source article
MashableGeneral Tech

ChatGPT 80s trend: Prompts to try it yourself

Try the ChatGPT ’80s photo trend with prompts for retro portraits, arcade dates and roller-rink nights. Here’s how to give your photos a throwback makeover.

Read source article
ChatGPT 80s trend: Prompts to try it yourself
AWS Machine Learning Blog

Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

Learn how to deploy Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight model, on Amazon SageMaker HyperPod with vLLM. This walkthrough covers cluster provisioning, NVFP4 quantization, and an OpenAI-compatible endpoint with built-in reasoning, tool calling, and native MTP speculative decoding.

Read source article