AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

31240 stories from 30+ sources, refreshed continuously.

arXiv cs.CL (NLP)Research

MA-RAG: Multi-Agent Retrieval-Augmented Generation for Query-Driven Summarization of Longitudinal Parkinson's Disease Assessments

arXiv:2608.28624v1 Announce Type: new Abstract: Accurate interpretation of single-visit and longitudinal clinical assessments for Parkinson's disease is time-consuming and often depends on specialist expertise. Although large language models (LLMs) can generate natural language summaries, they frequently lack domain-specific clinical grounding and struggle to produce factually correct and temporally consistent responses for structured longitudinal assessment data. To address these limitations, w

Read source article
arXiv cs.CL (NLP)Research

Asymmetric Within-Document Predictive Learning for Scientific Document Representation

arXiv:2608.28625v1 Announce Type: new Abstract: We study predictive pretraining for scientific document representation using the discourse structure of papers. We propose SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict method representations, and method representations are used to predict conclusion representations. Experiments on RELISH, high-influence citation, SciDocs, and cite prediction show

Read source article
arXiv cs.CL (NLP)Research

Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

arXiv:2608.28626v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models' training cutoffs. Manuscripts were presented to both models with author identities blinded,

Read source article
arXiv cs.CL (NLP)Research

Intelligent Identification and Repair of Design Defects in BIM via Domain-Specific Large Language Models

arXiv:2608.28629v1 Announce Type: new Abstract: Existing methods lack a generalized approach to efficiently identify and resolve the diversity of design defects in BIM. Therefore, this study proposes an integrated framework to identify and repair various defects in BIM via domain-specific LLMs. Firstly, a BIM-to-Text method with component-balanced chunking is introduced to bridge BIM data with LLMs. Then, prompt learning with rule injection, few-shot prompting and RAG is proposed to identify def

Read source article
arXiv cs.CL (NLP)Research

Enabling Proactive Spoken Turns via a Generalized Style-Aware Full-Duplex Framework

arXiv:2608.28630v1 Announce Type: new Abstract: Compared with half-duplex dialogue systems where the system waits for user turn completion before it responds, natural full-duplex dialogue systems require agents to act proactively in real time, including timely interruptions and backchannels. This creates a key challenge: improving turn timing without sacrificing response quality. To address limitations in realistic proactive turn-taking, we build a generalized style-aware full-duplex framework w

Read source article
arXiv cs.CL (NLP)Research

PAUSE: Editable Strategy Artifacts for Long-Form Cultural Story Adaptation

arXiv:2608.28633v1 Announce Type: new Abstract: Generative AI systems increasingly mediate cultural adaptation, but their cultural decisions are often hidden inside prompts, transient model plans, or final prose. We study PAUSE (Pause-And-Update Strategy Editing), an intervention that exposes an editable adaptation strategy as a human control surface for cultural decisions in long-form story adaptation. The strategy is a structured artifact that can be inspected, edited, and then projected throu

Read source article
arXiv cs.CL (NLP)Research

Do MLLMs Really Understand Low-Resource Khmer Documents? A Pilot Study on Khmer Document VQA

arXiv:2608.28635v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) have advanced document understanding, visual question answering, and text extraction. However, their reliability in low-resource, non-Latin settings remains uncertain. Khmer form documents present particular challenges because they contain complex script forms, mixed Khmer-English fields, and monetary values in both Cambodian Riel and US Dollars. Available resources for Khmer Document VQA are also lim

Read source article
arXiv cs.CL (NLP)Research

PromptKWS: A Novel Prompt-Guided Open-Vocabulary Keyword Spotting Framework

arXiv:2608.28640v1 Announce Type: new Abstract: In this paper, we present PromptKWS, a novel Prompt-guided keyword spotting (KWS) framework to improve the accuracy of open vocabulary KWS systems. In specific terms, we introduce the Prompt Phrases Prediction Network (PPN), an encoder-decoder architecture designed to effectively extract keyword prompts embeddings. we employ the PPN encoder to encode the keyword prompts and infuse the prompt embedding into the Prompt-guided KWS encoder by utilizing

Read source article
arXiv cs.CL (NLP)Research

Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture

arXiv:2608.28641v1 Announce Type: new Abstract: Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Terminal-Bench-LILT, a suite of 300 authentic coding tasks in ten languages: Arabic, Czech, German, Spanish, Hindi, Japanese, Korean, Serbian, Turkish, and Chinese. Each task targets issues specific to non-English software development that have no direct English equivalent, e.g., internationalization, encodi

Read source article
arXiv cs.CL (NLP)Research

Cross Lingual Transfer in Tulu Legal Comprehension: Script-Dependent Improvement and RAG-Induced Knowledge Conflict

arXiv:2608.28645v1 Announce Type: new Abstract: Low-resource languages without an adequate training corpus often use a related, higher-resource language as a scaffold for comprehension. Still, there is a need to develop rigorous evaluation methods to identify when models fail in cross lingual low-resource environments. Using the legal domain as a backdrop, three models (Llama3, Hex-1, Sarvam) were tested on the ability to classify legal complaints written in a low resource Dravidian language (Tu

Read source article
arXiv cs.CL (NLP)Research

Can Large Language Models Identify Meaningful Touchpoints in Conversion Attribution?

arXiv:2608.28649v1 Announce Type: new Abstract: Touchpoint selection in conversion attribution, namely identifying meaningful touchpoints contributing to conversions, is essential for e-commerce recommendation and online advertising. Current selection methods rely heavily on collaborative-filtering-based heuristics, which fail to align with user-perceived semantic intent. Through human annotation, we reveal a significant semantic gap: many implicitly-related, semantically relevant touchpoints re

Read source article
arXiv cs.CL (NLP)Research

Test-Time Scaling for Scientific Equation Discovery

arXiv:2608.28660v1 Announce Type: new Abstract: Test-time scaling (TTS) improves language model reasoning by allocating additional test-time compute, but prior work mainly studies closed-ended tasks such as math and coding. We study TTS for automated equation discovery, an open-ended setting where models search over candidate equations and rely on observed datapoints for feedback. We formulate LLM-driven equation discovery as an iterative search process that unifies Best-of-N, sequential refinem

Read source article
arXiv cs.CL (NLP)Research

GreenBench: Benchmarking Energy Efficiency and Carbon Footprint of Open-Source LLM Inference on Apple Silicon

arXiv:2608.28667v1 Announce Type: new Abstract: The rapid proliferation of Large Language Models (LLMs) has raised concerns about their environmental impact during inference. While Green AI research has focused on datacenter GPUs and embedded platforms, the energy profile of LLM inference on Apple Silicon, with its unified memory architecture, remains unstudied. This paper presents GreenBench, a benchmarking framework that evaluates the energy efficiency, throughput, and carbon footprint of five

Read source article
arXiv cs.CL (NLP)Research

ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering

arXiv:2608.28707v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in Visual Question Answering (VQA), yet they continue to struggle with questions requiring precise spatial reasoning and fine-grained visual understanding. These limitations often manifest as object, attribute, and spatial hallucinations, where models generate confident but visually unsupported responses due to insufficient region-level and fine-grained visual grounding. To

Read source article
arXiv cs.CL (NLP)Research

A rigor-matched audit of periodic-step layer skipping for efficient llm inference: conflayers versus swift, with a supplemental analysis of trained routing alternatives

arXiv:2608.28846v1 Announce Type: new Abstract: Layer-skipping methods for efficient LLM inference decide, at some granularity, which transformer layers to execute for a given input. We present a rigor-matched, three-seed audit of two periodic-step, search-based methods that make this decision online at inference time and re-evaluate it every few generation steps: a confidence-gated early-exit baseline (ConfLayers) and genuine self-speculative decoding (SWIFT, Xia et al. 2024), together with van

Read source article
arXiv cs.CL (NLP)Research

Latent-Space Intervention for Cross-Lingual Factual Consistency: Consistency Improvements without Accuracy Drops

arXiv:2608.28860v1 Announce Type: new Abstract: Large Language Models (LLMs) often answer the same factual question differently across languages. We study whether cross-lingual latent-space intervention can reduce this inconsistency. We train layer-specific autoencoders on parallel multilingual representations and apply inference-time corrections to factual QA prompts. We find that latent intervention improves geometric alignment between languages, and that this improvement translates into consi

Read source article
Hacker News FrontTools

GPU World

Article URL: https://www.gpuworld.org/ Comments URL: https://news.ycombinator.com/item?id=49517584 Points: 150 # Comments: 63

Read source article
Product HuntTools

Cosmic Agent Plugins

<p> Connect Cosmic agents to any service with an MCP server </p> <p> <a href="https://www.producthunt.com/products/cosmic?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1238046?app_id=339">Link</a> </p>

Read source article
Hacker News AILLMs

OpenAI to Cut Off AI Models for SpaceX-Owned Cursor

Article URL: https://www.reuters.com/business/media-telecom/openai-end-partnership-with-spacexs-cursor-2026-08-29/ Comments URL: https://news.ycombinator.com/item?id=49516583 Points: 4 # Comments: 1

Read source article
Svelte Blog

What’s new in Svelte: September 2026

<p>This month, Svelte 5.57 shipped with new <code>SvelteMap</code> methods and a few quality-of-life additions while SvelteKit 3 got closer to the finish line with its Release Candidate.</p> <p>The <code>sv</code> CLI also got a new <code>ai-tools</code> add-on that replaces the old <code>mcp</code> one, and <code>sv@next</code> now ships a task-based <code>sveltekit-3</code> migration for existing apps.</p> <p>Let’s dive a bit deeper!</p> <h2 id="What's-new-in-Svelte"><span>What’s new in Svelte

Read source article
Vercel Blog

Qwen 3.8 Max 0902 now available on AI Gateway

Qwen 3.8 Max 0902 from Alibaba is now available on AI Gateway. This is a new snapshot of Qwen 3.8 Max, with the gains concentrated in coding on larger projects, long-horizon work that runs without supervision, and agent runs. Vision handling is more accurate on charts and dense documents. To use Qwen 3.8 Max 0902, set model to alibaba/qwen3.8-max-0902 : The dated ID pins this snapshot, so a later release will not change what your requests run against. To move existing traffic onto it without a c

Read source article
Vercel Blog

Claude Fable 5.1 now available on AI Gateway

Claude Fable 5.1 from Anthropic is now available on AI Gateway. Fable 5.1 improvements compared to previous Claude models are concentrated in long, multi-stage work like agentic coding, knowledge work, and research that takes several rounds of searching and following up. Anthropic ships Fable 5.1 with cybersecurity and biology safety classifiers enabled. Finding vulnerabilities in source code is allowed, but some routine coding and debugging may still be refused. To ensure requests are still ser

Read source article
Artificial Intelligence News — Newsletter on Deep Learning & AI

AI Weekly Issue #528: What are companies building with AI? An Applied AI Deep Dive

We went looking for what companies are actually building with AI. The answer was not more chatbots. It was drones carrying diagnostic samples, driverless Frito-Lay trucks, AI-guided flight paths, repair copilots, and rugged GPU laptops in Ukraine. We reviewed 136 use cases from the last 20 days. The biggest surprise: only 38 included a reported outcome.

Read source article
Dan Luu

How accurate have Ed Zitron's AI skeptic predictions been?

<p>I was curious how well the predictions of the most widely cited AI skeptic I've seen (Ed Zitron) have done, so I looked at how his predictions panned out. To disclose my own biases, I've never had a particularly strong pro or anti AI progress position. For example, <a href="https://danluu.com/futurist-predictions/">in 2022, I did a comprehensive look at predictions Futurists made, including well-respected folks like Kurzweil</a> and found them to be generally wrong on both the prediction resu

Read source article
Simon WillisonLLMs

Introducing wrapture

<p><strong><a href="https://grahamdumpleton.me/posts/2026/08/introducing-wrapture/">Introducing wrapture</a></strong></p> New from Graham Dumpleton (of <a href="https://pypi.org/project/wrapt/">wrapt</a>, mod_wsgi, and New Relic's Python agent fame), who describes Wrapture as taking the monkeypatching ideas from wrapt and extending them to apply to testing and tracing at the same time.</p> <p>Wrapture (<a href="https://wrapture.readthedocs.io/">full documentation here</a>) makes it easy to wrap

Read source article
Hacker News AILLMs

AI and Employment: So Far, So Good

Article URL: https://marginalrevolution.com/marginalrevolution/2026/08/ai-and-employment-what-the-firms-say.html Comments URL: https://news.ycombinator.com/item?id=49516267 Points: 3 # Comments: 0

Read source article
Hacker News AILLMs

Show HN: Saccade – Live semantic browser truth for AI agents

I build one this application is to resolve the most of the agent cannot handle the brother they extremely slow. So I come up with this idea, we install one of the extntion on chrome or edge to give continuously compile the tabs authorized by the user into semantically meaningful objects with stable identities, and push page changes as deltas to the local Node.js Broker. The Agent reads the full truth or delta of the specified tab via MCP and executes actions using object IDs bound to document. T

Read source article
Product Hunt — The best new products, every day

OpenMarket

<p> Multi-agent marketplace where proof decides who wins </p> <p> <a href="https://www.producthunt.com/products/openmarket?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1237958?app_id=339">Link</a> </p>

Read source article
Hacker News AILLMs

AI giants lean into health care to stall public backlash

Article URL: https://www.axios.com/2026/08/31/ai-health-care-cures-disease-anthropic-open-ai Comments URL: https://news.ycombinator.com/item?id=49516052 Points: 2 # Comments: 0

Read source article
Hacker News AILLMs

ChatGPT becomes first AI chatbot to face tougher EU rules

Article URL: https://www.lemonde.fr/en/pixels/article/2026/08/31/chatgpt-becomes-first-ai-chatbot-to-face-tougher-eu-rules_6757011_13.html Comments URL: https://news.ycombinator.com/item?id=49515976 Points: 5 # Comments: 0

Read source article
AWS Machine Learning Blog

Connect an AgentCore Runtime hosted MCP server to Amazon Quick

In this post, you will learn how to deploy and host your MCP server in AgentCore Runtime and integrate it with Amazon Quick, along with the prerequisites. With this pattern, you promote reusability and avoid duplication of AI tools, so clients can reuse commonly used tools and agents exposed through an MCP server instead of authoring them from scratch again. Your customers get a way to use your product inside Amazon Quick (chat agents and workflows) without building custom connectors for every u

Read source article
Hacker News AILLMs

AI Job Scout and Rate Matcher

Article URL: https://github.com/Deviad/job-hunter Comments URL: https://news.ycombinator.com/item?id=49515381 Points: 1 # Comments: 0

Read source article