AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

33443 stories from 30+ sources, refreshed continuously.

arXiv cs.LGResearch

Bandwidth Selection in Kernel Density Estimation for Model Calibration

arXiv:2606.29925v1 Announce Type: new Abstract: As deep learning models are increasingly deployed in high-stakes applications, providing well-calibrated uncertainty estimates has become as critical as achieving high predictive accuracy. While Kernel Density Estimation (KDE) has emerged as a smooth and continuous alternative to traditional binning for quantifying miscalibration, its reliability is heavily dependent on the choice of the kernel bandwidth. Standard selection techniques, such as Maxi

Read source article
arXiv cs.CL (NLP)Research

Multi-Agentic System Leveraging Open-Source LLMs to Mitigate Disinformation Threats

arXiv:2606.30259v1 Announce Type: new Abstract: In contemporary societies, the threat of disinformation has reached alarming levels, exacerbated by the proliferation of electronic communication, social media, and advancements in artificial intelligence. As a result, there is an urgent need to develop effective countermeasures to mitigate this menace. However, the sheer scale of the problem renders manual fact-checking and human-based verification inadequate, underscoring the necessity for automa

Read source article
arXiv cs.CL (NLP)Research

REAR: Test-time Preference Realignment through Reward Decomposition

arXiv:2606.30339v1 Announce Type: new Abstract: Aligning large language models (LLMs) with diverse user preferences is a critical yet challenging task. While post-training methods can adapt models to specific needs, they often require costly data curation and additional training. Test-time scaling (TTS) presents an efficient, training-free alternative, but its application has been largely limited to verifiable domains like mathematics and coding, where response correctness is easily judged. To e

Read source article
arXiv cs.CL (NLP)Research

MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

arXiv:2606.30406v1 Announce Type: new Abstract: Modern large language models (LLMs) rely on reinforcement learning during post-training to push specific capabilities, yet integrating multiple capabilities into one model remains hard. Existing methods, such as Off-Policy Finetune and Mix-RL, are either inefficient or lose performance. In this work, we propose Multi-teacher On-Policy Distillation (MOPD), a post-training paradigm for combining the capabilities of multiple domain RL teachers: we fir

Read source article
arXiv cs.CL (NLP)Research

Field Order Should Not Matter: Permutation-Invariant Embedding Model Fine-Tuning for Structured Metadata Retrieval

arXiv:2606.30473v1 Announce Type: new Abstract: We study retrieval over catalogs of structured metadata, where each record is a small schema whose fields answer different kinds of query. Embedding a record with a text encoder first serializes its fields into a string, which forces a choice of field order. We show this choice, usually treated as an implementation detail, silently controls retrieval quality once the encoder is fine-tuned. A standard fine-tune loses 7.4 nDCG@10 points when the inde

Read source article
arXiv cs.CL (NLP)Research

SIMAX: A Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation

arXiv:2606.30491v1 Announce Type: new Abstract: Background. The widespread deployment of ambient digital scribes is driving large-scale capture of clinician-patient dialogues. Human coding of clinical communication data remains costly, inconsistent, and difficult to scale, motivating AI-driven communication coding systems. However, evaluating these systems requires real-world dialogues and human-coded labels, both hard to obtain at scale. Methods. We developed SIMAX (Scalable and Interpretable F

Read source article
arXiv cs.CL (NLP)Research

TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech

arXiv:2606.30543v1 Announce Type: new Abstract: With the proliferation of speech AI agents, understanding emotional entrainment in conversational interaction has become increasingly important. Emotional entrainment is shaped by social relationships and conversational context, influencing affective coordination over time. We introduce DyadEE, a dataset for emotional entrainment detection in dyadic speech interactions, containing both emotionally entrained conversations and synthetic interactions

Read source article
arXiv cs.CL (NLP)Research

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task?

arXiv:2606.30556v1 Announce Type: new Abstract: Traditional automatic evaluation methods have been shown to be unsuitable for modern Chinese poetry because of the distinct nature of this literary genre. Human evaluation remains reliable, but is expensive and not applicable to large-scale data. In this paper, we propose Poller (Poetry LLM Evaluator), a novel method leveraging large language models (LLMs) to evaluate the poetry understanding task. Specifically, our method requires LLMs to play the

Read source article
arXiv cs.CL (NLP)Research

Morphing into Hybrid Attention Models

arXiv:2606.30562v1 Announce Type: new Abstract: Hybrid attention models improve long-context efficiency by retaining only a subset of full-attention layers and replacing the remaining layers with linear attention. However, the effectiveness of Transformer-to-hybrid conversion critically depends on which layers preserve full attention. Existing hybrid layer selection methods typically rely on heuristic strategies such as fixed placement patterns or layerwise scoring, implicitly treating layer imp

Read source article
arXiv cs.CL (NLP)Research

Uncertainty-Aware Generation and Decision-Making Under Ambiguity

arXiv:2606.30578v1 Announce Type: new Abstract: With rapidly improving capabilities, Large Language Models (LLMs) are increasingly used in many complex real-world tasks. Beyond requiring in-depth knowledge and reasoning skills, many of these tasks exhibit a high degree of subjectivity and require that the outputs of the model can be trusted. While a lot of progress has been made to train better models, decision-making algorithms have received less attention. In this work, we present and evaluate

Read source article
arXiv cs.CL (NLP)Research

$M^3 QuestionIng$: Multi-modal Multi-span Medical Question Answering

arXiv:2606.28329v1 Announce Type: cross Abstract: The growing adoption of AI in healthcare, particularly in preventive care, highlights the critical need for accessibility and precision in Medical Question Answering (MedQA). In recent years, significant efforts have been made to develop multi-span medical question-answering systems, where the answer to a query may span multiple sections or paragraphs of a source document. However, existing systems fall short of aligning with real-world scenarios

Read source article
arXiv cs.CL (NLP)Research

LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution

arXiv:2606.28335v1 Announce Type: cross Abstract: We argue, with systematic empirical evidence, that a large language model's political ideology is not a fixed point, but a conditional distribution $\mathbb{P}($position$\mid$context$)$ over a real political space. We evaluate nine current LLMs using a unified measurement framework anchored by VAA-CHES projection models, which map responses onto three validated dimensions (lrgen, lrecon, galtan) across six contextual axes. Our findings reveal hig

Read source article
arXiv cs.CL (NLP)Research

PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation

arXiv:2606.28344v1 Announce Type: cross Abstract: Augmenting large language models (LLMs) with retrieved web text has become a dominant paradigm, yet the web is not natively textual: existing systems depend on complex parsing pipelines that linearize HTML and discard layout, visual structure, and formatting. We introduce PixelRAG, a new retrieval-augmented method that represents websites in their native visual form and performs retrieval and reading entirely in pixel space, enabling an end-to-en

Read source article
arXiv cs.CL (NLP)Research

Auditing LLM-Governed Social Robots with Culture-Specific Moral Gradients

arXiv:2606.28345v1 Announce Type: cross Abstract: LLM-governed social robots increasingly decide who receives real-world assistance first. As prioritization norms vary across cultures by age, status, and group size, failure to calibrate pluralistically can scale into unequal access. Yet LLM moral audits remain English-centered, rarely test embodied contexts, leaving pluralistic calibration as an urgent diagnostic gap amid intensifying LLM-robot deployment. We introduce a gradient-based audit fra

Read source article
arXiv cs.CL (NLP)Research

Sifei at SemEval-2026 Task 8: Hybrid Retrieval and Query Rewriting for Multi-Turn RAG

arXiv:2606.28352v1 Announce Type: cross Abstract: Multi-turn retrieval-augmented generation (RAG) is challenging due to evolving user intent, conversational noise, and strict context limits. We propose a training-free hybrid retrieval pipeline for SemEval-2026 Task 8 that combines dense and sparse retrieval with controlled query rewriting and cross-encoder reranking. On the official test set of Task A, our system achieves 0.5453 nDCG@5, ranking third among 38 teams and outperforming the stronges

Read source article
arXiv cs.CL (NLP)Research

How Do LLMs Cite? A Mechanistic Interpretation of Attribution in Retrieval-Augmented Generation

arXiv:2606.28358v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) aims to enhance the trustworthiness of Large Language Models (LLMs) by grounding their outputs in external documents, often using inline citations for verifiability. However, the faithfulness of these citations -- whether the model genuinely uses a source to generate an answer -- remains a critical, unverified assumption. This paper offers the first mechanistic account of how a large language model decides whe

Read source article
arXiv cs.CL (NLP)Research

LUMEN: Cost-Transparent Multi-Agent Pipeline for Automated Systematic Review and Meta-Analysis

arXiv:2606.28362v1 Announce Type: cross Abstract: Systematic reviews and meta-analyses (SR/MA) remain the gold standard for evidence synthesis, yet completing one typically requires 67 weeks and substantial expert effort. Recent large language model (LLM) systems have demonstrated strong performance on individual SR phases - screening (otto-SR: 96.7% sensitivity), extraction (Gartlehner et al.: 91.0% accuracy), and search (TrialMind: 0.83 recall) - but no study has reported what it actually cost

Read source article
arXiv cs.CL (NLP)Research

LEDGER: Scaling Agentic Document Editing with Dependency-aware Graph Retrieval

arXiv:2606.28379v1 Announce Type: cross Abstract: We introduce LEDGER to tackle the novel context engineering challenge of agentic document editing, where localized edits to long, structured documents must be applied efficiently without breaking cross-references or semantic consistency. LEDGER constructs a lightweight dependency graph that explicitly models document structure, including hierarchical organization, explicit references, implicit dependencies, and semantic relationships. For each ed

Read source article
arXiv cs.CL (NLP)Research

LoRA-Tuned Large Language Models for Dementia Detection via Multi-View Speech-Derived Features

arXiv:2606.28445v1 Announce Type: cross Abstract: Early detection of dementia enables timely intervention, and reflecting cognitive impairment, spontaneous speech offers a non-invasive screening modality. Conventional approaches often focus on a single representational dimension -- such as acoustic descriptors, pause modeling, automatic speech recognition (ASR) transcripts, or multimodal fusion -- limiting integrative reasoning across heterogeneous cognitive symptoms. We propose a low-rank adapt

Read source article
arXiv cs.CL (NLP)Research

GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

arXiv:2606.28514v1 Announce Type: cross Abstract: Multimodal models are increasingly deployed to solve tasks collaboratively with humans or other artificial agents. Existing benchmarks show that these models possess many of the required component capabilities, but the conditions that coincide in collaboration, including time pressure, information asymmetry, and imperfect communication, are usually studied in isolation. We introduce GPTNT, a benchmark built on the cooperative video game Keep Talk

Read source article
Hacker News: Show HN

Show HN: We made an Audio ML sharing platform

Hey everyone we recently built a platform for people to share their audio ml creations and we are hoping to create it so people in the community can find demos of their models and eventually monetize their trained models. We are a team of small creators who are trying to improve the space for open source ML so the future is not dominated by closed source giants like ElevenLabs or chatgpt. if you have any feedback please feel free to share. Comments URL: https://news.ycombinator.com/item?id=48728

Read source article
Product HuntTools

Cursor for iOS

<p> Build with coding agents from anywhere </p> <p> <a href="https://www.producthunt.com/products/cursor?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="http://www.producthunt.com/r/p/1184233?app_id=339">Link</a> </p>

Read source article
Product Hunt — The best new products, every day

Cursor for iOS

<p> Build with coding agents from anywhere </p> <p> <a href="https://www.producthunt.com/products/cursor-for-ios?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1184233?app_id=339">Link</a> </p>

Read source article
Product HuntTools

WorkBuddy

<p> Produce sharpened results faster with a team of AI experts </p> <p> <a href="https://www.producthunt.com/products/workbuddy-2?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1184211?app_id=339">Link</a> </p>

Read source article
Product HuntTools

Flowly

<p> A personal AI agent that runs on your desktop and iPhone </p> <p> <a href="https://www.producthunt.com/products/flowly-6?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1184210?app_id=339">Link</a> </p>

Read source article
Hacker News AILLMs

Why AI is like a (Clever Hans) Horse [video]

Article URL: https://www.youtube.com/watch?v=0GQ2RP-25gM Comments URL: https://news.ycombinator.com/item?id=48727856 Points: 3 # Comments: 0

Read source article
Hacker News AILLMs

Moonshot AI (kimi) launches a credit card

Article URL: https://www.kimi.com/aicard Comments URL: https://news.ycombinator.com/item?id=48727792 Points: 4 # Comments: 1

Read source article
Hacker News AILLMs

AI money is going to swamp the midterms this year

Article URL: https://www.ft.com/content/8f872761-76d9-43c7-bbcd-d313c4732e81 Comments URL: https://news.ycombinator.com/item?id=48727708 Points: 7 # Comments: 0

Read source article
Hacker News LLMLLMs

Show HN: TinyAgents – a Rust based recursive LLM harness

Hey Guys, I would like to showcase Tiny Agents, which is an entirely Rust-based RLM. RLMs are a recent and fantastic new innovation on how LLMs can be reimagined by designing the LLM agent itself to define how it will orchestrate or create its own sub-agent. I also could not find any good LLM harness workflow in Rust, so I built this one, taking massive inspiration from LangGraph and LangChain but also incorporating a Rust-based transpiler to build on-demand LLM graphs all in Rust following the

Read source article
Hacker News AILLMs

Why Won't Europe Build AI Data Centers in Iceland?

Article URL: https://mrkt30.com/why-wont-europe-build-ai-data-centers-in-iceland/ Comments URL: https://news.ycombinator.com/item?id=48727538 Points: 29 # Comments: 25

Read source article
Hacker News: Show HN

Show HN: Agentic Orchestrator, a TUI for long-running coding agents

Hello Folks! Agentic Orchestrator is a terminal tool that takes complex feature requests and builds them by orchestrating coding agents through a series of phases that emulate a full-fledged engineering flow: requirements clarification, research, design, multi-phase planning, implementation, and review. It is a single pane of glass for all your features and exposes post-publish utilities such as resolving merge conflicts and responding to review comments. The key design choice is that this is de

Read source article