AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

31240 stories from 30+ sources, refreshed continuously.

The DecoderBusiness

Terence Tao says AI could trigger math's biggest crisis since Gödel

In a new essay, Terence Tao warns that AI could push mathematics into a crisis on par with the foundational upheaval around 1900. What's being tested this time isn't mathematical truth but the values of the field: what counts as a contribution, what gets rewarded, and who did the work. His rule of thumb: a proof that no human can explain should be considered incomplete. The article Terence Tao says AI could trigger math's biggest crisis since Gödel appeared first on The Decoder .

Read source article
Terence Tao says AI could trigger math's biggest crisis since Gödel
Product Hunt — The best new products, every day

fx (by Vercel)

<p> Vercel's tiny, open-source coding agent </p> <p> <a href="https://www.producthunt.com/products/fx-by-vercel?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1227471?app_id=339">Link</a> </p>

Read source article
Moz

Should We Stop Writing With AI?

Should you stop writing with AI as detection, spam, and sameness grow? Learn how expert-led AI content builds trust, audience ownership, and lasting value for years.

Read source article
OpenAI BlogLLMs

Introducing Intelligence Age

Introducing Intelligence Age, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual freedom.

Read source article
OpenAI BlogLLMs

Introducing Intelligence Age

Introducing Intelligence Age, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual freedom.

Read source article
Product Hunt — The best new products, every day

ProtoNote

<p> Share AI-built prototypes, get feedback pinned to the page </p> <p> <a href="https://www.producthunt.com/products/protonote?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1227394?app_id=339">Link</a> </p>

Read source article
Hacker News LLMLLMs

LLM Reasoning Traces Are Not Audit Records

Article URL: https://rye.ai/blog/cot-faithfulness-reasoning-traces-not-audit-logs/ Comments URL: https://news.ycombinator.com/item?id=49371034 Points: 1 # Comments: 1

Read source article
UX Daily - User Experience Daily

Are AI-Generated Synthetic Users Replacing Personas? What UX Designers Need to Know

AI-generated personas sound like a dream: faster insights, lower costs, happier stakeholders. But there’s a catch—if you build for fake users, you risk losing the real ones. The choice isn’t just about speed. It’s about trust, accuracy, and your reputation as a thoughtful, strategic designer. A traditional persona is built on user research. Researchers gain a deep understanding of user needs, motivations, and behaviors and create a one-page summary that gives teams focus and promotes empathy. Co

Read source article
Product Hunt — The best new products, every day

MiniMax Design

<p> Your own agent team for open-ended creation </p> <p> <a href="https://www.producthunt.com/products/minimax?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1227288?app_id=339">Link</a> </p>

Read source article
The Guardian AIBusiness

It used to be our imperfect bodies that made us insecure. With AI, it’s our minds as well

<p>For most of modern history, technology sought to imitate humans. Humans increasingly seek to imitate technology</p><p>Two faces that appeared on my Instagram feed in recent months gave me pause. One belonged to <a href="https://www.youtube.com/watch?v=6WkoyUEIzvI">John Travolta at Cannes</a>. The other to <a href="https://www.instagram.com/reel/DOjUQAzCPy-/">Carla Bruni on a date night</a> with her husband, former French President Nicolas Sarkozy. Both looked recognisably themselves, yet, the

Read source article
It used to be our imperfect bodies that made us insecure. With AI, it’s our minds as well
arXiv cs.CL (NLP)Research

Persona-Guided LLM Agents for Task-Oriented Dialogue

arXiv:2608.18085v1 Announce Type: new Abstract: Prior work has shown that large language models (LLMs) can express diverse personality traits in open-ended text generation. However, it remains unclear whether they can do so in a goal-directed dialogue without compromising task completion, and whether adapting to the user's personality improves the interaction quality. We study these questions in task-oriented dialogue (TOD), where a system helps a user accomplish a goal via multi-turn interactio

Read source article
arXiv cs.CL (NLP)Research

Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining

arXiv:2608.18089v1 Announce Type: new Abstract: Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa. This suggests that the refusal mechanism is present in the residual stream but fails to activate for low-resource inputs. Recovering it normally requires labelled target-language data and retraining, neither of which is available at scale for most African languages. We introduce Latent Space Refusal Anchoring (LSR-A

Read source article
arXiv cs.CL (NLP)Research

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

arXiv:2608.18091v1 Announce Type: new Abstract: As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs -- the tendency to favor one's own outputs -- raises growing concerns about evaluation reliability. However, it has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated. As a result, existing measurements cannot separate genuine self-preference from these confounds. We address this by changing the object

Read source article
arXiv cs.CL (NLP)Research

Abliteration Mitigation via Refusal Aliases

arXiv:2608.18093v1 Announce Type: new Abstract: Abliteration, the removal of refusal capabilities from large language models by projecting weight matrices orthogonal to an extracted refusal direction, has emerged as a prominent safety concern through its ability to bypass post-training alignment using only a small set of contrastive prompts. We find that existing defenses commonly overlook the cause of abliteration; that is, how easily the refusal direction can be extracted. To hinder this proce

Read source article
arXiv cs.CL (NLP)Research

Backdoor Learning in Language Models and Vision-Language Models

arXiv:2608.18095v1 Announce Type: new Abstract: Recent advances in deep learning have significantly enhanced the capabilities of Natural Language Processing (NLP) and Vision-Language Models (VLMs). However, these advancements come with increased vulnerabilities, notably through backdoor attacks that pose severe security threats. This thesis addresses two critical dimensions of Trustworthy AI and Efficient Multimodal Representation Learning: (1) security through analyzing, detecting, and designin

Read source article
arXiv cs.CL (NLP)Research

FrenchNews-7: Benchmarking Cross-Publisher French News Editorial Desk Classification

arXiv:2608.18097v1 Announce Type: new Abstract: We present FrenchNews-7, a cross-publisher France-based French-language news editorial desk classification benchmark combining a large multi-outlet corpus, a URL-derived seven-class taxonomy, and a fine-tuned CamemBERT classifier. Labels are assigned via a hybrid pipeline combining publisher URL slugs with LLM annotation for structurally ambiguous cases, audited through an inter-rater study (2 humans + 2 LLMs; pairwise $\kappa \geq 0.766$, human--h

Read source article
arXiv cs.CL (NLP)Research

Fractional Decay KV-Cache: Ownership-Aware Memory Management for Improved Inference Relevancy in Dialog Systems

arXiv:2608.18098v1 Announce Type: new Abstract: Key-value (KV) caching is essential for efficient autoregressive inference in transformer based dialog systems, yet existing strategies treat all cached entries uniformly or apply coarse eviction heuristics that fail to adapt as dialog topics evolve. We propose Fractional Decay KV-Cache (FD-KVC), a novel algorithm that maintains a dual-channel scoring mechanism for each cached KV pair: a cumulative attention channel that tracks aggregate importance

Read source article
arXiv cs.CL (NLP)Research

Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS)

arXiv:2608.18100v1 Announce Type: new Abstract: AI systems now shape how hundreds of millions of people learn about cultures other than their own. When someone asks one of these systems about the Middle East, they do not receive neutral facts. They receive a representation shaped by the frameworks embedded in training data, and that data is overwhelmingly Western and English-language. This paper asks whether that representation is Orientalist in Said's sense: whether it denies agency to Middle E

Read source article
arXiv cs.CL (NLP)Research

Stability-Aware Feature Design for Robust Watermark Detection in Machine-Generated Text

arXiv:2608.18102v1 Announce Type: new Abstract: The widespread adoption of large language models (LLMs) has intensified the demand for principled methods to distinguish human from machine-generated text. Watermarking provides a promising avenue, yet existing detectors exhibit sharp performance deterioration under multiple paraphrasing and when applied to shorter texts. We introduce Pattern Stability Score (PSS), a novel detection framework that leverages local statistical features and stability

Read source article
arXiv cs.CL (NLP)Research

DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models

arXiv:2608.18103v1 Announce Type: new Abstract: Background: Mechanistic elucidation of traditional Chinese medicine (TCM) compound formulas remains a central challenge in the modernization of TCM. Conventional approaches, including data mining and network pharmacology, are insufficient for achieving deep integration between classical TCM theory and modern scientific research. In addition, direct question-answering using general-purpose artificial intelligence large language models is limited by

Read source article
arXiv cs.CL (NLP)Research

StocksTalk: A Voice-Enabled Conversational Agent for Structured Query Generation over Web Data

arXiv:2608.18105v1 Announce Type: new Abstract: StocksTalk is a voice-enabled conversational system for transforming spoken financial screening requests into executable and validated structured queries over real-world market data. The system combines streaming speech recognition, retrieval-augmented constraint extraction, schema-grounded LLM-based SQL generation, rule-based validation, and human-in-the-loop verification within an interactive dashboard. Unlike traditional template-driven financia

Read source article
arXiv cs.CL (NLP)Research

Different Facets of Verbalised Overconfidence: an Interpretability Study

arXiv:2608.18106v1 Announce Type: new Abstract: Large language models tend to overconfidence, giving assertive answers when the evidence suggests hedging or abstention. Using controlled reasoning scenarios that manipulate logical necessity and possibility, we study this behavior in Qwen3-4B, across three ways to express uncertainty: verbal epistemic markers, abstention, and numeric confidence scores. Our results confirm this tendency toward overconfidence, particularly when the model is prompted

Read source article
arXiv cs.CL (NLP)Research

Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals

arXiv:2608.18107v1 Announce Type: new Abstract: We investigate whether large language models (LLMs) systematically discriminate in candidate evaluations based on applicant name ethnicity and/or institutional prestige and geographic location. Three factorial experiments are reported (4,320 API calls, four LLMs, five professional domains). Study 1 (3x4 design) finds a statistically robust institution-tier gradient of +0.297 points on a 10-point scale (95% bootstrap CI: +0.175 to +0.422), while nam

Read source article
arXiv cs.CL (NLP)Research

Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation

arXiv:2608.18108v1 Announce Type: new Abstract: Large language models are being incorporated into sensitive and important decision-making processes across nearly all fields. While prior work studies model bias around inputs and scenario framing, models can also behave in unexpected and undesirable ways due to context accumulated over their deployment. In this work, we study a medical example in which a model is asked to assign resource-allocation probabilities to two people given brief clinical

Read source article
arXiv cs.CL (NLP)Research

Temporal Multi-Signal Fusion for Token-Level Hallucination Detection

arXiv:2608.18115v1 Announce Type: new Abstract: Token-level hallucination detectors score each token independently from a single signal, and fail exactly when the generating model is confidently wrong. This paper instead treats hallucination as a temporally extended span and detects it by sequence labeling: each token is scored from a 33-dimensional feature stream that fuses text statistics, Natural Language Inference (NLI) entailment, and language model surprisal, with no access to model intern

Read source article
arXiv cs.CL (NLP)Research

You Are What You Prompt: Prompt Quality, Domain Shift, and Uncertainty in Agrifood Vision-Language Models

arXiv:2608.18116v1 Announce Type: new Abstract: Vision-language models enable zero-shot classification through natural language prompts, but performance is sensitive to prompt formulation, especially in specialized domains. Zero-shot Prompt Ensembling (ZPE) addresses this by weighting prompts by discriminative signal, yet its behavior under domain shift remains unexplored. We evaluate ZPE in the agrifood domain using CLIP and SigLIP across four datasets and four prompt pools, spanning in-distrib

Read source article
arXiv cs.CL (NLP)Research

Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models

arXiv:2608.18132v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are typically built through a multi-stage pipeline consisting of cross-modal alignment, supervised fine-tuning (SFT), and preference optimization. This pipeline assumes that adapting an LLM to a new modality requires extensive task-specific supervision. However, pretrained LLMs already possess strong reasoning and instruction-following abilities. As LLMs evolve rapidly, an important question remains: can we

Read source article
arXiv cs.CL (NLP)Research

The Deontic Gap: Large Language Models and the Modal Language of Obligation

arXiv:2608.18144v1 Announce Type: new Abstract: Modal auxiliaries such as must, should, and have to mark necessity and obligation within the contexts of speaker authority and interpersonal stance. We examine whether large language models (LLMs) reproduce contemporary human patterns of deontic modal usage. Across three primary corpora, an external benchmark, two controlled replications, and a naturalistic eleven-model replication, AI-generated text consistently underuses positive deontic modals (

Read source article
arXiv cs.CL (NLP)Research

When Do LLMs Actually Help? Evaluating LLMs as Data Quality Annotators

arXiv:2608.18158v1 Announce Type: new Abstract: LLMs have been increasingly used to catch data quality issues automatically, but we know very little about how consistent these judgments actually are. This study tests an LLM on two e-commerce data quality tasks, entity matching and brand mislabeling, against rule based baselines and human verified ground truth, under both zero-shot and few-shot prompting. On entity matching while using the Abt Buy benchmark (2,194 labeled pairs), a simple rule ba

Read source article
arXiv cs.CL (NLP)Research

Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation

arXiv:2608.18164v1 Announce Type: new Abstract: Safety evaluations of large language models (LLMs) predominantly rely on text-based adversarial prompts, potentially overlooking vulnerabilities arising from alternative input representations. This work examines emoji-augmented prompts as a test case for this gap, evaluating 50 prompts across four open-source LLMs (Mistral 7B, Qwen 2 7B, Gemma 2 9B, Llama 3 8B). Results show substantial variation in robustness: Gemma 2 9B and Mistral 7B exhibit non

Read source article
Vercel Blog

How v0 authenticates to Snowflake without exposing the user's OAuth token

AI-generated applications often need to authenticate to external services on behalf of their users. That creates a problem: generated code shouldn't have access to the user's credentials. We faced that decision when building the v0 Snowflake integration . It lets users connect Snowflake, inspect schemas, query data, and generate applications that run against their warehouses. That generated code has to authenticate to Snowflake, but it is written by a model and runs without human review, and pro

Read source article
Hacker News LLMLLMs

Misconfigured Admin System Prompts Can Invert an LLM's Safety Layer

Article URL: https://medium.com/@aadvait.cr/how-misconfigured-admin-system-prompts-can-invert-every-single-llm-safety-layer-1085a1d79b55 Comments URL: https://news.ycombinator.com/item?id=49369538 Points: 1 # Comments: 0

Read source article
The Guardian AIBusiness

UC Berkeley professor admits to using AI to edit op-ed on students’ math skills

<p>Zvezdelina Stankova says she used AI to ‘help edit’ an article about some of her students being ‘five to eight years’ behind</p><p>A math professor at the University of <a href="https://www.theguardian.com/us-news/california">California</a>, Berkeley, criticizing a “severe” math deficiency among students in an <a href="https://sfstandard.com/opinion/2026/08/15/uc-berkeley-sat-test-blind-admissions-math-scores/">op-ed for the San Francisco Standard</a>, admitted to <a href="https://www.dailyca

Read source article
UC Berkeley professor admits to using AI to edit op-ed on students’ math skills
Artificial Intelligence News — Newsletter on Deep Learning & AI

AI Weekly Issue #524: What AI models are actually coming in the next six months?

If you use AI at work, the tools you rely on could change again before February. OpenAI, Google, Meta, Anthropic, several Chinese labs, and a group of world-model startups are all preparing or rumored to be preparing new releases. Some have announced dates. Others have only appeared in testing reports, leaks, or investor comments. This issue sorts those signals into a practical list: what is likely to ship, what will probably slip, and which releases might actually be worth changing your plans f

Read source article
Apple Machine LearningResearch

Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions

Cross-lingual knowledge transfer is critical for building high-performing multilingual language models for languages with insufficient training data. When target language data is scarce, the knowledge required for many downstream tasks involving scientific reasoning, commonsense inference, and world knowledge must be acquired primarily from the high-resource language, making effective knowledge transfer essential. Existing methods for improving such cross-lingual knowledge transfer require large

Read source article
Apple Machine LearningResearch

Scaling Laws for Mixture Pretraining Under Data Constraints

As language models scale, the amount of data they require grows – yet many target data sources, such as low-resource languages or specialized domains, are inherently limited in size. A common strategy is to mix this scarce but valuable target data with abundant generic data, which presents a fundamental trade-off: too little target data in the mixture underexposes the model to the target domain, while too much target data repeats the same examples excessively, yielding diminishing returns and ev

Read source article
Apple Machine LearningResearch

Progressive Refinement: An Iterative Pseudo-Labeling Approach for Mandarin-English Code-Switching ASR

Code-switching (CS), alternating languages within the same utterance, poses significant challenges for automatic speech recognition (ASR) due to limited CS training data. This paper applies an iterative pseudo-labeling training approach to CS-ASR for the first time, demonstrating its effectiveness in leveraging unlabeled data to improve CS-ASR performance. The approach comprises three phases: pseudo-label generation, two-stage bilingual model training, and iterative improvements. It begins by ge

Read source article