AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

34130 stories from 30+ sources, refreshed continuously.

arXiv cs.LGResearch

Towards Understanding The Calibration Benefits of Sharpness-Aware Minimization

arXiv:2505.23866v2 Announce Type: replace Abstract: Deep neural networks have been increasingly used in safety-critical applications such as medical diagnosis and autonomous driving. However, many studies suggest that they are prone to being poorly calibrated and have a propensity for overconfidence, which may have disastrous consequences. In this paper, unlike standard training such as stochastic gradient descent, we show that the recently proposed sharpness-aware minimization (SAM) counteracts

Read source article
arXiv cs.LGResearch

PERRY: Policy Evaluation with Confidence Intervals using Auxiliary Data

arXiv:2507.20068v3 Announce Type: replace Abstract: Off-policy evaluation (OPE) methods estimate the value of a new reinforcement learning (RL) policy prior to deployment. Recent advances have shown that leveraging auxiliary datasets, such as those synthesized by generative models, can improve the accuracy of OPE methods. Unfortunately, such auxiliary datasets may also be biased, and existing methods for using data augmentation within OPE lack principled uncertainty quantification. In high stake

Read source article
arXiv cs.LGResearch

Shapley-Inspired Feature Weighting in $k$-means with No Additional Hyperparameters

arXiv:2508.07952v2 Announce Type: replace Abstract: Clustering algorithms often assume all features contribute equally to the data structure, an assumption that usually fails in high-dimensional or noisy settings. Feature weighting methods can address this, but most require additional parameter tuning. We propose SHARK (Shapley Reweighted $k$-means), a feature-weighted clustering algorithm motivated by the use of Shapley values from cooperative game theory to quantify feature relevance, which re

Read source article
arXiv cs.LGResearch

Discovering New Theorems via LLMs with In-Context Proof Learning in Lean

arXiv:2509.14274v3 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated significant promise in formal theorem proving. In this study, we investigate the ability of LLMs to discover novel theorems and produce verified proofs. We propose a pipeline called Conjecturing-Proving Loop (CPL), which iteratively generates mathematical conjectures and attempts to prove them in Lean 4. A key feature of CPL is that each iteration conditions the LLM on previously generated theorems

Read source article
arXiv cs.LGResearch

How Does the Pretraining Distribution Shape In-Context Learning? A Fundamental Trade-Off

arXiv:2510.01163v2 Announce Type: replace Abstract: The factors driving the performance of in-context learning (ICL) in large language models (LLMs) remain poorly understood despite ICL's surprising effectiveness, enabling models to adapt to new tasks from only a handful of examples. To clarify and improve these capabilities, we characterize how the statistical properties of the pretraining distribution (e.g., tail behavior, coverage) shape ICL. We develop a theoretical framework that encompasse

Read source article
arXiv cs.LGResearch

Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning

arXiv:2510.04773v2 Announce Type: replace Abstract: As Large Language Models (LLMs) demonstrate remarkable capabilities learned from vast corpora, concerns regarding data privacy and safety are receiving increasing attention. LLM unlearning, which aims to remove the influence of specific data while preserving overall model utility, is becoming an important research area. One of the mainstream unlearning classes is optimization-based methods, which achieve forgetting directly through fine-tuning,

Read source article
arXiv cs.LGResearch

Consistent Zero-Shot Imitation with Contrastive Goal Inference

arXiv:2510.17059v2 Announce Type: replace Abstract: Zero-shot imitation learning requires an agent to reproduce expert behavior from a single demonstration without additional environment interaction or gradient updates at test time. We introduce Contrastive Inverse Reinforcement Learning (CIRL), a self-supervised framework for pre-training zero-shot imitation agents. Our methods rests on a key observation that many useful tasks can be summarized by a single goal state. We can thus convert the mu

Read source article
arXiv cs.LGResearch

Weight Space Representation Learning via Neural Field Adaptation

arXiv:2512.01759v3 Announce Type: replace Abstract: We investigate the potential of weights to serve as effective representations, focusing on neural fields. Our key insight is that constraining the optimization space through a pre-trained base model and low-rank adaptation (LoRA) can induce structure in weight space. Across reconstruction, generation, and analysis tasks on 2D and 3D data, we find that multiplicative LoRA weights achieve high representation quality while exhibiting distinctivene

Read source article
arXiv cs.LGResearch

Auto-exploration for online reinforcement learning

arXiv:2512.06244v2 Announce Type: replace Abstract: The exploration-exploitation dilemma in reinforcement learning (RL) is a fundamental challenge to efficient RL algorithms. Existing algorithms for finite state and action discounted RL problems address this by assuming sufficient exploration over both state and action spaces. However, this yields non-implementable algorithms and sub-optimal performance. To resolve these limitations, we introduce a new class of methods with auto-exploration, or

Read source article
arXiv cs.LGResearch

MINIF2F-DAFNY: LLM-Guided Mathematical Theorem Proving via Auto-Active Verification

arXiv:2512.10187v3 Announce Type: replace Abstract: LLMs excel at reasoning, but validating their steps remains challenging. Formal verification offers a solution through mechanically checkable proofs. Interactive theorem provers (ITPs) dominate mathematical reasoning but require detailed low-level proof steps, while auto-active verifiers offer automation but focus on software verification. Recent work has begun bridging this divide by evaluating LLMs for software verification in ITPs, but the c

Read source article
arXiv cs.LGResearch

Learning with Monotone Adversarial Corruptions

arXiv:2601.02193v2 Announce Type: replace Abstract: We study the extent to which standard machine learning algorithms rely on exchangeability and independence of data by introducing a monotone adversarial corruption model. In this model, an adversary, upon looking at a "clean" i.i.d. dataset, inserts additional "corrupted" points of their choice into the dataset. These added points are constrained to be monotone corruptions, in that they get labeled according to the ground-truth target function.

Read source article
arXiv cs.LGResearch

Data- and Variance-dependent Regret Bounds for Online Tabular MDPs

arXiv:2602.01903v3 Announce Type: replace Abstract: This work studies online episodic tabular Markov decision processes (MDPs) with known transitions and develops best-of-both-worlds algorithms that achieve refined data-dependent regret bounds in the adversarial regime and variance-dependent regret bounds in the stochastic regime. We quantify MDP complexity using a first-order quantity and several new data-dependent measures for the adversarial regime, including a second-order quantity and a pat

Read source article
arXiv cs.LGResearch

A Geometry-Aware Efficient Algorithm for Compositional Entropic Risk Minimization

arXiv:2602.02877v2 Announce Type: replace Abstract: This paper studies optimization for a family of problems termed $\textbf{compositional entropic risk minimization}$, in which each data's loss is formulated as a Log-Expectation-Exponential (Log-E-Exp) function. The Log-E-Exp formulation serves as an abstraction of the Log-Sum-Exponential (LogSumExp) function when the explicit summation inside the logarithm is taken over a gigantic number of items and is therefore expensive to evaluate. While e

Read source article
arXiv cs.LGResearch

Limitations of SGD for Multi-Index Models Beyond Statistical Queries

arXiv:2602.05704v2 Announce Type: replace Abstract: Understanding the limitations of gradient methods, and stochastic gradient descent (SGD) in particular, is a central challenge in learning theory. To that end, a commonly used tool is the Statistical Queries (SQ) framework, which studies performance limits of algorithms based on noisy interaction with the data. However, it is known that the formal connection between the SQ framework and SGD is tenuous: Existing results typically rely on adversa

Read source article
arXiv cs.LGResearch

Two Stages of Folding: Convergent Mechanisms in AI Protein Folding Trunks

arXiv:2602.06020v3 Announce Type: replace Abstract: How do protein structure prediction models fold proteins? We investigate this question through causal interventions on the folding trunks of ESMFold, OpenFold, and Boltz-1. Across all three models, we find a shared two-stage computational structure. In the first stage, early blocks initialize pairwise biochemical signals: features like charge propagate from sequence into pairwise representations through architecture-specific pathways. In the se

Read source article
Product HuntTools

Second Brain for AI v2

<p> AI memory that connects the dots across every tool </p> <p> <a href="https://www.producthunt.com/products/second-brain-cloudflare?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1180472?app_id=339">Link</a> </p>

Read source article
Hacker News AILLMs

Bernie Sanders Wants a U.S. Sovereign Wealth Fund for AI

Article URL: https://www.forbes.com/sites/jamesbroughel/2026/06/22/bernie-sanders-wants-a-us-sovereign-wealth-fund-for-ai/ Comments URL: https://news.ycombinator.com/item?id=48668581 Points: 4 # Comments: 2

Read source article
r/MachineLearningResearch

Xperience-10M Download Help [D]

<!-- SC_OFF --><div class="md"><p>Hi, </p> <p>I <strong>really really</strong> need access to Xperience-10M for a deadline which is very soon. </p> <p><a href="https://huggingface.co/datasets/ropedia-ai/xperience-10m">https://huggingface.co/datasets/ropedia-ai/xperience-10m</a></p> <p>Unfortunately, it looks like the owners have stopped approving people to download the dataset. I filled out the form a few weeks ago but have heard nothing back. Several others have also commented on the HF saying

Read source article
Hacker News Ask

You all think it's normal to sit behind a laptop all day

The moment i saw the first llm, i knew the future of tech is keyboardless, actually i was trying to get there before gpt but failed. The fact that its 2026 and most of the industry doesn't see this is baffling to me. When both hardware and software giants are innovating in the same direction (neuralink, apple vision, voice models...etc), You all think using a keyboard and a mouse all day is normal for primates - but it's completely unnatural! it's killing us slowly, both physically & socially. L

Read source article
Hacker News AILLMs

This One's Not AI

Article URL: https://blog.tacoda.dev/this-ones-not-ai-992c95537790 Comments URL: https://news.ycombinator.com/item?id=48668240 Points: 2 # Comments: 2

Read source article
Hacker News Ask

Ask HN: Where is our profession (programmer) going?

I had been running a small (3 people) software company for about 4 years. Since closing down, I recently hung out at a friend's company to see what they were working on (15 ppl). To preface: I'm a heavy user of Claude (rarely write code by hand), but what I'm seeing in person has been rather shocking to me, and I wanted to calibrate with others. In particular: - the code is not the source of truth anymore; it's ask claude to write, and ask claude to explain - LoC, abstractions, and all those "so

Read source article
OpenAI BlogLLMs

How agents are transforming work

A new OpenAI research paper shows how AI agents are transforming work, enabling longer, more complex tasks and expanding productivity across roles.

Read source article
Product Hunt — The best new products, every day

Paybond CLI

<p> Safe agent spend from the terminal </p> <p> <a href="https://www.producthunt.com/products/paybond?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1180424?app_id=339">Link</a> </p>

Read source article
OpenClaw Commits

fix(cron): preserve enabled-with-defaults failure alert through store…

<pre style='white-space:pre-wrap;width:81ex'>fix(cron): preserve enabled-with-defaults failure alert through store roundtrip (fixes #96589) (AI-assisted) (#96615) Summary: - The PR preserves `failure_alert_disabled === 0` as the enabled-with-defaults failure-alert state and adds focused codec roundtrip tests. - PR surface: Source +2, Tests +54. Total +56 across 2 files. - Reproducibility: yes. At source level, current main encodes `failureAlert: {}` with `failure_alert_disabled = 0`, then decode

Read source article
Hacker News: Show HN

Show HN: Promptctl – Git for your AI prompts

Article URL: https://github.com/naya-ai/promptctl Comments URL: https://news.ycombinator.com/item?id=48667544 Points: 2 # Comments: 0

Read source article
OpenClaw Commits

fix(agent): emit model.usage diagnostic for HTTP ingress traffic (#96…

<pre style='white-space:pre-wrap;width:81ex'>fix(agent): emit model.usage diagnostic for HTTP ingress traffic (#96152) * fix(agents): emit model.usage diagnostic for HTTP ingress traffic * fix(agents): emit model.usage diagnostic for HTTP ingress traffic * fix(agent): add regression tests and refactor ingress model.usage diagnostic emission * fix(agent): resolve oxlint curly and no-useless-fallback-in-spread violations</pre>

Read source article
TechCrunch AIBusiness

Europe is pushing back on Washington’s chip war

As ASML CEO Christophe Fouquet told TechCrunch in May, what China can currently buy are older-generation deep ultraviolet tools — gear first shipped about a decade ago — the same machines the MATCH Act would now put off limits.

Read source article
Hacker News AILLMs

We'll fight the platform war against Big AI

Article URL: https://www.anildash.com/2026/06/23/fight-ai-platform-war/ Comments URL: https://news.ycombinator.com/item?id=48667112 Points: 3 # Comments: 0

Read source article