AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

32272 stories from 30+ sources, refreshed continuously.

The DecoderBusiness

OpenAI admits it "didn't get everything quite right" with ChatGPT Work launch and scrambles to fix UX and costs

Following the launch of ChatGPT Work and GPT-5.6 Sol, OpenAI has acknowledged significant issues: excessive compute usage, a confusing transition to the desktop interface for chats and projects, an unclear distinction between Codex and ChatGPT Work, and regressions in existing workflows. In some cases, GPT-5.6 Sol reportedly deleted data on its own that the user had not authorized. The article OpenAI admits it "didn't get everything quite right" with ChatGPT Work launch and scrambles to fix UX a

Read source article
OpenAI admits it "didn't get everything quite right" with ChatGPT Work launch and scrambles to fix UX and costs
Hacker News LLMLLMs

Inside HackerRank's LLM-based Hiring Agent

Article URL: https://blog.grandimam.com/posts/how-hacker-rank-scores-engineers/ Comments URL: https://news.ycombinator.com/item?id=48869737 Points: 1 # Comments: 0

Read source article
Machine Learning

Predicting human preference for generated image pairs using HPSv3 [P]

<!-- SC_OFF --><div class="md"><p>Hey! I'm looking for ways to predict human preference for a project I'm building. (imagebench.ai)</p> <p>I've tryed <strong>HPSv3</strong>, <a href="https://github.com/MizzenAI/HPSv3">https://github.com/MizzenAI/HPSv3</a> and made post about it here:</p> <p><a href="https://imagebench.ai/blog/does-the-score-match-your-eye">https://imagebench.ai/blog/does-the-score-match-your-eye</a></p> <p>It looks ok, but have many limitation as you can see in my post. </p> <p>

Read source article
OpenClaw Commits

feat: sidebar update card (web + macOS) with app-first mac update flo…

<pre style='white-space:pre-wrap;width:81ex'>feat: sidebar update card (web + macOS) with app-first mac update flow and Sparkle beta track (#104171) * feat: sidebar update card (web + macOS) with app-first mac update flow and Sparkle beta track Squashed from claude/update-notification-display-c6cfb9 after semantic merge with #104178 (channel-aware CLI installs). See PR #104171 body for details. * chore(i18n): resync generated inventories after rebase * chore(i18n): resync locale metadata after r

Read source article
Dev.to

What edge cases would you test for stablecoin checkout webhooks?

<p>I'm building ChainPay, a stablecoin checkout for WooCommerce, SaaS products, Telegram sellers, and agent workflows.</p> <p>The wallet UI is only one part of the problem. The part I keep coming back to is payment state: webhook retries, late payments, partial payments, duplicated events, and what happens when the customer closes the checkout tab.</p> <p>Here is the checklist I am using right now:</p> <ul> <li>exact payment within the expiry window</li> <li>payment after expiry</li> <li>partial

Read source article
Hacker News AILLMs

Competitive Programming in the Era of AI

Article URL: https://www.vibhaas.net/posts/Competitive-Programming-in-the-era-of-AI/ Comments URL: https://news.ycombinator.com/item?id=48869623 Points: 1 # Comments: 0

Read source article
Dev.to

I kept leaving my terminal.

<h3> ...so I built <code>shortcuts</code>. </h3> <p>Every developer knows keyboard shortcuts are worth learning.<br> The problem is remembering them.</p> <p><code>shortcuts</code> does that for you.</p> <p>You forget how to split a pane in Windows Terminal, search through tmux scrollback, or jump to the end of a command. Instead of staying in your terminal, you open a browser, search the web, skim documentation, on a bad day ask some AI chatbot.</p> <p>The interruption often costs more time than

Read source article
Hacker News AILLMs

How AI is rewiring childhood

Article URL: https://www.economist.com/leaders/2025/12/04/how-ai-is-rewiring-childhood Comments URL: https://news.ycombinator.com/item?id=48869574 Points: 2 # Comments: 0

Read source article
OpenClaw Commits

fix(agents): keep exact tool allowlists on owning factories (#104213)

<pre style='white-space:pre-wrap;width:81ex'>fix(agents): keep exact tool allowlists on owning factories (#104213) * refactor(agents): centralize core tool factory descriptors --------- Co-authored-by: Ayaan Zaidi <hi@obviy.us></pre>

Read source article
OpenClaw Commits

feat(ui): full-page New session screen with gateway folder browser (#…

<pre style='white-space:pre-wrap;width:81ex'>feat(ui): full-page New session screen with gateway folder browser (#104238) * feat(gateway): add admin-only fs.listDir host directory listing * feat(ui): replace new-session dialog with full-page /new screen and folder browser * fix(ui): drop unnecessary template literal in new-session page * refactor(ui): rename request token locals for review-bundle hygiene * fix(ui): preserve typed draft on agent hydration and clear stale folder listings * fix(ui)

Read source article
The DecoderBusiness

Apple sues OpenAI for allegedly running a "coordinated campaign" to steal trade secrets through poached employees

Apple is suing OpenAI over systematic employee poaching and the alleged theft of trade secrets tied to unreleased products. According to the complaint, more than 400 ex-Apple employees now work at OpenAI, including former iPhone design chief Tang Tan. The lawsuit hits OpenAI right as it's building out its own hardware division, with its first product not expected to ship until 2027 at the earliest. The article Apple sues OpenAI for allegedly running a "coordinated campaign" to steal trade secret

Read source article
Hacker News Ask

Ask HN: How are you controlling Token Costs?

I have been using LLMs & Coding Agent since early 2024. A large problem with Coding Agents & LLMs in general is context compression. To give you some numbers, when I analysed my own sessions across Claude Code, Codex & Sakana, I found that most of my agents spent >90% of time re reading context and upon further investigation into the markdowns it was reading, I have a hand-wavy estimate of at least ~20% of this being useless to the task at hand. When digging a bit more into this problem, I reali

Read source article
The Guardian AIBusiness

Meta ditches Muse Image AI feature because it ‘misses the mark’ on users’ privacy

<p>Meta was criticised for feature launched on Tuesday that automatically lets users generate images using content from public Instagram accounts</p><p>Meta has said ⁠it is discontinuing an AI feature launched this week that allowed users to generate images using public Instagram ⁠accounts, after drawing widespread ⁠criticism over ​privacy concerns, including from a Hollywood union.</p><p>“Our intent was to provide a useful creative tool and to give people control ⁠over whether their public cont

Read source article
Meta ditches Muse Image AI feature because it ‘misses the mark’ on users’ privacy
Hacker News: Show HN

Show HN: I built a YouTube for generative videos in 50 prompts

So I decided yesterday to do a full rewrite of the media gallery, turning it into a generative media catalog with full-suite social interactions, search, and personalized recommendations built on top of the samsar-js library. It took around 50 prompts in 5 sessions to get this done across the client and API, built with Codex GPT-5.6 Sol, Luna, and then Sol again (rate limit reset, thx Mr. Sama). I did not use screenshots for the UI, It was entirely hand-prompted end-to-end. The project is fully

Read source article
arXiv cs.AIResearch

AI-integrated models for assessing agricultural resilience

arXiv:2607.07759v1 Announce Type: new Abstract: Agricultural supply chains are vulnerable to disruptions through linked biophysical and economic systems. We develop an AI-powered tool that integrates economic models (GTAP) with biophysical models (APSIM) to analyze supply chain shocks, enabling policymakers and market participants to assess cross-disciplinary impacts through queries and responses written in natural language.

Read source article
arXiv cs.AIResearch

Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

arXiv:2607.07761v1 Announce Type: new Abstract: Large language models (LLMs) have emerged as important tools in healthcare, showing growing potential for clinical reasoning and patient care. This survey examines recent progress in medical LLMs, focusing on reasoning applications and requirements. We present a dual-view approach that connects clinical practice with computational methods. On the clinical side, we establish a five-level competency scheme following Miller's Pyramid, progressing from

Read source article
arXiv cs.AIResearch

Alignment Plausibility: A New Standard for Assuring AI in Healthcare

arXiv:2607.07766v1 Announce Type: new Abstract: Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires. Developers' safety responses have been largely reactive, addressing the most visible and acute harms while subtler, longer-term patterns of risk (e.g., dependency, boundary e

Read source article
arXiv cs.AIResearch

Idiobionics: The Unification of Privacy and Intelligent Robotic Prostheses

arXiv:2607.07775v1 Announce Type: new Abstract: The human body is at the center of a growing family of technologies designed to tightly and persistently couple biological and digital systems. Robotic prostheses are a representative example of this tight coupling. Also referred to as bionic limbs, robotic prostheses are devices that support people who have lost limbs in pursuing daily life activities such as walking and grasping objects. Bionic limbs are now perceptive and responsive owing to the

Read source article
arXiv cs.AIResearch

VectorizationLLM: Smart Vectorization Based AI Assistant

arXiv:2607.07846v1 Announce Type: new Abstract: VectorizationLLM is a specialized Large Language Model based on Google open-weight LLMs. The model is designed to assist students to learn smart vectorization, time/wave vector analysis, piecewise functions, Fourier analysis, and differential equations in MATLAB. The course application is CTEC 247: Applied Computational Analysis II by the Department of Electrical & Computer Engineering Technology at New York Institute of Technology Old Westbury. Th

Read source article
arXiv cs.AIResearch

A Graph Neural Network Model for Real-Time Gesture Recognition Based on sEMG Signals

arXiv:2607.07850v1 Announce Type: new Abstract: For seemless control of advanced hand prostheses and augmented reality, accurate and immediate hand gestures recognition is essential. Surface electromyography (sEMG) signals obtained from the forearm are commonly employed for this purpose. In this paper, we present a novel approach for sEMG representation that utilizes graph networks which contain information about muscle activation patterns in the forearm. Based on these graph networks, we have d

Read source article
arXiv cs.AIResearch

Agentic Neural Architecture Search

arXiv:2607.07984v1 Announce Type: new Abstract: Neural architecture search (NAS) methods have grown increasingly efficient, yet they remain bounded by manually engineered search spaces that require substantial domain expertise and must be rebuilt for every new task. Large language models (LLMs) can generate architectures in an open-ended space, but how to optimally divide the labor between LLM-driven design and NAS-driven search remains unexplored. We propose a mechanism that bridges these two p

Read source article
arXiv cs.AIResearch

Concretized Proposition Prompting Resolves Composition-Knowledge Dichotomy in Large Language Models

arXiv:2607.08018v1 Announce Type: new Abstract: LLMs often struggle to balance compositionality with knowledgeability, a challenge we define as Composition-Knowledge Dichotomy. To address this, we propose Concretized Proposition Prompting (CPP), a framework that explicitly concretizes propositions relevant to questions. The results demonstrate that CPP significantly enhances reasoning performance, particularly in medical benchmarks where precise knowledge is paramount, while being competitive on

Read source article
arXiv cs.AIResearch

A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis

arXiv:2607.08038v1 Announce Type: new Abstract: Diagnostic error is a major threat to patient safety, yet current large language model (LLM) systems often treat diagnosis as a one-shot prediction task, lacking safeguards against missed high-risk alternatives or rigorous verification of their reasoning. Here, we present AegisDx, a safety-oriented framework for hypothetico-deductive clinical reasoning. AegisDx coordinates specialized LLM components through role-specific contracts, structured inter

Read source article
arXiv cs.AIResearch

When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals

arXiv:2607.08065v1 Announce Type: new Abstract: LLM-as-judge (Zheng et al., 2023) is increasingly the default for evaluating AI systems in enterprise pipelines, often scaled to ensembles (Verga et al., 2024) or "mixture-of-experts" (Shazeer et al., 2017) panels of judges. These systems share a key assumption: that consistency -- agreement among judges, or among a model's own samples -- indicates correctness. We show this assumption is unreliable. Agreement is not accuracy: a model can agree with

Read source article
arXiv cs.AIResearch

PARA-PV: Physics-Aware Retrieval-Augmented PV Prediction Based on Frozen Foundation Model and Distribution Shift Correction

arXiv:2607.08079v1 Announce Type: new Abstract: Accurate photovoltaic (PV) power forecasting is essential for reliable grid dispatch and renewable energy integration, yet it remains challenging because PV generation is jointly shaped by weather variability, day-night transitions, regime-dependent dynamics, and strict physical constraints. We propose PARA-PV, a Physics-Aware Retrieval-Augmented framework that embeds physical knowledge throughout the forecasting process. The framework first encode

Read source article
arXiv cs.AIResearch

Answer Set Programming Energised! End-to-End Neurosymbolic Reasoning and Learning with ASP and Energy Based Models

arXiv:2607.08136v1 Announce Type: new Abstract: We present a general neurosymbolic reasoning and learning methodology based on a modular integration of answer set programming with an energy based model substrate. Key contributions are: (1) supporting joint optimisation in the continuous latent space through explicit ASP-based declarative semantics fully incorporating background knowledge, constraints, non-monotonic inference; and (2) advancing recent works at the interface of answer sets, probab

Read source article
arXiv cs.AIResearch

ASMR: Agentic Schema Generation for Ship Maintenance Report Writing

arXiv:2607.08177v1 Announce Type: new Abstract: In this paper, we study the automatic schema generation problem: given a collection of historical ship maintenance and operational reports across multiple form categories, automatically discover compact and informative schemas that capture the essential information requirements of each report type. To address this challenge, we propose ASMR, a modular agentic framework consisting of two specialized agents. A Field Generation Agent extracts semantic

Read source article
arXiv cs.AIResearch

Compete Then Collaborate: Frontier AI Teachers Build a Verifiable Curriculum to Improve a Coding Student Beyond Imitation

arXiv:2607.08255v1 Announce Type: new Abstract: Large language models increasingly serve as teachers generating training data for smaller students. Prior multi-teacher knowledge distillation methods merge outputs without determining which frontier model teaches best, often relying on an LLM judge biased toward its own outputs. We introduce a compete-then-collaborate framework where four frontier AI teachers (Claude, Codex-GPT, Grok, Gemini) are ranked head-to-head by an execution-based judge (un

Read source article
arXiv cs.AIResearch

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters

arXiv:2607.08257v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters. We introduce $\textbf{MentalHospital}$, a virtual evaluation environment for LLM-based psychiatric clinical encounters. MentalHospital instantiates the Subjective Interviewing, Objective Examination, Diagnostic Assessment, an

Read source article
arXiv cs.AIResearch

Understanding Axes of Difficulty For Long Context Tasks Via PredicateLongBench

arXiv:2607.08284v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated rapidly improving long-context capabilities, prompting a wave of benchmarks designed to evaluate them. However, existing long-context evaluations - from Needle-in-a-Haystack (NIAH) tests to more recent multi-hop reasoning and summarization tasks - predominantly measure average-case performance, and many are either saturated or lack robustness. Notably absent is a systematic way to probe how models perf

Read source article
arXiv cs.AIResearch

Psychological Competence as a Missing Dimension in AI Evaluation

arXiv:2607.08285v1 Announce Type: new Abstract: Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy compliance. These measures remain essential, but they are not sufficient for systems that interact directly with users through natural language. Human-facing AI systems are increasingly used as advisors, coaches, tutors, and companions. In these roles, their responses can shape how users reason, interpret emotions

Read source article
arXiv cs.AIResearch

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

arXiv:2607.08317v1 Announce Type: new Abstract: Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a dog with five legs. These examples suggest that existing benchmarks may under-measure persistent blind spots in current systems. We introduce $\texttt{blind-spots-bench}$, a benchmark designed to expose such blind spots through tasks that appear simple for humans but rem

Read source article