Tool routing for a local LLM agent: grep to embeddings to GBNF grammar
Article URL: https://eris-system.dev/blog/tool-routing Comments URL: https://news.ycombinator.com/item?id=49624340 Points: 1 # Comments: 0
Article URL: https://eris-system.dev/blog/tool-routing Comments URL: https://news.ycombinator.com/item?id=49624340 Points: 1 # Comments: 0
I spend a good amount of time chasing the LLM for evidence that its answer is actually accurate. To be honest, I can usually get there, but it works great for fresh work and gets messy on months-old work. And it's not just me; a lot of people have raised the same thing. So for the last 6 months I've been trying to fix it. Today I can say it's working for me — it flags me in advance. Obviously not perfect yet, and that's where I need help: I need people to test it and tell me what else needs fixi
Article URL: https://xcancel.com/markchen90/status/2097400166554993041 Comments URL: https://news.ycombinator.com/item?id=49622858 Points: 2 # Comments: 0
Article URL: https://www.awanderingmind.blog/posts/2026-08-08-what-llm-coding-agents-have-taken-from-me.html Comments URL: https://news.ycombinator.com/item?id=49622154 Points: 2 # Comments: 0
<blockquote cite="https://mathstodon.xyz/@tao/117237320796901560"><p>I wrote recently about how the collection of good, fruitful open problems is now being mined in a non-renewable fashion, leading to the potential scenario of these problems becoming scarce. [...]</p> <p>We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may n
<p><a href="https://openai.com/index/navier-stokes-solution/">On the Navier–Stokes Millennium Prize Problem</a> introduces an impressive result from OpenAI, who used an unreleased model to produce a resolution to <a href="https://en.wikipedia.org/wiki/Navier–Stokes_existence_and_smoothness">the Navier–Stokes existence and smoothness problem</a>, one of the seven <a href="https://en.wikipedia.org/wiki/Millennium_Prize_Problems">Millennium Prize Problems</a> that have been subject to a $1,000,000
<p><strong><a href="https://openai.com/index/introducing-chatgpt-images-2-5/">Introducing ChatGPT Images 2.5</a></strong></p> OpenAI's image generation models are apparently used "more than 3 billion images across ChatGPT Images and the GPT‑Image models in the API". This latest release improves their instruction-following ability across multiple turns, responds faster, and "is better at preserving the subjects in your reference photos".</p> <p>There are two new model IDs in the API: <code>gpt-im
Article URL: https://support.microsoft.com/he-il/windows/deployment/updates-lifecycle/windows-update-is-now-carbon-aware Comments URL: https://news.ycombinator.com/item?id=49616971 Points: 1 # Comments: 0
See how an MIT researcher uses GPT-5.6 Sol with Codex to autonomously run quantum computing experiments, analyze results, and calibrate qubits.
Explore how more capable, affordable AI can expand the work people and businesses can accomplish—and make growth more economical.
ChatGPT Images 2.5 helps turn your ideas, sketches, and reference photos into more personalized, polished images that better reflect your ideas.
We’re sharing an AI-generated solution to the Navier–Stokes Millennium Prize Problem, including a writeup and a formal proof in Lean.
Apply now for OpenAI’s $5 million grant program supporting independent research on how generative AI affects teen development, well-being, and safety.
OpenAI is expanding support for journalism with tools, training, and partnerships for students, educators, journalists, and news organizations.
<p><strong>Release:</strong> <a href="https://github.com/simonw/llm/releases/tag/0.35">llm 0.35</a></p> <blockquote> <ul> <li>New OpenAI model: <code>gpt-6-astra</code> for <a href="https://openai.com/index/gpt-6-astra/">GPT-6 Astra</a>.</li> </ul> </blockquote> <p>Tags: <a href="https://simonwillison.net/tags/openai">openai</a>, <a href="https://simonwillison.net/tags/llm">llm</a>, <a href="https://simonwillison.net/tags/gpt-6-astra">gpt-6-astra</a></p>
<blockquote cite="https://openai.com/index/an-alien-mind/#scalable-defense"><p>The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI. [...]</p> <p>We will need powerful, aligned AI for defense; to secure infrastructure, to protect against rogue agents in real time, and to invent entirely new protective measures. This will be a primary focus of OpenAI’s deployment efforts.</p> <p>At the same ti
<p><strong>Tool:</strong> <a href="https://tools.simonwillison.net/video-compressor">Video compressor</a></p> <p>I recorded a short demo video of <a href="https://simonwillison.net/2026/Sep/7/equal-earth/">my Equal Earth</a> animation on my phone and wanted to publish an optimized version of that video (using FFMPEG) on my blog, so I had Claude Fable 5.1 in Claude Code for web <a href="https://claude.ai/code/session_01QHTdJZ4xg6TZfDXCmuvAE9">build me this tool</a> using the WebAssembly build of
<p><strong>Tool:</strong> <a href="https://tools.simonwillison.net/equal-earth">Mercator ↔ Equal Earth</a></p> <p>I got curious about the Equal Earth map projection that was recently <a href="https://www.theguardian.com/world/2026/sep/04/un-vote-world-map-mercator-equal-earth-africa">voted on at the UN</a> so I had GPT-6 Astra (medium) in ChatGPT Work <a href="https://chatgpt.com/share/6a9ee520-c82c-83ea-8111-2f7050c08638">build me</a> this animated transition between Mercator and Equal Earth us
OpenAI, AIRPPU and WAN-IFRA launch an AI program to help Ukrainian news organizations strengthen innovation, resilience, and independent journalism.
<p><strong><a href="https://openai.com/index/research-acceleration-view-inside-openai/">Research acceleration: The view inside OpenAI</a></strong></p> Apparently today is RSI day at OpenAI, for Recursive Self-Improvement - I think it's their new AGI. Both this piece and the new essay <a href="https://openai.com/index/an-alien-mind/">An Alien Mind</a> (by Chief Scientist Jakub Pachocki) talk about it, and this one doesn't even bother to expand the acronym.</p> <p>Included are details on how OpenA
Hey With LLMs having infinite attack resources compared to human researchers, how safe are we? We already know that you don't need top tier frontier models to create an attack vector. You can just loop a mid-sized model until it bears result. Or run a small swarm of agents to achieve the same goal How long until a criminal actor with running exo[1] over 2 mac studios breaches password managers? I understand that the storage of passwords might be encrypted, but they might find vulnerabilities at
This has been bothering me for a while. I feel like I have a decent conceptual grasp of what LLMs are doing with written text. But it seems like they also unlocked a bunch of progress in understanding and generating images, audio, and video. I can’t twist my brain into understanding the connection. Is the boom in generated non-text content also built on LLMs, or is it just correlated with it because a bunch of excitement drove investment into the industry? I’m hoping for an ELI-non-ai-but-cs-maj
Built this small project as a proof of concept to see how fun it could be and how low latency it would be to interact with the local agent in a 3D environment. Any feedback is welcome! Comments URL: https://news.ycombinator.com/item?id=49586853 Points: 1 # Comments: 0
I kept finding LLM pricing comparison posts that were already stale, so I built a scraper that snapshots the full OpenRouter catalog (425 models, 58 providers) every day and diffs it against the previous day. It's been running unattended for about a week now: https://costpertoken.dev One thing that fell out of the data surprised me: on the exact same model, a coding-agent-shaped call (large context in, code out) costs roughly 33x more per call than a bulk-classification-shaped call. I checked th
Article URL: https://twitter.com/RTomMcCoy/status/2094870131939684488 Comments URL: https://news.ycombinator.com/item?id=49586590 Points: 9 # Comments: 0
The CVE-2026-35029 privilege escalation in LiteLLM is a reminder that middleware is not something you can install once and forget. What is your view on using self-hosted LLM gateway? I wrote this post - https://leanroute.dev/blog/self-hosting-an-llm-gateway Would like to hear your views. Comments URL: https://news.ycombinator.com/item?id=49586552 Points: 1 # Comments: 2
Sup HN. I built VernLLM cuz every LLM gateaway I looked at, meant adding a network hop just to get rate limiting, multi provider fallback, circuit breaking and other features. Needing to take a whole seperate service just to deploy, monitor and trust with my API keys is just annoying since my whole tech stack was just TypeScript and Node. VernLLM does the same job within your LLM calls, but its in process instead. Its supports OpenAI-compatible APIs, Anthropic, Gemini and Bedrock providers. I wo
Jakub Pachocki reflects on increasingly capable AI and the challenge of keeping it aligned. He calls for stronger safeguards and international coordination.
Inside OpenAI, coding agents are reshaping AI research. Explore early data on agent usage, experiment velocity, task complexity, and research acceleration.
Article URL: https://github.com/llm-as-a-verifier/llm-as-a-verifier Comments URL: https://news.ycombinator.com/item?id=49583167 Points: 2 # Comments: 0
Shall we start indicating the time as "before llm" and "after llm"? I've just talked with someone who wants to break into the industry and i've wrote something like "i have few collogues who entered through contributing to open source, but it was before LLMs". LLMs seem to be such big thing and their impact is rolling throughout industries that maybe we shall start using ALLM to indicate the new time ;) Comments URL: https://news.ycombinator.com/item?id=49583058 Points: 1 # Comments: 0
Article URL: https://arxiv.org/abs/2502.05202 Comments URL: https://news.ycombinator.com/item?id=49582850 Points: 1 # Comments: 0
<p><strong><a href="https://www.youtube.com/watch?v=bOC3DisEOfg">Introducing GPT-6 Astra for developers</a></strong></p> Blink and you'll miss it, but there's a familiar creature at <a href="https://www.youtube.com/watch?v=bOC3DisEOfg&t=119">1m59s</a>:</p> <blockquote> <p>Across the board, Astra has more attention to detail, better understanding of the user's prompt, and can build more sophisticated outputs. In particular, it excels at building 3D models. I've seen it make incredible renderings
Article URL: https://nightrun.io/ Comments URL: https://news.ycombinator.com/item?id=49581115 Points: 2 # Comments: 0
Article URL: https://www.theopenlake.com/blog/openlake-leads-mlperf-storage-v3-0 Comments URL: https://news.ycombinator.com/item?id=49578727 Points: 35 # Comments: 1
<p><strong>TIL:</strong> <a href="https://til.simonwillison.net/llms/blender-coding-agents-macos">Using Blender with coding agents on macOS</a></p> <p>I've been having fun with Blender in ChatGPT Codex on my Mac recently. Getting it to work with coding agents is really easy: install the full Mac application from <a href="https://www.blender.org">blender.org</a> and run a prompt like this:</p> <blockquote> <p><code>Use the already install /Applications/Blender to render a scene of a pelican ridin
Article URL: https://think-twice.me/is-ai-biased-or-sexist-making-chatgpt-evaluate-itself/ Comments URL: https://news.ycombinator.com/item?id=49576559 Points: 1 # Comments: 0
As part of my job, I have worked across both width and depth. I have worked on the core architecture of deep learning models as well as designed large-scale machine learning systems. I have made this observation that AI Native transformation is all about who lays down a better structure around their core intelligence. Nothing else matters more. Even the context length issue can be handled through robust structure for the time being. I often term it as probabilism filled into determinism. Comment
Article URL: https://catalins.tech/ai-fatigue/ Comments URL: https://news.ycombinator.com/item?id=49576243 Points: 3 # Comments: 2
Article URL: https://www.tomshardware.com/pc-components/cpus/amd-unveils-threadripper-halo-station-an-ai-workstation-packing-96-cores-and-dual-liquid-cooled-mi350p-accelerators-the-most-powerful-workstation-in-the-world-can-run-trillion-parameter-models-says-amd Comments URL: https://news.ycombinator.com/item?id=49576187 Points: 4 # Comments: 0