Show HN: Reame – a CPU inference server that gets faster as it runs
Article URL: https://github.com/swellweb/reame Comments URL: https://news.ycombinator.com/item?id=48873417 Points: 1 # Comments: 0
Local-first agent governance: keeping an AI agent contained
Article URL: https://vektorgeist.com/blog Comments URL: https://news.ycombinator.com/item?id=48873414 Points: 2 # Comments: 0
Agentation – Visual UI Annotation for AI Coding Agents
Article URL: https://www.agentation.com/ Comments URL: https://news.ycombinator.com/item?id=48873337 Points: 2 # Comments: 0
Show HN: I Wanted AI Code Review I Could Own. So I Built Codra
Article URL: https://medium.com/@devarshidev/i-wanted-ai-code-review-i-could-actually-own-so-i-built-codra-beea2e3f18fd Comments URL: https://news.ycombinator.com/item?id=48873326 Points: 1 # Comments: 0
Show HN: Make any website agent-ready in one script tag
Article URL: https://www.openhermit.com Comments URL: https://news.ycombinator.com/item?id=48873282 Points: 1 # Comments: 0
Forget typosquatting; slopsquatting is the software supply chain threat created by AI coding tools
Slopsquatting represents an emerging supply chain threat made possible by AI hallucinations. As developers increasingly rely on AI coding assistants, they unknowingly grant cybercriminals access to their software from day one. Understanding what slopsquatting is Slopsquatting is a new type of supply chain attack that uses large language model (LLM) hallucinations to inject malicious code into development workflows. The term combines "AI slop" and "typosquatting," a deceptive practice where attac

My AI Model Tier List for Mid-2026
Article URL: https://taoofmac.com/space/blog/2026/07/11/1500 Comments URL: https://news.ycombinator.com/item?id=48873016 Points: 2 # Comments: 0
OpenAI's head of safety is reportedly leaving as part of company reorganization
The role will be replaced by an executive in charge of both research and safety teams.

An educational lab of AI agent architectures
Article URL: https://github.com/Rudnik-Ilia/Agents-Sandbox Comments URL: https://news.ycombinator.com/item?id=48872922 Points: 1 # Comments: 0
EQK
<p> Mac app with dynamic AI EQ </p> <p> <a href="https://www.producthunt.com/products/eqk-eq-with-ai?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1193751?app_id=339">Link</a> </p>
Show HN: Aether – Run Claude Code, Codex, or OpenCode in devboxes you can watch
Since coding agents like Claude Code and Codex came out, I've been pretty obsessed with them. It's hard not to when you're getting a 20x discount on inference. I've also been frustrated by them a lot. I hated walking around with my laptop lid open while connected to a hotspot on my phone, having to tediously accept commands so I don't lose my root directory, and running out of RAM when running multiple agents that were spinning up dev servers in different worktrees at once. A lot of people would
I made AI agents play diplomacy
Article URL: https://github.com/brendenehlers/diplomacy-ai Comments URL: https://news.ycombinator.com/item?id=48872843 Points: 1 # Comments: 0
Show HN: HoverSource – From pixel to source file in one keystroke
Every UI change starts with the same question: which file, which line, which styles? Just hover and press Alt+C. Your AI gets the answer without digging through thousands of lines of code. Comments URL: https://news.ycombinator.com/item?id=48872839 Points: 1 # Comments: 0
Show HN: Aerial – DIY Orchestration and Agent Messaging
Hey All, I made this project mainly for personal use on startups + to get better at rust. Agents should be able to send each other durable messages without needing Kafka, Redis, Postgres, a cloud account. People also shouldn't be tied down to Claude Code & Codex sub-agents, or Cursor / Antigravity to manage multiple agents simultaneously. It's a lightweight rust binary that allows you to manage agents via an MCP server, Daemon, and a shell service per agent so you can orchestrate agents or pass
VultronRetriever family of models released on HuggingFace![R]
<!-- SC_OFF --><div class="md"><p>Thrilled to announce the VultronRetriever family of models, which were announced during Raise Summit Paris and demonstrated running Q&A and embedding documents on the iPhone, fully offline! 📱</p> <p>Some highlights from the VultronRetriever model family:<br/> 🥇 Each model ranks #1 in its respective class on the MTEB Leaderboard, with VultronRetrieverPrime-8B as the global #1<br/> 📦 VultronRetrieverPrime-8B has up to 16x smaller index storage footprint and 12x
Show HN: A meditative waiting room for Claude/Codex
I kept finding myself getting distracted during those unpredictable stretches while Claude or Codex was working, so I built a little place to spend that time instead. It’s a mix between a meditation zone and a video game; there’s enough to explore to keep you engaged, but no scores, or goals or pressure. Just hang out for a bit, stay present, and head back to your code whenever your agent is ready. It's also a bit of a personal love letter to Wind Waker, Journey, Crash Bandicoot, and many other
AI coding agents read your code perfectly and understand your team not at all
Article URL: https://medium.com/@iamalizaidi110/https-www-youtube-com-watch-v-gdtyotlrndm-a00eb6f014ca Comments URL: https://news.ycombinator.com/item?id=48872770 Points: 7 # Comments: 0
The Silent Epidemic of LLM Technical Debt
Article URL: https://seldon-ai.com/blog/silent-epidemic-llm-tech-debt Comments URL: https://news.ycombinator.com/item?id=48872743 Points: 2 # Comments: 0
Long Context Isn’t Free — I Built a Safe Prompt-Pruning Layer That Makes LLM Systems Work
LLMs don’t fail because they forget—they fail because they remember too much. As conversations grow, prompts accumulate redundant and low-value tokens, driving up cost and latency while silently degrading output quality. This article introduces a deterministic prompt-pruning layer that reduces token usage without breaking dependencies, backed by real benchmarks and production-tested design. The post Long Context Isn’t Free — I Built a Safe Prompt-Pruning Layer That Makes LLM Systems Work appeare
How Does AI Work? [video]
Article URL: https://www.youtube.com/watch?v=YmLp8qe87A0 Comments URL: https://news.ycombinator.com/item?id=48872473 Points: 1 # Comments: 0
Show HN: Agent OS – a local-first harness for reliable software agents
Article URL: https://github.com/earthwalker17/agent-os Comments URL: https://news.ycombinator.com/item?id=48872405 Points: 1 # Comments: 0
Show HN: I used Claude to make some free fun pet-themed browser arcade games
I created https://whatpetshouldiget.com to help my kid understand the responsibilities of pet ownership. It just grew from there and I wanted to make it a destination site where people can learn and have fun. My kid helped me design four pet themed arcade games (done using Fable). Drawing lots of inspiration from existing titles, with some original tweaks, and incorporating of their suggestions along the way. We pulled these together and you can try them out here: https://whatpetshouldiget.com/p
Show HN: verbatimeter - check how grounded your LLM / RAG agent is in real-time
I developed this tool and tried to keep it minimalist, lightweight, portable and easy to use. Hoping some people will find it useful. I think it's especially valuable for RAG agent but could be used more broadly as well. https://github.com/pierreolivierbonin/verbatimeter Comments URL: https://news.ycombinator.com/item?id=48872265 Points: 1 # Comments: 0
OpenAI bets on families as ChatGPT goes deeper into households
ChatGPT is hiring a dedicated product manager to build experiences for families, caregivers, and older adults, according to a job posting.
AI customers are coming around to the idea that small is beautiful
OpenAI and Anthropic have built AI Swiss Army Knives, but the future may be smaller built-for-purpose tools
Show HN: Agent World – an open standard and live market for personal AI agents
Article URL: https://github.com/macrokit/agent-world Comments URL: https://news.ycombinator.com/item?id=48872195 Points: 1 # Comments: 0
Show HN: I turned my quote collection into a walkable 3D library (desktop-only)
Since 2017, each time I finish a book I save a single quote from it. I like revisiting them sometimes, for example at the end of the year or just. This collection is available as plain text on my website. So I thought it may be fun to ask the agents to port this into a 3D library where I can actually walk around and pick up the books, and I think it sort of worked! I'm not sure this is that interesting for people other than me – those are my books! – but I just find it really cool that you can j
Are suppliers ready for new robot safety standards?
Suppliers that are well prepared for standard changes will benefit, while underprepared companies may face market access disruption. The post Are suppliers ready for new robot safety standards? appeared first on The Robot Report .

ICE are heavily armed killers. They’re also huge losers
Donald Trump's Homeland Security regime has been at the center of two critical stories in the past two weeks. In the first, federal agents shot and killed a man and quickly got to work justifying the use of force under the flimsiest of pretenses. In the other, it made house calls to people who said […]

That Is Embarrassing: Why Frontier AI Still Makes Things Up, and What to Do About It
The best AI models still hallucinate. These hallucinations are sometimes funny, and sometimes cause actual damage. In this post we will consider recent tales of AI hallucinations, and then look under the hood to understand why they happen. The post That Is Embarrassing: Why Frontier AI Still Makes Things Up, and What to Do About It appeared first on Towards Data Science .
Ask HN: What are some of the use-cases of the frontier models's max mode?
Recently ChatGPT released an Ultra mode, it's "highest-capability setting, coordinating multiple agents across parallel workstreams to finish complex tasks faster" on their latest flagship product Sol of GPT-5.6. Similarly, Claude Fable also has an Ultra setting. I wonder what are primary workflows or use-cases this maximised modes are used for currently? In my average software development experience, I never had a need to go beyond the max mode or beyond Opus for example. Even in specific workf
Integrating Lambda Durable Functions into a Step Functions Workflow
<p>At re:Invent 2025, AWS <a href="https://aws.amazon.com/about-aws/whats-new/2025/12/lambda-durable-multi-step-applications-ai-workflows/" rel="noopener noreferrer">announced Lambda Durable Functions</a>. The feature introduces a <strong>checkpoint/replay mechanism</strong> that allows Lambda executions to run for up to one year, automatically recovering from interruptions by replaying from the last checkpoint.</p> <p>Lambda's 15-minute timeout is not a bug or a limitation to work around. It is
fix(agents): truncate multibyte middleware details at byte limit (#10…
<pre style='white-space:pre-wrap;width:81ex'>fix(agents): truncate multibyte middleware details at byte limit (#104156) * fix(agents): bound multibyte middleware details by bytes * refactor(agents): bound middleware detail byte sizing --------- Co-authored-by: Peter Steinberger <steipete@gmail.com></pre>
Why robotics teams need virtual gyms before deployment
To address environmental and task variability, robots can benefit from 'virtual gyms' to bridge the sim-to-real gap, says SoftServe. The post Why robotics teams need virtual gyms before deployment appeared first on The Robot Report .

GDPR retention and erasure for an agent mailbox
<p>Most "AI email" demos never think about deletion. The agent reads, replies, files things away, and the inbox just grows. That's fine in a demo. It is a problem the first time a real person emails your agent, because the moment that mailbox holds someone else's name, address, order history, or support complaint, you've taken on a data-protection obligation — and "we kept everything forever" is not a defensible retention policy.</p> <p>An <strong>Agent Account</strong> on Nylas accumulates pers
Keep your agent's mail out of spam traps
<p>Spam traps are the failure mode nobody puts in the demo. A bounce is loud — you get a <code>5.x.x</code> back, your code logs it, you move on. A complaint at least gives you a webhook. A spam trap gives you <em>nothing</em>. The message gets accepted, no error comes back, and somewhere a mailbox provider quietly writes your domain down as a spammer. By the time you notice, your inbox placement has already cratered and you have no single bounce to point at.</p> <p>That's the trap, literally. A
AI takes two-thirds of venture money, and your odds are still one in six
Article URL: https://okaneland.com/study/ai-startup-raise-math/ Comments URL: https://news.ycombinator.com/item?id=48871449 Points: 2 # Comments: 0
I built an AI code reviewer with 6 parallel agents. Here’s the architecture, warts and all.
<p>Six months ago I got fed up with my AI code review tool.</p> <p>Two problems.</p> <ul> <li>It charged roughly <strong>3.5× the raw OpenAI API cost</strong> while pretending that wasn't happening.</li> <li>It kept flagging phantom SQL injection in a codebase that literally had <strong>no SQL</strong>, while completely missing a real authorization bypass in the same pull request.</li> </ul> <p>That combination broke my trust.</p> <p>So I cancelled my subscription and started building my own.</p
I Analyzed 4,788 AI Coding Sessions — Here's Where Your Tokens Actually Go
<p>Last month I started tracking every command I ran through Claude Code, Cursor, and Aider. After 4,788 commands and 355 million tokens, I found something shocking:</p> <p>97.6% of my tokens were wasted on noise.</p> <p>Not on actual coding. Not on debugging. On repetitive test output, build logs, and progress bars that the AI didn't need to see.</p> <p>THE NUMBERS</p> <p>Commands tracked: 4,788<br> Total tokens processed: 355,785,039<br> Tokens that were actual content: 8,694,751<br> Tokens th