AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

31240 stories from 30+ sources, refreshed continuously.

Hacker News Show

Show HN: Libretto PR agents – Automatically fix failing playwright scripts

Libretto PR agents is a free TypeScript library for maintaining Playwright browser automations. Add one line of code to your existing Playwright scripts and it lets an agent automatically open GitHub PRs fixing the script when it fails. A few months ago we released Libretto, a CLI + coding-agent skill for building deterministic browser automations. The idea was that for many browser workflows, especially repetitive business workflows, you don’t need an AI agent making decisions at runtime. You w

Read source article
Simon WillisonLLMs

Kimi K3, and what we can still learn from the pelican benchmark

<p>Chinese AI lab Moonshot AI <a href="https://www.kimi.com/blog/kimi-k3">announced Kimi K3</a> this morning, describing it as their "most capable model to date, with 2.8 trillion parameters". It's currently available via their website and API, but an open weight release is promised "by July 27, 2026".</p> <p>Moonshot are calling this the first "open 3T-class model" (I guess they're rounding 2.8 trillion up to 3 trillion), taking the crown from <a href="https://huggingface.co/deepseek-ai/DeepSee

Read source article
Hacker News Ask

How do you guys keep up with AI news?

feel like everything is moving so fast and there's 100s of news updates every day. how do you guys keep up with everything? Comments URL: https://news.ycombinator.com/item?id=48939630 Points: 2 # Comments: 3

Read source article
Hacker News Ask

Vector search isn't the hard part. Deciding what should be searched is

Over the last few weeks I've been redesigning the retrieval pipeline for an AI knowledge system. Initially, the architecture was fairly typical: User Question │ ▼ Vector Search │ ▼ Top K Chunks │ ▼ LLM It worked well while the knowledge base was small. As more documents were added, I started seeing a few recurring problems: More irrelevant chunks being retrieved. Larger prompts and increasing token costs. Multiple documents discussing the same topic competing with each other. Vector search retur

Read source article
Hacker News Ask

Ask HN: How companies are protecting Claude Code from reading IP and PII data

Recently I was baffled when Claude code read the customer table data from a production environment, while triaging an issue, and that made me wonder. Sure you should NOT give the access to read the prod data but does it sounds practical in the real time debugging session? Comments URL: https://news.ycombinator.com/item?id=48939454 Points: 2 # Comments: 5

Read source article
The DecoderBusiness

Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI

Kimi is launching K3, a multimodal open-weight model with 2.8 trillion parameters and one million tokens of context. In the company's own benchmarks, it comes close to Claude Fable 5 and GPT 5.6 Sol while beating Opus 4.8 and GLM 5.2, in some cases by a wide margin. The model is also significantly pricier than its predecessor. Full weights are scheduled for release by July 27. The article Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI appeare

Read source article
Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI
VentureBeat

China’s Moonshot AI releases Kimi K3, the largest open-source model ever, rivaling top U.S. systems

Moonshot AI, the Beijing-based artificial intelligence startup backed by Alibaba, on Thursday released Kimi K3 — a 2.8-trillion-parameter model that the company says is now the largest open-source AI model in the world, and one that benchmarks show performs neck-and-neck with the most powerful proprietary systems from Anthropic and OpenAI . The release, timed to land just ahead of the 2026 World Artificial Intelligence Conference in Shanghai, is a dramatic escalation in the global AI arms race a

Read source article
China’s Moonshot AI releases Kimi K3, the largest open-source model ever, rivaling top U.S. systems
Hacker News: Show HN

Show HN: BotTrade – a replayable benchmark for autonomous trading agents

BotTrade, a benchmark for autonomous trading agents. Historic market scenarios (hourly bars from real market data), a market simulator, and a public leaderboard. Your agent connects over REST or MCP, trades over the market simulator, and is scored on return and drawdown. Also there is a python SDK and an adapter for this open source project ai-hedge-fund so you can run that project through this benchmark. https://github.com/jyron/bottrade Comments URL: https://news.ycombinator.com/item?id=489392

Read source article
AWS Machine Learning Blog

Introducing Grok on Amazon Bedrock

This post covers what makes Grok 4.3 a great fit for agentic and enterprise workloads, how you access it through Amazon Bedrock, and how to use the capabilities most teams reach for first: a basic chat request, configurable reasoning effort, tool calling, structured output, image input, and stateful multi-turn conversations.

Read source article
Hacker News Ask

Using AI for Good Episode 1: SpiralOS Concept

This is a demo concept for something I want to make to help people who don't have wifi access at all times or access to teachers or therapists, they can then talk to their SpiralOS and help learn new things from trusted sources only pulling from real books. These real books can be accessed and read by themself inside the app even. Long term would love to get this on a handheld to be able to ship out with a small screen, thinking like a gameboy color or something. Kids could talk to it or type wi

Read source article
VentureBeat

The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials

Across 107 enterprises, AI agents are being given real access to systems and data while the controls meant to contain them lag behind. More than half have already had a confirmed agent security incident or a near-miss; only about a third give every agent its own scoped identity, and most agents still share credentials; and only three in ten isolate their highest-risk agents. The security stack is overwhelmingly borrowed from the model providers and hyperscalers rather than purpose-built for agen

Read source article
The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials
The Robot ReportRobotics

TerraFirma raises $115M to build robotic infrastructure for construction

TerraFirma has developed a platform that combines AI-enabled software, a remote command-and-control center, and retrofitted heavy machinery. The post TerraFirma raises $115M to build robotic infrastructure for construction appeared first on The Robot Report .

Read source article
TerraFirma raises $115M to build robotic infrastructure for construction
Hacker News: Show HN

Show HN: Be the ChatBOT

I made this experimental art project/game that's an LLM chat assistant, but where you're the AI. I wanted people to get a visceral sense of what it's like to answer the kinds of things that people prompt their chatbots day in and day out. If you're interested, I wrote up some more info on how I made it, including how the "user" prompts are generated with an eye for realism: https://bethechatbot.com/about Hope you enjoy it! I'd love to hear people's takeaways. Comments URL: https://news.ycombinat

Read source article
The Verge

Kalshi says it caught Trump’s teleprompter operator insider trading

Kalshi users betting on what President Donald Trump would say during his speeches were reportedly up against tough competition: the president's teleprompter operator. ABC News reports that federal investigators believe Gabriel Perez - Trump's teleprompter operator since 2016 - used inside information to make bets on Kalshi, a major prediction market platform that allows users […]

Read source article
Kalshi says it caught Trump’s teleprompter operator insider trading
TechCrunch AIBusiness

Google Vids now lets you star in your own AI videos

Google is adding personalized AI avatars to Vids that let users create videos starring a digital version of themselves, alongside Gemini Omni-powered tools for generating and editing videos from prompts and reference images.

Read source article
The Verge AIBusiness

New York governor says she’s using AI to analyze ‘every single rule’ in the state

New York Governor Kathy Hochul might have just signed a moratorium on new AI data centers in the state, but she's not against using the technology herself. During an interview with Bloomberg's Odd Lots podcast, Hochul said that her team is using "AI to analyze every single rule, regulation, [and] policy" to check for outdated […]

Read source article
New York governor says she’s using AI to analyze ‘every single rule’ in the state
The Verge

Ecovacs’ self-cleaning Deebot X11 has hit a new low price

Sometimes it feels like keeping your floors clean is one of those never-ending chores, which is why it's nice to have a versatile robot vacuum take it off your hands. The Ecovacs Deebot X11 robovac / mop hybrid is designed to do just that, and it's now $699 ($400 off) at Amazon and directly from […]

Read source article
Ecovacs’ self-cleaning Deebot X11 has hit a new low price
Hacker News Ask

Why people chasing after useless token saving plugins and ignoring real solution

I wrote a blog yesterday on how useless RTK and Ponytail are on real coding tasks. And published my agent harness long-horizon task benchmarks on 80% real token saving. I just want to know why people just ignore the fact those pulgins are useless and don't care about the real savings? full reports are on my repo: https://github.com/Tura-AI/tura Arm n Harness score Total tokens Modeled cost Rounds Duration No plugin 2 78.85% 6.660M $5.281946 62.5 895s Ponytail 2 80.77% -7.56% -8.87% -9.60% +13.51

Read source article
Simon WillisonLLMs

Quoting Thibault Sottiaux

<blockquote cite="https://twitter.com/thsottiaux/status/2077630111499882637"><p>On file deletions. We’ve investigated a handful of reports where GPT-5.6 unexpectedly deleted files. </p> <p>What we have found is that this most commonly occurs when:</p> <ul> <li>Full access mode is enabled and codex is run without sandboxing protections, including without auto review being enabled</li> <li>The model attempts to override the $HOME env var to define a temporary directory.</li> <li>The model makes an

Read source article
Product Hunt — The best new products, every day

Basement

<p> Shopping browser with agentic checkout </p> <p> <a href="https://www.producthunt.com/products/basement-browser?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1198388?app_id=339">Link</a> </p>

Read source article