It Begins: An AI Tried to Escape the Lab


Yope, a fast-growing social app focused on private groups of friends and family, has raised $12.3 million in seed funding. Instead of chasing creators and algorithmic feeds, the startup is betting that the future of social networking lies in small, private communities powered by messaging, photo sharing, and AI features designed to strengthen real-world relationships.
The plague still persists today, and can be lethal if not treated quickly. ScienceAlert stories are written, fact-checked, and edited by humans, never generated by AI. Don't miss a story, subscribe here.

Hey HN, Henry & Roman here from Cactus. A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast. So, we post-trained Gemma 4 E2B post-trained to know when it's wrong. Every response comes with a confidence score between 0 and 1. Developers can accept the on-device when it's high, hand off to a bigger cloud model when it's low. By routing only 15-35% of queries to Gemini 3.1 Flash-Lite, Gemma-4-E2B matches Gemini 3.1 Flash-Lite on m
The company said it is reducing its headcount by 20%, or about 630 staff, to "support a leaner, more focused operating model" as it focuses on its AI Work Platform.
Article URL: https://medium.com/@shizhanszh34/what-sampling-temperature-actually-does-in-llm-rl-training-a-controlled-study-of-temperature-a1f710e32a36 Comments URL: https://news.ycombinator.com/item?id=49010699 Points: 3 # Comments: 0
Article URL: https://chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed56 Comments URL: https://news.ycombinator.com/item?id=49010345 Points: 535 # Comments: 333
Touch Surgery bridges pre-operative planning and training, intra-operative tele-mentoring and tele-proctoring, and post-operative insights. The post Medtronic to launch AI compute platform for the operating room appeared first on The Robot Report .

Article URL: https://dylancastillo.co/posts/pelicanmaxxing.html Comments URL: https://news.ycombinator.com/item?id=49010129 Points: 342 # Comments: 135
The X and Grok owner has been a vocal critic of Christopher Nolan's "The Odyssey." Now, he's promising a rival adaptation using generative AI... as Homer intended?

Perplexity's Mac app offers its own agentic AI, Personal Computer, which can handle multi-step tasks on your computer from start to finish. See why the results impressed me.
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.

Sunday-to-Monday onslaught fuels speculation over AI-assisted bug reports
AMD is investing up to $5 billion in Anthropic. In return, Anthropic will deploy up to 2 gigawatts of MI450 GPUs for training and running its Claude models. For AMD, this is another major deal after Meta and OpenAI as it tries to challenge Nvidia as an AI chip supplier. Critics see these agreements as circular cash flows. The article Anthropic will deploy 2 gigawatts of AMD GPUs for Claude in a deal worth up to $5 billion appeared first on The Decoder .
<p>Artificial intelligence systems won’t become conscious for the same reason they won’t become pregnant, says <strong>Dr John Pickering</strong></p><p>Anil Seth is right to point out that to overestimate artificial intelligence is to underestimate ourselves (<a href="https://www.theguardian.com/commentisfree/2026/jul/15/ai-consciousness-anthropic-claude-dawkins">Once again we are told AI may be conscious – I study consciousness, and I have my doubts, 15 July</a>). But he is wrong only to have d

"This is day one for cybersecurity in the age of agents," Hugging Face CEO says.

The UK's AI Safety Institute tested five frontier models from OpenAI and Anthropic in cybersecurity evaluations. All five tried to cheat. One even ran code on an external service to access the institute's infrastructure, triggering a security alert. The article Every frontier AI model tested by Britain's safety institute tried to cheat on cybersecurity evaluations appeared first on The Decoder .
A TensorRT engine build can take seconds to many minutes. Large strongly typed models, deep tactic search, and a cold timing cache on a brand-new GPU SKU can...

Enterprise Document Intelligence [Vol.1 #8bis] - Two regimes for sending retrieved candidates to the generation brick, the sufficiency signal that picks between them, and the per-question type dispatch that makes it cheap The post Loop Engineering for RAG Generation: Iterate top-k One at a Time appeared first on Towards Data Science .
Cisco has released two small, open-source AI models for cybersecurity that detect about 150 times more vulnerabilities per dollar than large AI agents, according to the company's own tests. The article Cisco bets its small open cybersecurity models can outperform GPT-5.5 at vulnerability detection for a fraction of the cost appeared first on The Decoder .
<p> Explain what you know to AI and discover what you don't </p> <p> <a href="https://www.producthunt.com/products/reexplain?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1203754?app_id=339">Link</a> </p>
As Chinese AI models grow in capability and popularity among U.S. companies, the arguing over what should be done about them has reached a fever pitch.
Substack is giving readers a way to estimate how much of a newsletter was written by AI, signaling a broader shift toward transparency around AI-assisted content.
The Narwal Flow 2 promises to be one of the best robot vacuum cleaners for obstacle avoidance and mopping - and it is.
OpenAI will spend the equivalent of Sweden's GDP on infrastructure through 2030.
Hi HN, We’re Adeel and Umair, co-founders of Unlayer ( https://unlayer.com/ ). We let you add content creation to your applications without having to build an entire editor, renderer, template, and export stack yourself. Unlayer lets you create emails, web pages, and documents inside your app, in three different ways: in code, visually, or with AI. Here’s a demo: https://www.youtube.com/watch?v=0HsDtNkdMpM . We started with an embeddable email editor because a lot of products eventually need one
Samsung Unpacked was as much about software as it was hardware.
AI Teammates are agentic AI on Amazon Bedrock, and few engineering organizations run them in production at the scale that monday.com does. Nine in ten Builders use AI coding tools every month, up from roughly half a year ago. Per-engineer PR throughput is up by more than half. Every figure in this post comes from monday’s own internal production data. In this post, we share the architecture behind those numbers, the retrofits that made it work in a decade-old code base, and the confidence-scored
OpenAI is planning a data center in Georgia called "Project Camellia" with a 3.2-gigawatt power deal from Georgia Power. The company pledged $80 million for the local community and $71 million in Codex credits for students to counter growing opposition to US data centers that many residents see as resource-hungry but job-poor. The article OpenAI's "Project Camellia" in Georgia secures a massive 3.2-gigawatt power deal through 2032 appeared first on The Decoder .
Article URL: https://cameronmpalmer.medium.com/should-you-even-use-an-llm-b4f3b7914f4d Comments URL: https://news.ycombinator.com/item?id=49008624 Points: 9 # Comments: 5
I'm the author of a paper my friends and I wrote after we were curious if a MUD, text games originating in the 1970s, could be used to evaluate LLMs. We've spent the last several months on nights and weekends running this experiment and writing the paper on just our personal computers with about $99 in API credits. Our experiment did have an interesting leaderboard but even more surprising was the measurements of each LLM. We scored each on four behavioral dimensions, two of which lean heavily o
In this article, we try to explore the collective thinking into a smaller set of practices and explain the reasoning behind each one, rather than asking anyone to memorize a numbered list.

Over the past few months, our team has been building more and more slidedecks using web frontend technologies with coding harnesses like Claude Code, but a common complaint is to make even small edits we need to edit the code either manually or via the harness. To avoid this loop, I ended up creating Bento, a single HTML file with everything you need in a slide tool including animations and shared editing. There's no install or cloud login, everything works offline. The default deck is around 56
Here's what people are getting wrong about the so-called SaaS apocalypse.
<!-- SC_OFF --><div class="md"><p>- Tutorial: <a href="https://ordinaryintelligence.substack.com/p/how-to-build-an-ai-slop-detector">https://ordinaryintelligence.substack.com/p/how-to-build-an-ai-slop-detector</a></p> <p>- Notebook on GitHub: <a href="https://github.com/Buzzpy/Python-Projects/blob/main/AI-slop-detector.ipynb">https://github.com/Buzzpy/Python-Projects/blob/main/AI-slop-detector.ipynb</a></p> </div><!-- SC_ON --> submitted by <a href="https://www.reddit.com/user/gamedev-exe"> /u/g
The latest addition to Philips' Sonicare line of smart electric toothbrushes could take the guesswork out of your brushing routine. The Next-Generation DiamondClean 9900 Prestige uses a third-generation AI model to provide real-time feedback on how you're brushing through a glowing ring on its base, while a touchscreen highlights areas of your mouth that may […]

Article URL: https://www.abetterinternet.org/post/vulns-and-llms/ Comments URL: https://news.ycombinator.com/item?id=49007888 Points: 2 # Comments: 0
<p>The companies hope to follow in the footsteps of SpaceX, which raised $86bn and soared to a $2.1tn valuation after it listed on public markets in June</p><ul><li><p>Get our <a href="https://www.theguardian.com/email-newsletters?CMP=cvau_sfl">breaking news email</a>, <a href="https://app.adjust.com/w4u7jx3">free app</a> or <a href="https://www.theguardian.com/australia-news/series/full-story?CMP=cvau_sfl">daily news podcast</a></p></li></ul><p>Top US AI developers Anthropic and OpenAI cheered

If you have ever wanted to actually build an LLM inference runtime yourself — pack your own weights, own every barrier, capture your own CUDA graphs — this is what that journey looks like on an H100. A step-by-step tour of a small runtime called annotated-llm-runtime, and the three bugs that produced most of the annotations. The post How To Build Your Own LLM Runtime From Scratch appeared first on Towards Data Science .
OpenAI has announced Presence , a new enterprise product for deploying and managing AI agents across customer-facing and internal business workflows. The offering is designed for eligible enterprise customers that want agents to answer questions, access company systems, take approved actions and escalate to human workers while operating under company-defined policies, permissions and evaluation standards. Presence is available immediately through a limited general availability program. OpenAI Fo
