Claude Sonnet 5: Anthropic's Most Agentic AI Model Arrives at a Reduced Price (2026)
Article URL: https://lucasaguiar.xyz/en/posts/claude-sonnet-5-2026/ Comments URL: https://news.ycombinator.com/item?id=48812163 Points: 3 # Comments: 0
Article URL: https://lucasaguiar.xyz/en/posts/claude-sonnet-5-2026/ Comments URL: https://news.ycombinator.com/item?id=48812163 Points: 3 # Comments: 0
You can save up to 73% off a range of Dreame's automated smart home products right now.

This study focuses on Text-to-Sounding-Video (T2SV) generation, which aims to generate a video with synchronized audio from text, with both modalities aligned to the text conditions. Despite progress in joint audio-video training, two critical challenges remain: (1) text conditioning is a bottleneck—shared captions (TV=TA) trigger modal interference, while a gap persists between dense training captions and concise inference user prompts, and (2) the optimal fusion mechanism for cross-modal featu
Multi-domain fine-tuning of large language models requires improving performance on target domains while preserving performance on constrained domains, such as general knowledge, instruction following, or safety evaluations. Existing data mixing strategies rely on fixed heuristics or adaptive rules that cannot explicitly enforce preservation of such capabilities. We propose DynaMiCS, a dynamic mixture optimizer that casts multi-domain fine-tuning as a constrained optimization problem. At each up
The web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection attempts remain limited to offline trajectories for supervised fine-tuning or a handful of simulated environments for RL training, thus failing to capture web diversity. We propose Weblica (Web Replica), a framework for constructing reproducible and scalable web environments. Our framework leverages 1) HTTP-level caching to capture and replay stabl
While large language models (LLMs) and coding agents are often applied to user interface (UI) development, developers find it difficult to reliably assess their proficiency in visual and interaction design. Existing evaluations either rely on human experts, who can accurately assess usability by testing critical flows but are slow and costly, or on automated judges, which are scalable but less accurate and opaque. We present FlowEval, a reference-based framework that measures whether a generated
Safety alignment in language models operates through two mechanistically distinct systems: refusal neurons that gate whether harmful knowledge is expressed, and concept neurons that encode the harmful knowledge itself. By targeting a single neuron in each system, we demonstrate both directions of failure — bypassing safety on explicit harmful requests via suppression, and inducing harmful content from innocent prompts via amplification — across seven models spanning two families and 1.7B to 70B
GitHub Tools now ships an eve toolset through the new @github-tools/sdk/eve subpath. One file in agent/tools/ can register every GitHub tool, or use a preset such as maintainer , so you can build a complete GitHub agent in nine lines of code. Safe by default: Every write tool, such as mergePullRequest , requires approval unless you opt out. Gate individual tools with always , once , or an input-dependent predicate; pauses survive restarts and deploys. Presets: code-review , issue-triage , repo-e
See how Australian Payments Plus uses ChatGPT Enterprise and Codex to move faster through payments complexity. AP+ saves time, improves quality, and keeps human judgment central.
MUFG uses ChatGPT Enterprise to build an AI-native organization, improve workflows, and deliver new AI-powered financial services at scale.
Hiya! So I've been playing around with having Claude make videos for a bit now even had some success posting the results to TikTok (and setup a whole pipeline so Claude can generate and post autonomously). With the release of Nano Banana 2 Lite, I was curious show fast I could make the generation, so last night I gave it a whirl and got down to around 30s for short-form video. It uses GLM-5.2 fast via Fireworks to generate the scripts and image prompts and, like I said, Nano Banana 2 Lite for th
<p><strong><a href="https://huggingface.co/tencent/Hy3">tencent/Hy3</a></strong></p> New Apache 2.0 licensed model from Tencent in China:</p> <blockquote> <p>Hy3 is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tencent Hy Team. Following the Hy3 Preview launch in late April, we gathered feedback from 50+ products and scaled up post-training with higher quality data. Today, we introduce Hy3, which outperforms similar-siz
An AI agent carried out the technical execution of a real-world ransomware attack for the first known time, but new details show a human still chose the victim, set up the infrastructure, and supplied stolen credentials — meaning it wasn't quite the fully autonomous cybercrime debut that last week's headlines suggested.
Robotics tech is changing fast, so for many it makes sense to rent a robot.

Model dev and AGI fearmonger Anthropic signs 20-year lease with TeraWulf
Article URL: https://github.com/ninjahawk/Subtext Comments URL: https://news.ycombinator.com/item?id=48811892 Points: 5 # Comments: 0
Training state-of-the-art large language models (LLMs) with billions of parameters requires distributed training across hundreds or thousands of GPUs. At this scale, hardware failures are not exceptional events—they are expected....

SK Hynix is experiencing a boom credited to AI. It will ride that to a multi-billion dollar US IPO, expected to take place on Friday.
<pre style='white-space:pre-wrap;width:81ex'>fix(openai): bound Codex OAuth token response body reads with readResponseWithLimit (#99479) * fix(openai): bound Codex OAuth token response body reads with readResponseWithLimit Replace unbounded response.arrayBuffer() in postTokenForm with readResponseWithLimit using a 1 MiB cap to prevent OOM from oversized token endpoint responses. Add real node:http loopback server tests. * fix(openai): wrap readResponseWithLimit result in Uint8Array for TS BodyI
<pre style='white-space:pre-wrap;width:81ex'>docs(provider): document Featherless AI setup</pre>
<pre style='white-space:pre-wrap;width:81ex'>feat(provider): add Featherless AI integration</pre>
Article URL: https://www.esri.com/en-us/software-engineering/blog/articles/ai-orchestration-in-the-browser Comments URL: https://news.ycombinator.com/item?id=48811674 Points: 2 # Comments: 0
<pre style='white-space:pre-wrap;width:81ex'>feat(models): add Claude Sonnet 5 support (#98254) * feat(models): add Claude Sonnet 5 support Co-authored-by: Ariel Bravy <ariel@vortexradar.com> * fix(models): align Sonnet 5 validation baselines * fix(models): satisfy Sonnet 5 CI contracts * fix(models): enforce Sonnet 5 provider contracts * docs(changelog): defer Sonnet 5 release note --------- Co-authored-by: Peter Steinberger <steipete@gmail.com> Co-authored-by: Ariel Bravy <ariel@vortexradar.co
<p>Scottish government to consider SNP national council motion for moratorium on all new datacentres</p><p>The Scottish government is about to consider a sweeping moratorium on building new datacentres, putting a key plank of the UK’s AI strategy at risk.</p><p>Last Sunday the Scottish National party (SNP)’s national council passed a motion to freeze all new datacentres in Scotland. That motion has been sent to the Scottish government to consider.</p> <a href="https://www.theguardian.com/uk-news

<pre style='white-space:pre-wrap;width:81ex'>fix: preserve preflight overflow token counts into recovery budgeting (#101181) Budget context engine assembly against the reserve and rendered prompt pressure, and carry the preflight estimated prompt tokens, prompt budget, and overflow tokens into the outer overflow recovery loop so compaction engines compact against the prompt OpenClaw actually rendered instead of a minimally over-budget guess.</pre>
<p> Gather people instantly by notifying or ringing their phones </p> <p> <a href="https://www.producthunt.com/products/say-who-say-why-kadoink-gathers-everyone?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1189761?app_id=339">Link</a> </p>
<p>The morning I announced what I'd been building, a comment showed up on the post. It was friendly. It opened with a compliment, agreed with me, and then walked through a few of the risk signals worth thinking about...who owns the code, what a change actually touches when it runs, whether it goes anywhere near auth or payments. Solid stuff. It was also every point I'd made in the post it was commenting on.</p> <p>Then it suggested I go build a tool that scores that risk automatically and flags
Article URL: https://old.reddit.com/r/BuyFromEU/comments/1up518w/proton_now_using_100_chinese_llms_drops_european/ Comments URL: https://news.ycombinator.com/item?id=48811481 Points: 12 # Comments: 0
<h2> The problem </h2> <p>My Mac hard-crashed with seven Claude Code CLI sessions open, across five repos. All gone.</p> <p>The built-in recovery is <code>claude --resume</code>, and it doesn't scale to seven:</p> <ul> <li>It only lists sessions for the <strong>directory you run it from</strong> — you have to remember which repos you were even in, then <code>cd</code> into each one.</li> <li>It <strong>takes over the shell</strong> you run it from — so you're spawning windows and retyping paths
Article URL: https://www.vybe.build/blog/learn-what-not-to-tokenize Comments URL: https://news.ycombinator.com/item?id=48811403 Points: 7 # Comments: 3
Today, we’re excited to announce a deep-link integration between Hugging Face and Amazon SageMaker AI. Developers can now go from model discovery to hands-on experimentation in SageMaker Studio with a single selection.
Article URL: https://github.com/last9/gpu-telemetry Comments URL: https://news.ycombinator.com/item?id=48811382 Points: 2 # Comments: 0

<p>VEQRA is an existing Microsoft 365 automation platform I built that detects and routes enterprise incidents. But it couldn't answer the critical questions: Why did this happen? What is the financial impact? What should we do right now?</p> <p>The Qwen Cloud Global AI Hackathon was the opportunity to build the intelligence layer VEQRA was missing.</p> <h2> What it does </h2> <p>VEQRA AI orchestrates three specialized AI agents that resolve a critical enterprise incident in 13 seconds:</p> <ul>
Article URL: https://github.com/hritvikgupta/nimbus Comments URL: https://news.ycombinator.com/item?id=48811086 Points: 1 # Comments: 0
A single word may be putting them at risk. ScienceAlert stories are written, fact-checked, and edited by humans, never generated by AI. Don't miss a story, subscribe here.

I think coding benchmark results don't represent our messy reality. As they 1) Use purpose-built test harnesses We use Claude Code or Codex 2) Test one-shot tasks We work in sessions Sessions are messy, we start with a large primary task, then some cleanup, an adjacent fix here another over there. We start, stop, and change our minds. That changes both cost and quality. Cache TTLs expire. Context grows. I am working on creating one. Here's my rough plan: 1) Use Claude Code and Codex 2) Use sessi
Article URL: https://github.com/eranif/kennel Comments URL: https://news.ycombinator.com/item?id=48810920 Points: 2 # Comments: 0