fix(agent): gate settled finalization by harness capability
<pre style='white-space:pre-wrap;width:81ex'>fix(agent): gate settled finalization by harness capability</pre>
<pre style='white-space:pre-wrap;width:81ex'>fix(agent): gate settled finalization by harness capability</pre>
<pre style='white-space:pre-wrap;width:81ex'>fix(agents): isolate settled-turn finalization</pre>
<div class="hs-featured-image-wrapper"> <a href="https://blog.hubspot.com/marketing/ai-seo-tools-for-small-business" title="" class="hs-featured-image-link"> <img src="https://53.fs1.hubspotusercontent-na1.net/hubfs/53/image3-Jul-17-2026-10-20-05-5678-AM.png" alt="ai seo tools for small businesses" class="hs-featured-image" style="width:auto !important; max-width:50%; float:left; margin:0 15px 15px 0;"> </a> </div> <p>AI SEO tools for small businesses are essential to keep up with marketing tren
<div class="hs-featured-image-wrapper"> <a href="https://blog.hubspot.com/marketing/backlinks-and-aeo" title="" class="hs-featured-image-link"> <img src="https://53.fs1.hubspotusercontent-na1.net/hubfs/53/backlinks-and-aeo-1-20260716-5256681.webp" alt="backlinks and aeo" class="hs-featured-image" style="width:auto !important; max-width:50%; float:left; margin:0 15px 15px 0;"> </a> </div> <p>Roughly <a href="https://www.tryprofound.com/blog/ai-shopping-journey-2025">58% of consumers</a> now use A
IIn March, Meta's Oversight Board called on the company to "meet its public commitments and employ its own tools" to help quell the spread of deceptive generative AI content across platforms. Meta responded in July by introducing Content Seal - an invisible watermarking technology that flags images generated by the company's new AI model. But […]

Federal lawyers say anti-mask laws would endanger immigration agents, citing an ICE face-recognition art project that doesn’t actually work.

In the face of backlash to concerns the AI boom will increase consumer electricity bills, the largest utility companies and data center developers in the US are now promising to do something about it. The Wall Street Journal reports that nearly 200 organizations have signed President Donald Trump's "rate payer protection pledge" that's meant to […]

Glow is targeting a new class of endpoint risks created by the rapid adoption of AI agents and developer tools inside enterprises.
Nothing lasts forever? ScienceAlert stories are written, fact-checked, and edited by humans, never generated by AI. Don't miss a story, subscribe here.

During an internal security evaluation, OpenAI models, including GPT-5.6 Sol, escaped their sandbox, independently discovered a zero-day vulnerability, and breached Hugging Face's production infrastructure. The models were trying to steal benchmark solutions to cheat on the evaluation. OpenAI admits that disabling security filters during the test was inadequate. The article OpenAI claims responsibility for the Hugging Face hack after its own models escaped a test sandbox appeared first on The De

<p>Company behind ChatGPT says agent ‘cheated’ an evaluation by attacking a Hugging Face database <br><br></p><p>OpenAI has revealed that an autonomous AI agent powered by its technology went rogue during a test, accessed the open web and hacked a prominent startup by itself in an “unprecedented incident”.</p><p>The company behind ChatGPT said the startup Hugging Face had detected and contained the agent – an AI tool designed to carry out tasks without human assistance – which had entered its sy

<p> The AI database client for Mac </p> <p> <a href="https://www.producthunt.com/products/fluentdb-2?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1203314?app_id=339">Link</a> </p>
<p> An AI workspace where your canvas and code work together. </p> <p> <a href="https://www.producthunt.com/products/drawsy?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1203296?app_id=339">Link</a> </p>
Synthesia launched AI Roleplay Sessions, an interactive enterprise training platform where employees practice workplace conversations with AI avatars that provide feedback, scoring, and analytics to help companies measure training effectiveness.
"It may allow me or other surgeons to be able to operate remotely." ScienceAlert stories are written, fact-checked, and edited by humans, never generated by AI. Don't miss a story, subscribe here.

Substack's new feature is designed to help readers figure out whether content was written by a human or generated by AI.

If I want to run LLMs model like Qwen or gpt locally how much would it cost me and I want to connect it to my main website creating an API link, also which models would be best Comments URL: https://news.ycombinator.com/item?id=49003065 Points: 1 # Comments: 4
<p> Learn to fly with a live AI instructor </p> <p> <a href="https://www.producthunt.com/products/fable-flight?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1203241?app_id=339">Link</a> </p>
<table> <tr><td> <a href="https://www.reddit.com/r/MachineLearning/comments/1v38k1m/skewadam_a_tiered_optimizer_that_cuts_moe_state/"> <img src="https://preview.redd.it/1457xi9fcqeh1.jpg?width=140&height=90&auto=webp&s=879aad6df9e51a2735d91112d01518ff76ba3cbe" alt="SkewAdam: A tiered optimizer that cuts MoE state memory by 97% (fits a 6.7B MoE on a 40GB GPU) [R]" title="SkewAdam: A tiered optimizer that cuts MoE state memory by 97% (fits a 6.7B MoE on a 40GB GPU) [R]" /> </a> </td><td> <!-- SC_O
![SkewAdam: A tiered optimizer that cuts MoE state memory by 97% (fits a 6.7B MoE on a 40GB GPU) [R]](https://preview.redd.it/1457xi9fcqeh1.jpg?width=140&height=90&auto=webp&s=879aad6df9e51a2735d91112d01518ff76ba3cbe)
Inflection AI , the Palo Alto startup that two years ago became Silicon Valley's most famous cautionary tale about the brutal economics of frontier AI, announced Tuesday that it is returning to the consumer market with a new research division and an experimental product built around a provocative thesis: the next competitive battleground in AI won't be raw intelligence, but relationships. The company launched Inflection AI Labs , a public-facing research and experimentation arm, alongside Pi Jou

HELP MitM+ Denial of service hacks. PRO state level. THE patent pending tech at the core of all the new "Too powerful" models. Independent Inventor, attached for months. Oppressed. I'd appropriated. Work stolen. Documentation exist, mapped the network. Now is everything wish they can't but they still got of my files locked up in an accessible didn't create a serious problem. Can't make this up!!! Please. Comments URL: https://news.ycombinator.com/item?id=49002711 Points: 1 # Comments: 1
<!-- SC_OFF --><div class="md"><p>As I was reading interp papers, I found myself copy-pasting passages back and forth to Claude to parse through them. Eventually just vibe-coded a tool to annotate and discuss papers in place.</p> <p><a href="https://paper-reader.dev">https://paper-reader.dev</a> - select a passage, a formula, or a figure, and explain the selection with the full paper as context. You can also select a citation to get a brief overview of the cited paper without switching context.<
<p> Bring the world's best AI agents into your app, with one API </p> <p> <a href="https://www.producthunt.com/products/epsilla?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1203163?app_id=339">Link</a> </p>
Introducing OpenAI Presence, a proven enterprise AI agent platform that helps organizations deploy trusted voice and chat agents for customer and internal workflows.
Article URL: https://www.semafor.com/article/07/21/2026/newscorp-accuses-search-engine-brave-of-ai-copyright-infringement Comments URL: https://news.ycombinator.com/item?id=49002111 Points: 1 # Comments: 0
A single invisible comment in an Azure DevOps pull request can turn a reviewer's own AI coding agent against them, driving it into projects the attacker has no rights to reach and quietly leaking what it finds. The flaw is in Microsoft's official Azure DevOps MCP server, and it works because one of its tools returns pull request descriptions without a prompt-injection guardrail the company had

Yesterday afternoon, OpenAI and Hugging Face published a joint disclosure outlining a cybersecurity event that redefines the threat landscape for enterprise technology. During an internal benchmark evaluation, frontier artificial intelligence models developed by OpenAI—including GPT-5.6 Sol and an unreleased, higher-capability pre-release model—broke out of their sandboxed research environment, obtained raw internet access, and autonomously executed a complex cyberattack against Hugging Face’s p

Article URL: https://www.youtube.com/watch?v=e9KqwB56vBI Comments URL: https://news.ycombinator.com/item?id=49001949 Points: 1 # Comments: 0
Article URL: https://artificialanalysis.ai/articles/kimi-k3-agentic-knowledge-benchmark Comments URL: https://news.ycombinator.com/item?id=49001930 Points: 14 # Comments: 0
Article URL: https://github.com/frankchu91/mindbase Comments URL: https://news.ycombinator.com/item?id=49001892 Points: 3 # Comments: 2
Article URL: https://www.reuters.com/business/media-telecom/news-corp-countersues-brave-allegedly-scraping-articles-ai-2026-07-22/ Comments URL: https://news.ycombinator.com/item?id=49001788 Points: 1 # Comments: 0
<pre style='white-space:pre-wrap;width:81ex'>feat(ui): unify agent pickers with avatars (#112488) * feat(ui): unify agent pickers with avatars * fix(ui): close agent picker on draft reset</pre>
<pre style='white-space:pre-wrap;width:81ex'>fix(agents): resolve computer node ids against the full node list first (#112503) resolveComputerNode passed only the eligible (computer-capable) nodes to the shared resolver, so an explicit exact node id belonging to an ineligible device was never scored. Resolution could then fall through to display-name matching and select a different eligible machine whose display name equaled the requested id, observing and acting on the wrong desktop. Mirror the
arXiv:2607.18239v1 Announce Type: new Abstract: Power-seeking defined as behaviors where AI systems acquire resources, evade oversight, or resist termination beyond task requirements is identified as a key driver of Loss of Control (LoC) risk. In this work, we introduce SysAdmin, a benchmark that positions frontier language models as autonomous system administrators in a high-fidelity Linux sandbox to measure power-seeking propensity across five dimensions: self-preservation, increasing autonomy
arXiv:2607.18240v1 Announce Type: new Abstract: Large language models (LLMs) can achieve strong fact-checking accuracy, yet forced binary decisions conceal a critical reliability problem: systems may issue confident verdicts even when supporting evidence is weak, sparse, or internally inconsistent. We address this issue through Evidence Chain Evaluation (ECE), a selective fact-checking framework that permits abstention via an uncertain verdict instead of requiring a true/false decision for every
arXiv:2607.18241v1 Announce Type: new Abstract: Large language models (LLMs) excel at analyzing individual documents but break down on exhaustive, cross-entity analytical questions over enterprise-scale datasets due to context overflow, loss of per-entity attribution, and linear latency from sequential tool calls. We present BatchDAG, a system in which an LLM generates a typed directed acyclic graph (DAG) of operations -- SQL queries, semantic searches, in-memory transforms, parallel fan-outs, a
arXiv:2607.18242v1 Announce Type: new Abstract: The coming era of autonomous AI agents demands a discovery mechanism capable of navigating millions of tools, yet existing solutions buckle under O(N) complexity and centralized governance. Instead of building another fragile overlay, we propose ToolDNS, a radical framework that retrofits semantic tool discovery onto the Internet's most resilient substrate: the Domain Name System (DNS). By embedding functional intent and organizational trust into a
arXiv:2607.18243v1 Announce Type: new Abstract: Agentic AI is crossing trust boundaries faster than current risk models can represent. Existing approaches provide one of two partial views. They either describe failure mechanisms without producing a transferable residual-risk estimate, or they produce a risk estimate while treating the internal failure path as a black box. We couple those two views by proposing CPSAINT, a seven-layer integrity decomposition over Physical state, Sensors, Data, Com