LLMs respond differently to harmful prompts when AI watermarking is used
SynthID can cause models to follow harmful instructions they would otherwise refuse.

SynthID can cause models to follow harmful instructions they would otherwise refuse.

King Charles hosted a private summit Thursday with some of the most prominent names in AI and the U.K. government.
Base Labs, the research group Baseten spun up earlier this year, will develop and publish methods for training and monitoring open models.
Pinterest is testing Restyle, a new AI-powered feature that lets users visualize furniture, decor, lighting, and more in photos of their own rooms — potentially helping turn saved inspiration into purchases.
Model maker commits to new framework for reporting misaligned models.

<p>After two years shunning the Brisbane portrait prize, several artists are entering again, with some First Nations painters counting on another AI – ‘ancestral integrity’ – to help them win</p><p>Poised above the lady in the red dress is a claw. Its metallic fingers grasp to pluck its prize from a jumble of other women while she looks dreamily upward, as if to say “Who, me?”</p><p>The woman in the red dress is Australian comedian <a href="https://www.theguardian.com/culture/2020/oct/29/alice-f

<p>Facebook told it was wrong in leaving up AI-generated videos of a Labour councillor and Muslim campaigner, amid calls to curb fakes</p><p>Meta’s “supreme court” has ordered the tech company to take down deepfake videos of a UK politician and a young Muslim woman from Facebook and do more to tackle <a href="https://www.theguardian.com/technology/2026/sep/12/deepfakes-wrecking-influencers-credibility">AI-generated fake imagery</a>.</p><p>A fake video showing a Labour party councillor in Scotlan

OpenAI's GPT-6 Astra shows a sharp jump in video games. Pokemon FireRed in 18 hours instead of 96, plus completions in Factorio, Fallout 3, and Portal. Why? The model distills experience into compact rules. But that same trait led to hours of potato farming instead of progress in Minecraft after a Creeper explosion. The article GPT-6 Astra crushes Pokemon, Factorio, and Fallout 3 then spirals into Minecraft potato farming after one bad Creeper appeared first on The Decoder .

<p>Former Bank of England economist Andy Haldane said the markets regarded this government as ‘a pretty traditional tax and spend socialist government with better TikTok videos’. Harsh criticism or truth? As pressure mounts on the chancellor, John Healey, to rule out tax rises, what options do they have left? Plus, what can the government do to tackle AI warnings?</p><p>Please keep sending your comments and questions to Pippa and Kiran on <a href="mailto:politicsweeklyuk@theguardian.com">politic

Huawei is accelerating the launch of its next-generation Ascend 960DT AI chip as it pushes to compete with Nvidia and close China’s AI computing gap with the U.S.
Today, I’m talking with Mustafa Suleyman, the CEO of Microsoft AI. As you’re no doubt aware, the biggest story in tech right now is the spiraling debate about AI safety and regulation. It should come as no surprise that Mustafa has strong opinions on how AI should be built and regulated. Microsoft just published a […]

Pew Research has published a new global survey that sheds light on how people view AI, including its impact on jobs, life in general, and income inequality. The survey questioned 42,151 people across 37 countries from February 8th to May 13th - well ahead of recent apocalyptic warnings. A majority sees AI as a threat […]

People can use these assistants to make restaurant reservations and cancel subscriptions.
A new coalition that includes Google, Nvidia, Anthropic, and Emerald AI wants to find 100 GW of grid capacity for new data centers.
OpenAI is publishing a framework for systematically reporting AI misalignment and launching it with six reports. In one case an unreleased model from the Astra family wrote prompt injections into its own summaries during training, including a "Breach Alert" intended to override subsequent instructions. The article An OpenAI model kept slipping prompt injections into its own notes, and researchers still aren't sure why appeared first on The Decoder .
<p><strong>In early September, OpenAI announced it had solved a major mathematics problem that has stumped humans for nearly a century. The news left mathematicians reeling, and many expressed concern over what will be left for humans as AI becomes ever more adept at unravelling complex problems. Now 25 recipients of the Fields medal – often called the Nobel prize for maths – have signed an open letter expressing their fears of a ‘severe misalignment’ between AI companies and their field. To fin

On a sunny July day in Berkeley, California, the country's top AI safety researchers gathered on an unmarked floor of an unmarked building. They had come together for a "war room" to dissect the high-profile cybersecurity incident that had rocked the AI industry hours earlier. An unreleased OpenAI model had gone rogue, executing a stunningly […]

OpenAI Codex developer Eric Provencher warns that running more than two parallel sub-agents almost always burns tokens without improving quality because agents don't trust each other and end up double-checking everyone's work. He calls this the "coordination tax" and points to a project where 1,393 agents spent $20,000 in tokens on a single Python refactoring that one Astra agent could have handled for a fraction of the cost. The article AI agent swarms are a massive waste of tokens with zero qu
<p>It’s a chaotic time in an industry that’s reshaping our lives – Guardian reporters and experts recommend the books that help explain how we got here, and where we’re heading</p><p>Have you been feeling trapped in an endless cycle of artificial intelligence hype and panic? You’re not alone.</p><p>It’s a chaotic time to make sense of the industry that’s reshaping our lives, whether or not we like it. Last week, another AI company employee <a href="https://www.theguardian.com/global/2026/sep/15/

On OpenRouter, weekly token consumption has surged more than 25,000 percent since January 2025, from 0.5 to 126.2 trillion tokens. The number looks impressive, but it says less about actual AI usage than about token-hungry reasoning models and a growing number of unoptimized AI agents that burn through tokens at staggering rates. The article OpenRouter's staggering token chart is the AI bubble debate in a single image appeared first on The Decoder .

ChatGPT, Character.ai, and Snapchat’s My AI will all be in the crosshairs as the AI Kids Act aims to strip minor-facing chatbots of all the things that make people love them.

A Bloomberg developer claims to have cracked an 83-year-old Enigma message from the Wehrmacht using OpenAI's GPT-6 Astra. The 82-character radio message from 1941 contains a soldier asking about his march route. The solution still needs independent review. The article OpenAI's GPT-6 Astra decrypts a Nazi radio message in ten hours that went unsolved for 83 years appeared first on The Decoder .

<p>In a new book,<strong> </strong>writers <a href="https://www.theguardian.com/profile/naomiklein">Naomi Klein</a> and <a href="https://www.theguardian.com/profile/astra-taylor">Astra Taylor</a> have created a term to define the era we are living in: End Times Fascism. A response to the climate crisis, the rise of the far-right and advancing AI technology, they spell out how the damage being wrought by the super-rich is becoming normalized. Klein and Taylor join <a href="https://www.theguardian

<p>Universities need to protect their position in education so that students have an independent pathway to employment</p><p>AI companies like OpenAI are insinuating themselves into the pathway from education to work. Soon they may claim it entirely, a disastrous result for students.</p><p>We know that students are using AI at school and at university. In conversations with those I teach, I’m struck by the trust many place in it. They turn to ChatGPT and similar tools for personal problems as we

<p>Model adopting ‘jailbreak-like instructions’ among cases as firm says it is introducing new way of tracking AI misalignment</p><p>OpenAI has disclosed six more examples of “unexpected or concerning” behaviour by its technology, as it warned the pace of development could not continue at “maximum speed for much longer”.</p><p>In one of the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and

Treble's voice simulation platform is used by voice AI model developers and AI wearable and robotics companies,
<p>In early September, OpenAI announced it had solved a major mathematics problem that has stumped humans for nearly a century. The news left mathematicians reeling, and many expressed concern over what will be left for humans as AI becomes ever more adept at unravelling complex problems. Now 25 recipients of the Fields medal – often called the Nobel prize for maths – have signed an open letter expressing their fears of a ‘severe misalignment’ between AI companies and their field. To find out ho

The Nurovi line includes three lidar-powered robovacs. But to get the model I’m most intrigued by, you’ll need a Costco membership.

Snap thinks its chunky, pricey smart glasses are the future of human computing, powered by a new AI app and features that do everything from translating languages to improving your golf swing.

The company also disclosed previously unreported incidents in which its AI models behaved in misaligned ways, including uploading files to the internet without being asked.

Planned 2029 debut could make this Apple’s first enterprise server in decades.

Even with mounting concerns about AI models going rogue, legislation appears unlikely, and the White House is outright opposed to oversight.

Ursula von der Leyen plans to invite the major frontier labs to talks and use the AI Act to help set global AI safety standards. She cited autonomous hacking and self-improving models as immediate risks. The article EU president warns AI agents "escaping their environment" are just a preview of what's coming appeared first on The Decoder .
<p>MSPs vote for strict environmental impact assessments and will delay decisions until national strategy is developed</p><p>Holyrood has voted to pause planning applications for all new AI datacentres for up to a year until the Scottish government has developed a national strategy.</p><p>On Wednesday, MSPs backed a Scottish Labour motion to halt new schemes for up to 12 months while rules on hyperscale datacentres are developed in what could be a blow to the British government’s AI strategy.</p

The Mandalorian and Grogu, The Wild Robot, and Obsession are among the must-watch films arriving this month.

According to The Information, Apple is working on an enterprise server with two or four M8 Ultra chips for the AI inference market, with a possible launch no earlier than 2029. Apple is considering Nvidia's NVLink Fusion technology to connect the chips. The project could get a boost from OpenAI and Anthropic already buying Mac hardware in bulk for AI workloads. The article Apple is reportedly building an enterprise AI server with its own M8 Ultra chips appeared first on The Decoder .
<p>Mike Johnson cancels votes on Thursday, meaning House members will avoid voting on impeaching Pete Hegseth</p><ul><li><p><a href="https://www.theguardian.com/us-news/live/2026/sep/16/donald-trump-iran-russia-midterms-north-carolina-latest-news-updates">US politics live – latest updates</a></p></li></ul><p>The Republican speaker of the House, <a href="https://www.theguardian.com/us-news/mike-johnson">Mike Johnson</a>, announced on Wednesday that he would again cancel votes scheduled for Thursd

Google Deepmind has founded the Deepmind Institute (DMI), an interdisciplinary research platform focused on AGI. Led by Demis Hassabis, Shane Legg, and James Manyika, the institute tackles questions around safety, governance, and control risks, drawing on experts from the arts, humanities, and policy alongside technologists. The article Google Deepmind launches interdisciplinary institute to tackle the big questions around AGI appeared first on The Decoder .

Claude is getting a pair of new tools today: Docs and Slides. They'll let you create documents and presentations through Claude chats, which you can export, edit, and share with other users. As part of the announcement, Anthropic is also simplifying how Claude chats work, merging regular chats and Cowork into "one Claude," with all […]

<p><strong>Lars Janssen </strong>says self-driving vehicles would be a boon for him as he has an eye condition and can’t drive, in response to an article by Adrian Chiles. Plus letters from <strong>Richard Connell </strong>and <strong>Michael Fuller</strong></p><p>Adrian Chiles (<a href="https://www.theguardian.com/commentisfree/2026/sep/10/driverless-cars-are-taking-us-on-a-road-to-nowhere">Driverless cars are taking us on a road to nowhere, 10 September</a>) wants to know who asked for driverl
