Giving LLM the schema docs made its wrong SQL answers more plausible, not rarer
Article URL: https://quaesitor.eu/silent-failures/ Comments URL: https://news.ycombinator.com/item?id=49360009 Points: 1 # Comments: 0
Article URL: https://quaesitor.eu/silent-failures/ Comments URL: https://news.ycombinator.com/item?id=49360009 Points: 1 # Comments: 0
The ChatGPT-maker said training will be slowed for two weeks while it puts the upgrades in place.

You can have an impressive website, great content, and a seamless user experience. But you might still be invisible to search engines and AI tools. That’s because Google and ChatGPT don’t just look at your website. They also look at what other credible sources say about your brand to understand who you are and whether … The post Brand Mentions in 2026: How to Earn and Track Them Across the Web appeared first on Backlinko .
<p>A rise in lawsuits over AI use in employment decisions is raising questions about how companies hire and fire</p><p>For the last four years, Erin Kistler has applied for thousands of jobs at companies like <a href="https://www.theguardian.com/technology/paypal">Paypal</a>, <a href="https://www.theguardian.com/technology/microsoft">Microsoft</a> and <a href="https://www.theguardian.com/media/netflix">Netflix</a>, only to find her résumé disappear into a black hole. A product manager with nearl

Article URL: https://artificialanalysis.ai/agents/search-api Comments URL: https://news.ycombinator.com/item?id=49359449 Points: 1 # Comments: 0
<p> Run AI models on any device </p> <p> <a href="https://www.producthunt.com/products/nobodywho?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1226610?app_id=339">Link</a> </p>
<p> Skip the setup and run OpenClaw & Hermes, fully managed </p> <p> <a href="https://www.producthunt.com/products/cloudways?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1226604?app_id=339">Link</a> </p>

1min.AI gives you lifetime access to dozens of AI models, including GPT, Gemini, and more

A new, scalable technique could enable powerful radars and sensors based on quantum technology that works at room temperature.

<p>A historian charts the emergence of a democracy-crushing dystopia that is – in some ways – already with us</p><p>Pulitzer prize-winning US historian Jill Lepore’s new book addresses the threat posed to liberal democracy by artificial intelligence. According to Lepore, the extraordinary power of private tech companies led by men whose priorities may not align with the wellbeing of the Earth or its inhabitants is set to destroy civilisation as we know it.</p><p>Lepore’s concept of the Artificia

<p> Give every employee a secured, sandboxed pro assistant agent </p> <p> <a href="https://www.producthunt.com/products/onecli?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1226539?app_id=339">Link</a> </p>
Replit introduces Free Mode, powered by GPT-5.6 Luna, so anyone can turn ideas into working software without worrying about token costs.
Discover how customer reviews directly influence what Large Language Models (LLMs) say about your local business in AI-powered search results and better understand how sentiment, volume, recency, and keyword relevance in reviews shape AI-generated recommendations.
<p> Find and remove every trace AI leaves in your text </p> <p> <a href="https://www.producthunt.com/products/claude-watermark?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1226478?app_id=339">Link</a> </p>
<p> Run your parallel coding agents from your phone </p> <p> <a href="https://www.producthunt.com/products/superset-5?utm_campaign=producthunt-atom-posts-feed&utm_medium=rss-feed&utm_source=producthunt-atom-posts-feed">Discussion</a> | <a href="https://www.producthunt.com/r/p/1226440?app_id=339">Link</a> </p>
Article URL: https://block.xyz/inside/designing-ai-with-character-what-we-learned-building-berd Comments URL: https://news.ycombinator.com/item?id=49357223 Points: 1 # Comments: 0
Article URL: https://slopornot.simplicated.dev Comments URL: https://news.ycombinator.com/item?id=49357213 Points: 1 # Comments: 0
Article URL: https://ai-stock-research-kiaan.streamlit.app Comments URL: https://news.ycombinator.com/item?id=49357163 Points: 1 # Comments: 1
Article URL: https://github.com/paibyun9/EGA-V9 Comments URL: https://news.ycombinator.com/item?id=49357141 Points: 1 # Comments: 2
Article URL: https://alicebraincore.web.app Comments URL: https://news.ycombinator.com/item?id=49357139 Points: 1 # Comments: 0
Article URL: https://www.tomshardware.com/tech-industry/artificial-intelligence/ai-company-hiring-a-literal-pirate-to-salvage-sunken-treasure-found-by-artificial-intelligence-pays-up-to-usd500k-a-year-mining-80-million-pages-of-spanish-colonial-records-to-find-undiscovered-wrecks-and-lost-cargo Comments URL: https://news.ycombinator.com/item?id=49356812 Points: 3 # Comments: 0
Article URL: https://github.com/YOalphabet/YOalphabet Comments URL: https://news.ycombinator.com/item?id=49356777 Points: 3 # Comments: 1
The problem: I already had AI credits, but those credits were locked to one application. context : I was using both an agentic IDE and a Hostinger deployment agent. One day, I ran out of credits on the deployment agent. To keep using it, I either had to wait for credits to reset or upgrade to a higher subscription or buy tokens. At the same time, I already had a subscription for the IDE, but I could not use those credits on Hostinger. simply despite having credits, we cannot use them. Is this ac
Article URL: https://arxiv.org/abs/2608.16834 Comments URL: https://news.ycombinator.com/item?id=49356648 Points: 1 # Comments: 0
Article URL: https://plugos.net/plugclaw/ Comments URL: https://news.ycombinator.com/item?id=49356609 Points: 3 # Comments: 0
Article URL: https://arxiv.org/abs/2608.16753 Comments URL: https://news.ycombinator.com/item?id=49356591 Points: 1 # Comments: 0
arXiv:2608.16890v1 Announce Type: new Abstract: Clinical trial programming -- transforming study protocols into analysis-ready datasets under CDISC standards -- is a bottleneck in regulatory submissions, yet LLM-based code generation fails catastrophically on this task: across 11 single-shot attempts with five frontier models, none produces a valid subject-level analysis dataset. We introduce GxP-Agent, a multi-agent system that encodes regulatory process ordering as a directed acyclic graph (DA
arXiv:2608.16891v1 Announce Type: new Abstract: Agentic AI systems request tool actions that can modify files, send messages, launch jobs, or change workflow state. This shifts the safety problem from harmful text generation to harmful operational side effects. Prompt-level governance can shape model behavior, but it does not create an execution boundary. We introduce Aegis, a runtime governance system that treats model outputs as action proposals and mediates them through a trusted decision lay
arXiv:2608.16956v1 Announce Type: new Abstract: API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omission, output rail, service product, prompt, and price schedule. We study the reasoning-effort term through a registered paired contrast of Sonnet 5 with explicit high effort against the same model with effort omitted, using 30 AIME 2026 items and five calls per item. Every paid attempt was assigned one
arXiv:2608.16977v1 Announce Type: new Abstract: AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model reasoning is a limited resource, and expert mathematical review is even more sharply constrained. Allocating these scarce resources well is therefore central to making AI-assisted mathematical discovery efficient. In most current AI-for-math workflows, human effort is concentrated at the beginning and end, in selecting suitable researc
arXiv:2608.17007v1 Announce Type: new Abstract: Agent Skills can specify procedural and resource obligations for tool use, and language models instantiate them as concrete programs. However, when models turn this guidance into code for existing tool interfaces, even a semantically correct program may load an entire input and exceed the memory available to one tool call. We present SkillEffect, a checked-lowering runtime for computations with a recoverable source relation, an audited bounded impl
arXiv:2608.17053v1 Announce Type: new Abstract: A bounded agent may obtain information for a decision from its own past, from peers, or from both sources. Retaining task-relevant history can reduce later communication, while a peer message can supply what memory lacks. Under limits on both resources, how should an agent allocate its information budget? Given a fixed task and decision rule, the memory and message rate pairs attaining a performance threshold form an achievable region under specifi
arXiv:2608.17067v1 Announce Type: new Abstract: As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks. Existing defenses predominantly operate under white-box assumptions, relying on text encoder optimization, weight editing, or inference-time intervention, and fundamentally cannot scale to proprietary models. Black-box alt
arXiv:2608.17071v1 Announce Type: new Abstract: We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-specialized agents run in parallel and coordinate through conclusions-only shared memory, a deterministic benchmark guard, and read-only cross-agent state with plateau-triggered drafting. We evaluate \kernelarc{} on NVIDIA H100 and B200 GPUs using category-representative SOL-ExecBench workloads. The resulting implementations
arXiv:2608.17124v1 Announce Type: new Abstract: Combining the answers a large language model (LLM) samples for a question into one decision is a test-time information fusion problem, usually solved by majority voting. Voting is unreliable on difficult questions, where the sampled answers share correlated errors, so the wrong answer can win and drawing more samples makes the decision worse. Selecting a candidate by reading a correctness signal from the model's hidden states is a promising alterna
arXiv:2608.17128v1 Announce Type: new Abstract: A personal AI system needs a model of the user's goals, constraints, and ongoing commitments to plan and act on their behalf, and the quality of that model is bounded by what the system can observe. Broader observation does not by itself improve assistance because a bounded system must select and compress information for the task at hand. We argue that this observation bottleneck has a cooperative structure: the system builds a partial model of the
arXiv:2608.17150v1 Announce Type: new Abstract: To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matching content to a user's evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not explicitly model user knowledge so they neither produce realistic interactions across knowledge levels nor reflect how interactions unfold as that knowledge evolves. To close this ga
arXiv:2608.17170v1 Announce Type: new Abstract: Algorithm selection for constraint satisfaction problems requires extracting features that capture problem structure. Manually designing feature extractors demands deep domain expertise and quickly becomes a bottleneck when new problem classes appear. We present an automated approach that uses Large Language Models (LLMs) in an agentic check--fix--verify loop to synthesize executable Python scripts that act as interpretable, problem-specific featur
arXiv:2608.17183v1 Announce Type: new Abstract: Small Language Models (SLMs) are increasingly deployed in resource-constrained, privacy-sensitive settings, where safety and bias failures can cause security and societal risks. However, existing AI safety\slash security\slash compliance benchmarks are designed for large language models that may not transfer reliably to SLMs. We therefore ask: Can these benchmarks effectively and reliably evaluate SLMs? To answer this question, we conduct a large-s