Eliezer Yudkowsky: Will superintelligent AI end the world? [video]
Article URL: https://www.ted.com/talks/eliezer_yudkowsky_will_superintelligent_ai_end_the_world Comments URL: https://news.ycombinator.com/item?id=48870062 Points: 1 # Comments: 0
Article URL: https://www.ted.com/talks/eliezer_yudkowsky_will_superintelligent_ai_end_the_world Comments URL: https://news.ycombinator.com/item?id=48870062 Points: 1 # Comments: 0
Article URL: https://www.alexdelivet.com/insights/the-end-of-zombie-startup-land Comments URL: https://news.ycombinator.com/item?id=48870025 Points: 1 # Comments: 0
Following the launch of ChatGPT Work and GPT-5.6 Sol, OpenAI has acknowledged significant issues: excessive compute usage, a confusing transition to the desktop interface for chats and projects, an unclear distinction between Codex and ChatGPT Work, and regressions in existing workflows. In some cases, GPT-5.6 Sol reportedly deleted data on its own that the user had not authorized. The article OpenAI admits it "didn't get everything quite right" with ChatGPT Work launch and scrambles to fix UX a

Close your eyes and listen. ScienceAlert stories are written, fact-checked, and edited by humans, never generated by AI. Don't miss a story, subscribe here.

Article URL: https://blog.grandimam.com/posts/how-hacker-rank-scores-engineers/ Comments URL: https://news.ycombinator.com/item?id=48869737 Points: 1 # Comments: 0
<!-- SC_OFF --><div class="md"><p>Hey! I'm looking for ways to predict human preference for a project I'm building. (imagebench.ai)</p> <p>I've tryed <strong>HPSv3</strong>, <a href="https://github.com/MizzenAI/HPSv3">https://github.com/MizzenAI/HPSv3</a> and made post about it here:</p> <p><a href="https://imagebench.ai/blog/does-the-score-match-your-eye">https://imagebench.ai/blog/does-the-score-match-your-eye</a></p> <p>It looks ok, but have many limitation as you can see in my post. </p> <p>
<pre style='white-space:pre-wrap;width:81ex'>feat: sidebar update card (web + macOS) with app-first mac update flow and Sparkle beta track (#104171) * feat: sidebar update card (web + macOS) with app-first mac update flow and Sparkle beta track Squashed from claude/update-notification-display-c6cfb9 after semantic merge with #104178 (channel-aware CLI installs). See PR #104171 body for details. * chore(i18n): resync generated inventories after rebase * chore(i18n): resync locale metadata after r
<p>I'm building ChainPay, a stablecoin checkout for WooCommerce, SaaS products, Telegram sellers, and agent workflows.</p> <p>The wallet UI is only one part of the problem. The part I keep coming back to is payment state: webhook retries, late payments, partial payments, duplicated events, and what happens when the customer closes the checkout tab.</p> <p>Here is the checklist I am using right now:</p> <ul> <li>exact payment within the expiry window</li> <li>payment after expiry</li> <li>partial
Article URL: https://www.vibhaas.net/posts/Competitive-Programming-in-the-era-of-AI/ Comments URL: https://news.ycombinator.com/item?id=48869623 Points: 1 # Comments: 0
Article URL: https://upsun.com/blog/8-stages-ai-engineering-maturity/ Comments URL: https://news.ycombinator.com/item?id=48869595 Points: 2 # Comments: 0
<h3> ...so I built <code>shortcuts</code>. </h3> <p>Every developer knows keyboard shortcuts are worth learning.<br> The problem is remembering them.</p> <p><code>shortcuts</code> does that for you.</p> <p>You forget how to split a pane in Windows Terminal, search through tmux scrollback, or jump to the end of a command. Instead of staying in your terminal, you open a browser, search the web, skim documentation, on a bad day ask some AI chatbot.</p> <p>The interruption often costs more time than
Article URL: https://www.economist.com/leaders/2025/12/04/how-ai-is-rewiring-childhood Comments URL: https://news.ycombinator.com/item?id=48869574 Points: 2 # Comments: 0
<pre style='white-space:pre-wrap;width:81ex'>fix(agents): keep exact tool allowlists on owning factories (#104213) * refactor(agents): centralize core tool factory descriptors --------- Co-authored-by: Ayaan Zaidi <hi@obviy.us></pre>
<pre style='white-space:pre-wrap;width:81ex'>fix: make LaunchAgent plists launchd-readable (#104228)</pre>
<pre style='white-space:pre-wrap;width:81ex'>feat(ui): full-page New session screen with gateway folder browser (#104238) * feat(gateway): add admin-only fs.listDir host directory listing * feat(ui): replace new-session dialog with full-page /new screen and folder browser * fix(ui): drop unnecessary template literal in new-session page * refactor(ui): rename request token locals for review-bundle hygiene * fix(ui): preserve typed draft on agent hydration and clear stale folder listings * fix(ui)
Apple is suing OpenAI over systematic employee poaching and the alleged theft of trade secrets tied to unreleased products. According to the complaint, more than 400 ex-Apple employees now work at OpenAI, including former iPhone design chief Tang Tan. The lawsuit hits OpenAI right as it's building out its own hardware division, with its first product not expected to ship until 2027 at the earliest. The article Apple sues OpenAI for allegedly running a "coordinated campaign" to steal trade secret
Article URL: https://ypipe.com/ Comments URL: https://news.ycombinator.com/item?id=48869348 Points: 1 # Comments: 0
I have been using LLMs & Coding Agent since early 2024. A large problem with Coding Agents & LLMs in general is context compression. To give you some numbers, when I analysed my own sessions across Claude Code, Codex & Sakana, I found that most of my agents spent >90% of time re reading context and upon further investigation into the markdowns it was reading, I have a hand-wavy estimate of at least ~20% of this being useless to the task at hand. When digging a bit more into this problem, I reali
Article URL: https://github.com/jaurakunal/isitsecure Comments URL: https://news.ycombinator.com/item?id=48869100 Points: 2 # Comments: 0
Article URL: https://onepromptai.app Comments URL: https://news.ycombinator.com/item?id=48868975 Points: 1 # Comments: 1
<p>Meta was criticised for feature launched on Tuesday that automatically lets users generate images using content from public Instagram accounts</p><p>Meta has said it is discontinuing an AI feature launched this week that allowed users to generate images using public Instagram accounts, after drawing widespread criticism over privacy concerns, including from a Hollywood union.</p><p>“Our intent was to provide a useful creative tool and to give people control over whether their public cont

So I decided yesterday to do a full rewrite of the media gallery, turning it into a generative media catalog with full-suite social interactions, search, and personalized recommendations built on top of the samsar-js library. It took around 50 prompts in 5 sessions to get this done across the client and API, built with Codex GPT-5.6 Sol, Luna, and then Sol again (rate limit reset, thx Mr. Sama). I did not use screenshots for the UI, It was entirely hand-prompted end-to-end. The project is fully
arXiv:2607.07759v1 Announce Type: new Abstract: Agricultural supply chains are vulnerable to disruptions through linked biophysical and economic systems. We develop an AI-powered tool that integrates economic models (GTAP) with biophysical models (APSIM) to analyze supply chain shocks, enabling policymakers and market participants to assess cross-disciplinary impacts through queries and responses written in natural language.
arXiv:2607.07761v1 Announce Type: new Abstract: Large language models (LLMs) have emerged as important tools in healthcare, showing growing potential for clinical reasoning and patient care. This survey examines recent progress in medical LLMs, focusing on reasoning applications and requirements. We present a dual-view approach that connects clinical practice with computational methods. On the clinical side, we establish a five-level competency scheme following Miller's Pyramid, progressing from
arXiv:2607.07766v1 Announce Type: new Abstract: Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires. Developers' safety responses have been largely reactive, addressing the most visible and acute harms while subtler, longer-term patterns of risk (e.g., dependency, boundary e
arXiv:2607.07775v1 Announce Type: new Abstract: The human body is at the center of a growing family of technologies designed to tightly and persistently couple biological and digital systems. Robotic prostheses are a representative example of this tight coupling. Also referred to as bionic limbs, robotic prostheses are devices that support people who have lost limbs in pursuing daily life activities such as walking and grasping objects. Bionic limbs are now perceptive and responsive owing to the
arXiv:2607.07846v1 Announce Type: new Abstract: VectorizationLLM is a specialized Large Language Model based on Google open-weight LLMs. The model is designed to assist students to learn smart vectorization, time/wave vector analysis, piecewise functions, Fourier analysis, and differential equations in MATLAB. The course application is CTEC 247: Applied Computational Analysis II by the Department of Electrical & Computer Engineering Technology at New York Institute of Technology Old Westbury. Th
arXiv:2607.07850v1 Announce Type: new Abstract: For seemless control of advanced hand prostheses and augmented reality, accurate and immediate hand gestures recognition is essential. Surface electromyography (sEMG) signals obtained from the forearm are commonly employed for this purpose. In this paper, we present a novel approach for sEMG representation that utilizes graph networks which contain information about muscle activation patterns in the forearm. Based on these graph networks, we have d
arXiv:2607.07984v1 Announce Type: new Abstract: Neural architecture search (NAS) methods have grown increasingly efficient, yet they remain bounded by manually engineered search spaces that require substantial domain expertise and must be rebuilt for every new task. Large language models (LLMs) can generate architectures in an open-ended space, but how to optimally divide the labor between LLM-driven design and NAS-driven search remains unexplored. We propose a mechanism that bridges these two p
arXiv:2607.08018v1 Announce Type: new Abstract: LLMs often struggle to balance compositionality with knowledgeability, a challenge we define as Composition-Knowledge Dichotomy. To address this, we propose Concretized Proposition Prompting (CPP), a framework that explicitly concretizes propositions relevant to questions. The results demonstrate that CPP significantly enhances reasoning performance, particularly in medical benchmarks where precise knowledge is paramount, while being competitive on
arXiv:2607.08038v1 Announce Type: new Abstract: Diagnostic error is a major threat to patient safety, yet current large language model (LLM) systems often treat diagnosis as a one-shot prediction task, lacking safeguards against missed high-risk alternatives or rigorous verification of their reasoning. Here, we present AegisDx, a safety-oriented framework for hypothetico-deductive clinical reasoning. AegisDx coordinates specialized LLM components through role-specific contracts, structured inter
arXiv:2607.08065v1 Announce Type: new Abstract: LLM-as-judge (Zheng et al., 2023) is increasingly the default for evaluating AI systems in enterprise pipelines, often scaled to ensembles (Verga et al., 2024) or "mixture-of-experts" (Shazeer et al., 2017) panels of judges. These systems share a key assumption: that consistency -- agreement among judges, or among a model's own samples -- indicates correctness. We show this assumption is unreliable. Agreement is not accuracy: a model can agree with
arXiv:2607.08079v1 Announce Type: new Abstract: Accurate photovoltaic (PV) power forecasting is essential for reliable grid dispatch and renewable energy integration, yet it remains challenging because PV generation is jointly shaped by weather variability, day-night transitions, regime-dependent dynamics, and strict physical constraints. We propose PARA-PV, a Physics-Aware Retrieval-Augmented framework that embeds physical knowledge throughout the forecasting process. The framework first encode
arXiv:2607.08136v1 Announce Type: new Abstract: We present a general neurosymbolic reasoning and learning methodology based on a modular integration of answer set programming with an energy based model substrate. Key contributions are: (1) supporting joint optimisation in the continuous latent space through explicit ASP-based declarative semantics fully incorporating background knowledge, constraints, non-monotonic inference; and (2) advancing recent works at the interface of answer sets, probab
arXiv:2607.08177v1 Announce Type: new Abstract: In this paper, we study the automatic schema generation problem: given a collection of historical ship maintenance and operational reports across multiple form categories, automatically discover compact and informative schemas that capture the essential information requirements of each report type. To address this challenge, we propose ASMR, a modular agentic framework consisting of two specialized agents. A Field Generation Agent extracts semantic
arXiv:2607.08255v1 Announce Type: new Abstract: Large language models increasingly serve as teachers generating training data for smaller students. Prior multi-teacher knowledge distillation methods merge outputs without determining which frontier model teaches best, often relying on an LLM judge biased toward its own outputs. We introduce a compete-then-collaborate framework where four frontier AI teachers (Claude, Codex-GPT, Grok, Gemini) are ranked head-to-head by an execution-based judge (un
arXiv:2607.08257v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters. We introduce $\textbf{MentalHospital}$, a virtual evaluation environment for LLM-based psychiatric clinical encounters. MentalHospital instantiates the Subjective Interviewing, Objective Examination, Diagnostic Assessment, an
arXiv:2607.08284v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated rapidly improving long-context capabilities, prompting a wave of benchmarks designed to evaluate them. However, existing long-context evaluations - from Needle-in-a-Haystack (NIAH) tests to more recent multi-hop reasoning and summarization tasks - predominantly measure average-case performance, and many are either saturated or lack robustness. Notably absent is a systematic way to probe how models perf
arXiv:2607.08285v1 Announce Type: new Abstract: Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy compliance. These measures remain essential, but they are not sufficient for systems that interact directly with users through natural language. Human-facing AI systems are increasingly used as advisors, coaches, tutors, and companions. In these roles, their responses can shape how users reason, interpret emotions
arXiv:2607.08317v1 Announce Type: new Abstract: Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a dog with five legs. These examples suggest that existing benchmarks may under-measure persistent blind spots in current systems. We introduce $\texttt{blind-spots-bench}$, a benchmark designed to expose such blind spots through tasks that appear simple for humans but rem