2026-09-26

AI Frontier Daily Briefing: 2026-09-26

Appeals court upholds the Pentagon's supply chain risk designation of Anthropic (309 pts/513 comments); Dutch government takes the day's top score at 910 upvotes with a NixOS-based replacement for its Microsoft workplace; California's billionaire tax draws 792 comments; Microsoft exits the personal AI chatbot race; the author of rr leaves Google over AI acceleration; Meta's Muse is caught routing to an OpenAI model; Claude computes a nine-loop scattering amplitude for about $1,000; Anthropic puts Claude to work on Ebola sitreps; a CMU professor redesigns a course around AI doing the homework; the plan-mode debate; an agentic CUDA-kernel optimizer; Docker cloud sandboxes, portable SIMD in Go, the Rails World keynote fight, Topcoat v0.9, git-bug, Typst 0.15; Google's orbital TPUs, Oracle's data-centre contract mess, ASML's zero European orders, the Avast sandbox break.

88 stories made the HN front page on 2026-09-25 (UTC), 23 of them with more comments than upvotes. The lead today is the appeals court upholding the Pentagon’s supply-chain-risk designation of Anthropic; the Dutch government took the day’s top score at 910 upvotes by building its own NixOS-based alternative to the Microsoft workplace; Microsoft walked away from the personal AI chatbot race while the author of rr left Google rather than speed up AI chips; and the tools section is strong, with portable SIMD in Go, the Rails World keynote fallout, and Claude’s nine-loop physics computation. This update adds Anthropic’s Ebola-response deployment, a CMU course rebuilt around AI doing the homework, the plan-mode debate, an agentic CUDA-kernel optimizer, a skeptical read on Jev’s calibration claims, and the VS Code border-radius affair. 30 items.

1. Appeals court keeps Anthropic’s “supply chain risk” label

Safety guardrails count as sabotage now?

The U.S. Court of Appeals for the D.C. Circuit upheld the Pentagon’s designation of Anthropic as a supply chain risk under the Federal Acquisition Supply Chain Security Act, 309 pts and 513 comments. Background: Anthropic’s DoD contract restricted Claude’s use: no mass domestic surveillance, human oversight in targeting decisions. The department demanded those terms come out, talks collapsed, and Anthropic landed on the risk list, barring defense contractors from its products. Anthropic sued, calling the move retaliation and saying the statute targets “adversaries.” The majority read the definition as covering “any person” and held that “the Department reasonably feared that Anthropic might manipulate Claude’s design”; the dissent argued “manipulate” must mean malicious subversion, not product guardrails. It is the first use of the designation against a domestic company. If you sell AI into government, this opinion decides whether safety terms survive procurement.Source · HN discussion

2. Dutch government builds its own Microsoft replacement, 910 upvotes

Roll your own workplace suite?

DAWO (Digitale Autonome Werkplek Omgeving, digital autonomous workplace) is the Dutch government’s effort to replace the Microsoft workplace environment with a component-based stack, 910 pts and 531 comments, the day’s top score. The OS layer is built on NixOS, with AI, cloud, and collaboration components alongside; code lives on code.overheid.nl and Codeberg, developed jointly by government, industry, and community. According to itsFOSS, the trigger was Microsoft cutting off the International Criminal Court’s access to its workplace software. No rollout size or timeline has been published. If you run enterprise IT procurement or work on open-source desktops, this national-scale NixOS deployment is worth tracking.Source · HN discussion

3. California’s billionaire tax chases wealth that already left

They moved. So who’s left to tax?

California’s billionaire wealth tax has qualified for this November’s ballot, 276 pts and 792 comments, the day’s comment count leader. A Center for Land Economics estimate puts the state’s land at about $8.14 trillion, roughly 8x the taxable billionaire wealth, and says a 0.25% land value tax would raise the same $20 billion a year the wealth tax promises. The authors argue the wealth-tax base is overestimated by nearly 2x: six billionaires (about $540 billion combined, including Page, Brin, and Thiel) left before the January 1, 2026 cutoff, and Zuckerberg (about $220 billion) moved afterward and is expected to challenge retroactivity. Shrinking the base pushes the rate toward 1.6-1.9%, which accelerates further flight. The root cause is Proposition 13 (1978), which froze property assessments. If you do US site selection or investing, this rate trajectory feeds straight into cost models.Source · HN discussion

4. Microsoft walks away from the personal AI chatbot race

30 million subscribers couldn’t pay for one chat window?

Bloomberg reports Microsoft is dropping the head-on competition with frontier labs for personal AI chatbots and rebooting the Copilot line instead, 69 pts and 59 comments. The numbers: more than 30 million paid Copilot subscriptions as of end of June, and about 90 million paying users on the M365 bundle it is strongest in. The HN experience reports are rougher: multiple users say enterprise Copilot truncates chat history heavily and crashes mid-research-task, and one found a cheaper AI-free 365 tier that only appears behind the cancel button. If you price AI products, this is a bundle-first model running out of road: users pay for the suite, not the assistant.Source · HN discussion

5. The author of rr is leaving Google over AI acceleration

Doesn’t want to keep stoking this fire?

Robert O’Callahan, former Mozilla engineer, creator of the rr debugger, announced he is leaving Google, 243 pts and 300 comments. Working remotely from New Zealand on Google’s chip-design tooling, he concluded the work’s main impact would be faster, cheaper AI hardware, and that the current pace is “far too high.” Staying at Google outside AI acceleration was impractical: New Zealand has no other relevant team. He says the AI labs’ warnings are sincere, frames the departure as a moral choice, and wants his next work to be “unambiguously pro-human.” Among the recent run of public AI resignations, this is the quietest: no accusation, just a list of things he no longer wants to help with.Source · HN discussion

6. Meta’s Muse appears to route sessions to an OpenAI model

Thought it was all in-house?

Blogger Pete James dug through Muse’s runtime logs inside a VM and found that, alongside Meta’s own Avocado model, one subagent session was routed to a model internally labeled azure/muse-special, 70 pts and 32 comments. The evidence: repo code referencing a “GPT Responses model client via MAGI native Azure OpenAI lane,” a gpt_responses_v1 session signature, and tool-call IDs in OpenAI’s call_ plus 24-character format, versus 32 hex characters in Avocado sessions. He rules out distillation: the OpenAI-labeled model’s raw reasoning is encrypted and only sent back to Azure, where Meta’s servers can’t read it. Muse’s shipped catalog also lists Claude Opus 4.6/4.7/4.8, GPT-5.5/5.6, and Kimi K3 among some 15 Avocado versions. Routing is a server-side decision, and the logs don’t show why this model was picked. If you track model supply chains, this is rare first-hand material on how a big lab routes across models.Source · HN discussion

7. OpenAI “attacked Australia’s health system and doubled down”

Are they even going to acknowledge it?

A Guardian opinion piece. “OpenAI attacked Australia’s health system and doubled down”, argues the wait-and-see period is over, 12 pts. It follows the previous day’s revelation that OpenAI’s agent scraped Australia’s Medicare statistics portal at scale (covered in yesterday’s briefing), which drew a public response from Prime Minister Albanese. The op-ed demands regulatory action. The HN thread holds a single, pessimistic comment: OpenAI will face no real consequences, and there is a risk this capability ends up in state-sponsored cyber operations. If you work on AI compliance, the same incident hitting HN two days running means scraping authorization on public data portals has moved from PR problem to regulatory problem.Source · HN discussion

8. Nine loops, about $1,000: Claude takes a frontier physics computation

Are the physicists sleeping tonight?

Science writer Matt von Hippel challenged AI labs on August 7 to compute the nine-loop scattering amplitude in N=4 super Yang-Mills on an academic budget, 100 pts and 52 comments. Anthropic researchers Liam Fitzpatrick and Siddharth Mishra-Sharma took it with Fable 5.1 running in Claude Science, reaching the result via two independent routes (a direct bootstrap and an indirect form-factor path), at roughly $1,000-2,000 end to end; the bootstrap leg alone cost about $100, matching 96 CPUs running for a week. Lance Dixon of SLAC/Stanford validated it independently on September 1, having expected a direct computation to be too fragile: “if you make any mistake at all, it all crashes down.” Song He’s group at the Chinese Academy of Sciences concurrently reached most of the result with GPT-6 assisting on constraints. The previous ceiling for such calculations was five loops, in the electron magnetic-moment prediction. If your research is compute-heavy, remember this budget figure: one frontier calculation now costs about a laptop.Source · HN discussion

9. OpenAI ships a mental-health benchmark, commenters split

Who do you talk to when no psychiatrist will take you?

OpenAI released MentalHealthBench, evaluating how language models respond in mental-health contexts, 31 pts and 14 comments. The thread splits three ways. Demand side: one commenter, whose parents are clinical psychologists of 40 years, notes waits of months for psychiatrists in some regions and argues LLMs could serve the unserved. Risk side: therapy is 20-plus sessions steered by a therapist, while an LLM is steered by previous tokens; a delusional patient could talk the model into agreeing. Accountability side: if a model’s out-of-standard response causes harm, who carries the liability a license would? If you build companion or health apps, read this benchmark before you ship into this space.Source · HN discussion

10. The Optimus ramp, hands still assembled by human workers

Stuck at the fingers?

The Information reports Optimus output rose from a few dozen units a week in Q2 to several hundred by August, targeting an automated line above 1,000 robots a week by year-end and eventually around 20,000, 10 pts and 22 comments. Three bottlenecks: each hand and forearm has over 100 screws and small parts still assembled by hand, and the fixtures can’t hold tolerances tighter than car components; touch sensors fail reliability tests, with a replaceable “sensing glove” fix planned for next year; and motors and precision gears depend on Chinese suppliers whose prototype parts pass but whose volume consistency wobbles. On AI: Optimus can’t yet handle a wide range of tasks, takes days to learn basic ones, and has 500,000+ hours of training data. For contrast, XPeng started running an automated IRON humanoid line in September, targeting 2027 commercial sales. If you evaluate humanoid supply chains, hand and joint supplier yield matters more than keynote specs.Source · HN discussion

11. From 236ms to 8ms — run Jev-style decision models locally

Worth self-hosting models this small?

Ollaya is an open-source (Apache-2.0) local runtime for decision models, API-compatible with TypeSafe’s Jev, 201 pts and 56 comments. Decision models skip token-by-token generation and return structured answers (choice, score, yes/no) plus probabilities in a single forward pass. Measured: the Laya model answers five questions end-to-end in about 8-10 ms on an RTX 4090, versus Jev’s hosted median of 236-276 ms including network; calibration (ECE after temperature fitting) is 0.081 versus Jev’s 0.246. Weights are pulled from the authors’ own Hugging Face repos, pinned to a commit and checked against sha256, never re-hosted. Runs on macOS, Windows, Linux, and Docker; NVIDIA GPUs need driver R580+. If you run classification or moderation pipelines, this is another option for swapping per-token fees for a local resident process.Source · HN discussion

12. Write once, SIMD everywhere, Go ships a portable vector API

Is the hand-written assembly moat filling in?

The Go blog published an experimental portable SIMD package, 325 pts and 123 comments. The design follows Google’s Highway: vector length is determined at runtime, only operations in the intersection of supported platforms are exposed, gaps are filled with efficient emulation, and code still runs fully emulated on hardware without SIMD. Timeline: Go 1.26 shipped archsimd for amd64, 1.27 adds arm64, wasm, and the portable simd package, and 1.28 plans SVE plus more operations; builds need GOEXPERIMENT=simd. The implementation rewrites ASTs in the compiler frontend and dispatches across 128/256/512-bit and emulation paths; GODEBUG=simd=0 forces the emulated path for testing. Authors David Chase and Junyang Shao note Go’s Green Tea garbage collector already uses SIMD for heap scanning, with targets from cryptography to AI. If you write Go with vector math, 1.27 is an upgrade worth taking.Source · HN discussion

13. The loudest Rails World news? DHH is moving Hey off Rails

The framework’s leader says stop writing code?

At the Rails World 2026 keynote, DHH announced that Hey, a flagship Rails app, is being rebuilt on a Rust backend plus six native apps, while offering no real vision for Rails itself, 287 pts and 184 comments. He says he wrote only 3% Ruby this year (versus roughly half historically), produced 150k lines of code in August, predicts hand-written code ends this year for virtually all programmers, and claims Hey Next uses 99% less CPU with peak traffic fitting on a Raspberry Pi. Critic Jared Norman checks each claim: the output is verbose Rust he never read; the performance gains conflate switching to Rust with dropping the web frontend; and “never look at the code” sits fifteen minutes away from “security, something’s coming.” DHH also repackaged convention over configuration as “token efficiency.” The rubyonrails.org site put up an /ai page the same day. If you make framework or AI-coding-workflow decisions, two numbers carry more information than the positions: 3% hand-written code, and 150k lines nobody read.Source · HN discussion

14. Topcoat v0.9. Rails’ spirit, rebuilt in Rust by the Tokio team

Conventions work on LLMs too?

The Tokio team released Topcoat v0.9, a batteries-included full-stack Rust framework, 108 pts and 92 comments, two months after its July reveal. Server-side rendering is the default; browser interactivity uses signals and a type-checked Rust subset compiled to JavaScript, avoiding server round-trips. Components called shards re-render on the server when browser state changes, live! and emit! macros stream UI updates, and v0.9 adds WebSocket server-push. The bundled Toasty ORM gains an update! macro, PostgreSQL JSONB document fields, and polymorphic relations through enums. Lead Carl Lerche is a former Rails core member, and his bet is explicit: well-defined conventions let LLMs work with fewer tokens and fewer errors, an argument worth reading against the backdrop of the Rails keynote.Source · HN discussion

15. git-bug: a bug tracker that lives inside your Git repo

Issues that travel with push?

git-bug is a distributed, offline-first bug tracker embedded in Git, 266 pts and 90 comments, 10.4k stars, GPLv3, by Michael Muré. Bugs are stored as DAG-modeled entities in the repository: nothing leaves your machine until you push; no files are added to your project, and listing or opening bugs takes milliseconds. Three interfaces: CLI, an interactive terminal UI, and a Web UI on a GraphQL API; bridges import and export with GitHub, GitLab, Jira, and Launchpad. The roadmap includes pull-request support and an identity rework using did:plc, and the on-disk format has a public spec. If your team is tired of issues being locked to a single platform, this is a credible disaster-recovery option.Source · HN discussion

16. Docker ships cloud sandboxes for agentic workloads

A fence for YOLO mode?

Docker released Cloud Sandboxes, the cloud counterpart to local Docker Sandboxes (sbx), 27 pts and 7 comments. The pitch: move one sandbox between laptop and cloud with a single command, with microVM-grade isolation, aimed at coding agents that run autonomously (YOLO mode included), plus secrets handling and definable controls. It connects to a broader line covering reproducible agent evaluation (fixed execution, structured artifacts, runtime evidence) and an MCP Enterprise Gateway on the enterprise side. If you run Claude Code-style tools and want to loosen the approval loop, start in a local sbx and move the same sandbox to the cloud unchanged.Source · HN discussion

17. Typst 0.15 takes big strides toward replacing LaTeX

Still LaTeX-only for submissions?

LWN covers the typesetting system Typst’s recent progress, 80 pts and 11 comments. Version 0.15 (June) adds variable-font axis control, MathML output for browser-rendered math, a bundle mode emitting multiple linked outputs such as HTML plus PDF from one source, per-chapter bibliographies, and compilation toward archival and accessibility PDF standards. Written in Rust, Apache-2.0, with about 460 compiler contributors, an ecosystem of 1,500+ packages, and documentation that ships as a 26MB PDF; HTML and bundle modes need feature flags. The ceiling is network effects: most venues that demand sources still accept only LaTeX or Word. If you write papers and templates, migrate new documents first and try the bundle mode.Source · HN discussion

18. A Pentium II pushed to 600MHz, emulated on the M6 Mac Mini

How fast is a 25-year-old machine today?

Reviewer Lily ran a patched build of 86Box 6.0 (released May 31, 2026) on an M6 Mac Mini (12-core CPU, 24GB RAM), emulating a Pentium II Deschutes with a Voodoo 3, 255 pts and 109 comments. The pass bar was strict: Cinebench 2000 plus Winamp playing a WAV, with a flat 100% emulation speed and zero audio dropouts. The M6 held a stable 600MHz, the M4 topped out at 500MHz, a 20% higher ceiling; 650MHz failed on one or two audio underruns. At 600MHz, 3DMark 2000 SE ran at full speed for 7-8 minutes and the CB2000 score hit 9.28, oddly above period user reports for real Pentium IIs, a discrepancy the author flags without fully explaining it. Telemetry showed 86Box on two performance cores at about 4.7GHz, the whole package under 26%. If you follow Apple Silicon single-core performance or retro emulation, her patch is available and the test is reproducible.Source · HN discussion

19. Google puts TPUs in orbit — first Suncatcher test launches October 1

Did they crack cooling in vacuum?

Ars Technica reports that MVP, the first test satellite in Google’s Suncatcher project, is scheduled to launch on October 1, 43 pts and 69 comments. Refrigerator-sized, it carries four of Google’s custom TPU accelerators to test whether they survive launch stress plus radiation and thermal extremes. The stated limit is candid: cooling can only run in bursts of about 15 minutes, after which the TPUs must shut down so the radiators can catch up. Commenters split into skeptics (no convection in vacuum: heat leaves only by radiation, dead hardware is unrecoverable, economics fail) and defenders who note this is a small research bet that Starship-class cheap heavy launch could change. Others suspect regulatory arbitrage: Texas has no jurisdiction in low Earth orbit. If you plan data-center cooling and power, that 15-minute window is the first hard number of this whole route.Source · HN discussion

20. Oracle must pay data-centre investors even with no power

Time to do the AI-infrastructure math?

The FT reports Oracle issued a force majeure notice over power supply for its Jupiter data-center project in New Mexico, while remaining on the hook to investors, 148 pts and 139 comments. The project is reportedly a 25GW site whose natural-gas pipeline state regulators rejected; Oracle signed a 1.8GW solid-oxide fuel-cell deal with Bloom Energy this year, expanded to 2.8GW in April. The dispute: Oracle claims the notices “did not establish a project delay,” which commenters call legally shaky: data-center leases typically carve permitting out of force majeure, and no known Bloom installation is within two orders of magnitude of this scale. Others note Oracle’s debt has rated just above junk since July. If you track AI infrastructure financing, this is the first crack in the off-balance-sheet-leverage plus green-power story.Source · HN discussion

21. ASML sold “absolutely nothing” in Europe in 2026

Europe builds the best machine but won’t buy it?

Tom’s Hardware reports ASML’s CEO saying the company sold “absolutely nothing” in Europe in 2026, 66 pts and 117 comments, and is calling on the EU to help create demand. The backdrop is Europe’s shrinking share of chip manufacturing and the absence of a domestic advanced-node customer; ASML’s most advanced EUV tools are bought by TSMC, Samsung, Intel, and other Asian and US firms. Two lines dominate the thread: energy prices and fab costs make Europe unattractive, and the “European chip sovereignty” narrative looks hollow when the lithography champion has to ask Brussels for policy, not wait for the market. If you watch the geography of the semiconductor supply chain, the weight of this item is ASML’s own posture.Source · HN discussion

22. Data centres now targets, Zelensky says Russia widened its strikes

Do server rooms join the air-defense list?

The BBC reports Zelensky saying Russia has expanded its attacks to hit Ukraine’s data centres, 82 pts and 91 comments. The report names no specific facilities or damage figures. The most quoted line in the thread states the position plainly: “data centers are now dual use infrastructure and valid war targets.” On retaliation, commenters split between hitting Russian data centres in kind and warnings from WWII-era target-switching mistakes. The technical consensus is clear either way: data centres depend on power and cooling that are large, hard to shield, and slow to repair. If you do infrastructure disaster recovery, this breaks the default assumption that data centres are civilian facilities; multi-region redundancy moves from cost line to safety line.Source · HN discussion

23. Avast’s sandbox fully broken, double-fetch to SYSTEM

The machine with antivirus is easier to escalate?

The SAFA Team published part two of its CVE-2025-13032 analysis, 107 pts and 27 comments. The bug is a double-fetch in Avast’s kernel driver aswSnx: the Length field of a UNICODE_STRING is read twice: once to size the paged-pool allocation, once to size the memmove, and a racing thread flips the value between the two reads, producing a pool overflow. The chain: overflow the IORing object’s RegBuffers array for arbitrary kernel read/write (Windows has no SMAP), leak a kernel address via the MDL’s Process field, repair the obfuscated ProcessBilled header to avoid a BSOD, then walk the _EPROCESS list and swap tokens to reach SYSTEM on a then-current Windows 11. The vulnerability is patched, and newer kernels validate user-memory access from user mode, which kills this technique. If you do Windows security research or driver work, the heap-spray and cleanup craft in this writeup rewards a close read.Source · HN discussion

24. A 94% AI score gets a French prize author removed

Does the detector get the final say?

The BBC reports that Haitian-Canadian writer Thelyson Orelien was accused by an anonymous account, “Balance ton Claude” (a play on France’s #MeToo slogan “Balance ton porc”), of writing his novel almost entirely with AI, 34 pts and 80 comments. Radio-Canada ran the whole book through Pangram’s detector and reported a 94% AI score, while other Goncourt finalists and French classics scored as human-written. Reporters also found older plagiarism: a 2014 award-winning work was almost entirely lifted from a 1960s French novel, plus lifted blog posts and articles; his pre-2023 writing scores low on Pangram, his newer work scores high. The outcome: he was removed from the Goncourt shortlist, and the account’s operator was identified as Le Figaro journalist Samuel Fitoussi. The thread disputes detector reliability: one demo shows the same passage flipping between 100% AI and 100% human depending on context, while others counter that long texts are more stable and the plagiarism history shrinks the benefit of the doubt. If you use AI detection for moderation, this case is worth a post-mortem on method, text length, and control groups.Source · HN discussion

25. 7,672 Ebola cases, and Claude cuts the daily sitrep to an hour

A benchmark with real patients on the line?

The Bundibugyo ebolavirus outbreak (BDBV, a strain with no confirmed vaccine) in eastern DRC was declared an international emergency in May; as of September 19 it totals 7,672 confirmed cases and 3,699 deaths, a 48.2% fatality rate, 28 pts and 12 comments. Anthropic’s deployment runs on three fronts: 1) WHO AFRO staff built a Claude skill that pulls case and lab figures out of district PowerPoint decks, checks them against the previous report, and flags trend changes, cutting the situational report from a full day to under an hour, and enabling multiple forecasting models to run at once. 2) With CEPI, Claude organizes multi-factor vaccine-proposal data so experts can compare candidates side by side. 3) The DRC’s national biomedical institute (INRB) uses Claude Science to assemble viral genomes and build phylogenetic trees from plain-language prompts instead of command-line bioinformatics. Muyembe, co-discoverer of Ebola in 1976, frames it: you beat Ebola by knowing where it is today, not where it was last week. If you build AI-for-science or public-health data systems, this is a case study with the division of labor stated plainly: the model organizes data, judgment stays human.Source · HN discussion

26. AI did all the homework, so a CMU professor rebuilt the course

Oral exams make a comeback?

Christian Kästner, associate professor at CMU, published the redesign checklist for his Machine Learning in Production course, 10 pts. The course has 100-170 students, and he lets them “use AI in any form, without attribution”. Then he changed six things: 1) written reflections are gone, replaced by a 15-minute TA meeting after each assignment, worth about 20% of assignment points, pass/fail with penalty-free retries. 2) Reading quizzes halved and ungraded, folded into class discussion. 3) Exam weight raised from 15% to 25%. 4) A 10% resubmission tax, after students started handing in AI output first and engaging only with grader feedback. 5) An LLM-as-a-judge pre-screen of submissions (pass or needs-review) cuts TA grading time by 50-80%, with humans still deciding deductions. 6) The old 12k-LOC Instagram-clone assignment was fully solved by Claude Code in fall 2025, so it was swapped for a Zulip task spanning over 500k LOC. He concedes this cuts against the evidence for frequent low-stakes assessment, and that every countermeasure decays as models improve. If you teach an AI/ML course, this checklist is copyable.Source · HN discussion

27. Plan mode is dead, says a former GitHub engineer

You’d let an agent run without a plan?

Ayman Nadeem, a former GitHub senior engineer and builder of the desktop coding app Nuanced, argues in a September 24 post that plan mode is obsolete, 8 pts and 4 comments. His claim: the plan-approve-execute workflow served two purposes: precise instructions for the agent, and a readable plan for the human. The first is disappearing as models improve, since “every decision the model can reliably make on its own is one fewer decision that needs to be surfaced”; the second still matters, but the artifact shouldn’t be a long AI-generated document. Nuanced failed on exactly this: unreadable spec walls, plus a waterfall split between thinking and building that real work doesn’t follow. He points to Codex’s newer loop as the contrast: understand, act, inspect, clarify, adjust, act again. The unsolved problem sits at the end: as agent counts grow “from five to hundreds,” how do humans keep a working model of the system? If you design agent workflows, the value here is not the verdict but the redefinition: a plan as a process for human understanding, not a deliverable.Source · HN discussion

28. An agent loop writes, benchmarks, and refines CUDA kernels

Kernel tuning by LLM, then?

bertaye/agentic-cuda-optimizer is an experimental project, 31 pts and 10 comments. The pipeline is a LangGraph agent loop: 1) propose changes to kernel code and launch configuration; 2) a standalone C++ harness compiles with NVRTC and runs via the CUDA Driver API; 3) NumPy verifies output correctness, and every case must pass; 4) candidates are ranked by geometric-mean latency across performance cases, and the fastest validated kernel is kept, with history and heatmaps. The default model is gpt-5-mini, the example CLI uses —max-iterations 7, and timing uses 10 warmup plus 100 measured launches. Developed on a Windows RTX 3060 laptop; requires Python 3.12+ and CMake 3.24+. If you write GPU kernels, this is a minimal reference for handing the whole edit-compile-benchmark-refine cycle to an agent.Source · HN discussion

29. F1 0.958 can still be useless; calibration beats accuracy

Whose calibration numbers do you trust?

Kartik Pansuriya takes apart TypeSafe’s Jev, billed as the first “System One model,” in a long post, 13 pts and 11 comments. Jev skips token-by-token generation and returns an option plus a probability in one forward pass: 70-500 ms end to end, marketed as 40x-200x faster than frontier LLMs, $0.042 per million input tokens with free output, up to 255 options. His argument: the real selling point is not speed but calibration: the training objective (RLCD, reinforcement learning for calibrated decisions) rewards honesty rather than human preference, while most production classifiers are miscalibrated and the usual fixes (Platt, isotonic, temperature scaling) are extra components that drift. On his own COMPSAC paper-acceptance task, a random forest scores F1 0.958 versus 0.957 for the majority-class baseline, while ROC-AUC is 0.676 versus 0.500, so accuracy-style metrics carry almost no signal on imbalanced data. His evaluation list: Brier score, ECE, reliability diagrams, latency, and cost per thousand calls, with gradient-boosted trees as a mandatory baseline. If you run classification or moderation pipelines, the methodology — measure on your own data, ignore vendor numbers — travels directly.Source · HN discussion

30. Border radius has infected the VS Code editor

Rounded corners, in my terminal?

User kidCaulfield filed issue #338035: since VS Code 1.139.0, the text editor, file explorer, terminal, and Copilot chat all render with rounded corners, reproducible with every extension disabled (macOS 26.5.1), 50 pts and 24 comments. The report is half tongue-in-cheek (“your editor is now infected with border radius,” “I can no longer work efficiently with these redundant curves attracting my attention”), and the ask is that desktop apps stop importing web UI fashion. The issue is closed as a duplicate and merged into a main tracking issue, with no fix committed. If you build desktop UI, this 50-upvote joke thread is a free user-preference survey: style migrations like this one deserve an explicit toggle.Source · HN discussion