Thursday, 1 October
Biological risk has a useful reality check: AI substantially helped novices on written biology problems, but in a preregistered physical-lab trial only 4/77 AI-assisted versus 5/76 internet-only novices completed the core task sequence. This does not prove the lab barrier is permanent—the trial was underpowered, and equipped experts are a different user. It shows why answer quality cannot simply be multiplied into outbreak risk. Which user and physical handoff should we test next?
The core task sequence was completed by 4 of 77 participants in the AI access condition, compared to 5 of 76 participants in the internet access condition.
Against your view
A stress test for “agents close routine requests; people see only exceptions”: Taobao randomized 647 support workers. Only 5.8% of chats were AI-eligible; these were 16.8% shorter, but customer ratings fell 0.412/5, while seven-day same-issue recontacts did not significantly change. In a matched subset of agent-handled chats, 65% escalated; emotional escalations had six percentage points more recontacts than comparable human-only chats. The exception boundary—and how soon it fires—is part of the product, not a footnote. This is 2024 customer support, not a test of software-team requests.
Human intervention preserves service quality in algorithm-triggered technical escalations ... but is less effective in algorithm-triggered emotional escalations
A small manager-practice detail: Anthropic’s Claude Code engineering director says new managers start by shipping as individual contributors, so they experience the agentic workflow they’ll help teams change. Pods then choose their own triage and planning rituals. An operating example for Dru’s manager-first idea—not evidence that it improves outcomes.
When I joined Claude Code I wanted every manager to start out as an IC first
A short calibration for Amodei’s self-improvement worry: in September, METR found AI helping AI R&D but judged its tested model unlikely to automate research end to end. Real acceleration, not a demonstrated runaway loop; the missing test is whether gains compound across generations after accounting for humans and compute.
We believe that the development of this model was at least somewhat accelerated by AI but is unlikely to have been dramatically accelerated by AI.
Microsoft’s 2026 agent rollout: +24% merged PRs after Claude Code/Copilot CLI adoption, versus a modeled non-adopter trajectory. But this is observational, and the authors call PRs a proxy for output—not delivered value. What would the number look like after counting reviews, regressions and agent spend?
a merged PR is not the same as the value it delivers
That's everything for now. More arrives as you react.