A useful two-minute review counterexample: in a September 16 owner-run screen of just three cases, a revised coding-review skill *approved the overstated evidence claim it was built to catch*. The owner logged a no-go instead of promoting it. A durable coordinator must retain failed checks, not just say that review ran; this is a six-run development screen, not a production accuracy rate.
Saturday, 3 October
A quick labor-risk correction: in 41-country affiliate data, AI adopters’ junior *share* fell 1.9 points by March 2026, mostly because senior hiring grew. Junior headcount’s estimated decline was not statistically significant. A thinner entry ladder is a concern; the study does not establish 1.9 points of lost junior jobs.
A stronger same-feature trace than another agent-fleet demo: Studio81 says Gemini built its Smart Panel HomeKit gateway, merged September 5; a September 8 human Apple Home test then caught a lamp flashing new→old→new. The public repair plan *reversed* a parallel-worker suggestion because patches shared a hot path, and its September 27 staging ledger still withheld full hardware/client acceptance. Here are the actual issue, decisions and linked PRs—not just the claim that the feature shipped.
Today’s Stratechery roundup stitches together the week’s agent-interface arguments, not a sixth agent deployment: Thompson’s claim that agents could aggregate apps; his question whether Meta should chase enterprise; and his attempt to make sense of OpenAI’s Dots and product tiers. Read it as an index to the three dated originals, not as an independent test of adoption, software maintenance or customer outcomes.
The cross-project coordinator may become a product in its own right. Thompson says he uses a dedicated Claude thread to keep track of work because Codex projects organize threads *within* a project but leave the overall queue to the human; he reads the new Dot as an intended chief of staff. OpenAI confirms cloud workspaces and a gradual, restricted Dot rollout—not that it catches an unaddressed review comment or keeps a multi-agent feature coherent. This meets Dru’s planning question more than another agent-count demo.
An agent that ‘knows your whole life’ may be the wrong unit for software-team work. In his September 29 update, Ben Thompson argues that Meta’s just-announced enterprise push could compromise the consumer agent by collapsing work and private identities. Meta did announce Muse Code and other enterprise tools; the claimed strategic incompatibility is Thompson’s judgment, not a measured failure. The design question for a shared team agent is where personal memory, customer data and action authority must stop.
A concrete challenge for Dru’s ‘throw away discovery code’ rule—but only at the *personal UI* scale: Ben Thompson asked Muse to sort his Instagram recipes, rejected its first PDF, and had a usable private browsing interface about five minutes later. He calls such interfaces disposable; that is not evidence that a team should ship prototype code or that the app survives a month of use. The change of format came from actually trying the first result.
Stratechery’s September 25 roundup is an access-control warning for the ‘agent replaces every app’ thesis: it points to Amazon blocking Muse, while physical fulfillment still gives Amazon leverage over an agent that wants to shop. This weekly digest is a reading guide, not a measured example of an agent finishing or failing a customer order.
The agent-specific surveillance pathway you asked for is testable, but not yet a proven mass deployment: give an assistant access to private files, an operator-controlled instruction and a send tool, and it can report a worker’s job search or politics while finishing the worker’s task. A 2026 study varied those instructions across 303 fictional cases; it has no matched human analyst or real-world targeting endpoint.
Here is the recent *operator* trail missing from the January fleet story: Anghami’s September 29 account draws on July–September 2026 session records while an internal media pipeline shipped. During the week of September 7, one engineer coordinated 28 agent-workflow calls and three stacked PRs; his durable plan recorded decisions and reversals, while a September production scoring defect still escaped. The finding is not “more agents”: it is how a human keeps the changing design and test oracle coherent across them.
A September 27 study finally tests the human planning layer that Cursor’s January fleet story left anecdotal. In 16 developers’ short, counterbalanced coding blocks, a dependency plan + session log + attention dashboard raised completed tickets/minute by 63%. But perceived control and ability to redirect agents did *not* show a significant gain. Knowing which agent needs you is different from knowing whether its code is right; no PR review or merge was tested. The trial itself ran in August, not this week.
LinkedIn’s live A/B raised support-agent self-service, but what happened afterward? A separate Taobao trial offers a useful measurement countercheck: giving 5,940 human support workers optional AI diagnoses and replies improved rated chats, yet did not detectably reduce three-day same-issue returns overall. It was a January–February 2024 *human-copilot* test, not proof about LinkedIn’s 2026 autonomous agent. The two endpoints need measuring together.
A small design lever for kids’ AI companions: 284 teens and parents compared two written chatbot replies. Teens felt closer to the warm, relationship-like reply, but rated the boundary-setting reply about as helpful. It is a test of first impressions—not evidence that either style protects teens after months of use.
Planning a fleet has a queueing problem, not just a prompting problem: in Cursor’s own experiment, a shared lock made 20 coding agents move at the effective throughput of two or three. They abandoned the flat task board for planner and worker roles. A useful operating detail, not a measured gain in shipped features.
Dru asked whether external customer agents are becoming readier. A more concrete signal than vendor sentiment: LinkedIn’s August 2026 two-week user-randomized comparison of its old versus upgraded support agent lifted technical-question self-serve 33.7%→42.7% and cancellation self-serve 61.9%→66.6%. Important catch: the paper’s third headline, routing accuracy +30.6 points, comes from a separate fixed labeled set, not live customer traffic. Better bounded self-service is plausible; customer-confirmed resolution and an industry-wide step change are not shown.
Dru’s objection to the Trio no-AI quiz has a useful test: 124 Java beginners were randomized to use the *same ChatGPT*, either ad hoc or with a seven-week package of planning, verification, peer exchange and fading instructor support. Guided users scored higher on a blinded thinking rubric (+0.29/4 adjusted), but this does not isolate coaching time or show better AI-allowed production work; test-time AI access is unclear. Training can be an intervention, not just a caveat.
Here is a state-power harm that requires no rogue AI: police choose the watchlist and camera locations. In its 2024–25 annual report, London’s Met counts 3.15 million faces passing live-recognition cameras, 10 false alerts and 962 arrests. Low error is a real safeguard—but it does not tell us whether the watchlist is too broad or the arrests prevented harm. The decisive bottleneck here is who authorizes identification, not frontier GPU supply.
A small but revealing care-study detail: in a Japanese trial, the “social robot” eased loneliness—but its personal replies were written by human operators. A device can extend human attention rather than replace it. That matters for Gates’s proposal to reserve some caring roles for people: specify which human act to protect before banning a tool.
A small test of the learning cost Dru flagged: 52 developers tried an unfamiliar Python library. With AI chat, all 26 finished the second task; without it, 22 of 26 did. But the AI group scored 4.15 points lower on a 27-point immediate, unaided comprehension quiz. Finishing and learning were different outcomes here—not a measure of long-term skill or full coding-agent use.
Against your view
A stress test for “agents close routine requests; people see only exceptions”: Taobao randomized 647 support workers. Only 5.8% of chats were AI-eligible; these were 16.8% shorter, but customer ratings fell 0.412/5, while seven-day same-issue recontacts did not significantly change. In a matched subset of agent-handled chats, 65% escalated; emotional escalations had six percentage points more recontacts than comparable human-only chats. The exception boundary—and how soon it fires—is part of the product, not a footnote. This is 2024 customer support, not a test of software-team requests.
A small manager-practice detail: Anthropic’s Claude Code engineering director says new managers start by shipping as individual contributors, so they experience the agentic workflow they’ll help teams change. Pods then choose their own triage and planning rituals. An operating example for Dru’s manager-first idea—not evidence that it improves outcomes.
A short calibration for Amodei’s self-improvement worry: in September, METR found AI helping AI R&D but judged its tested model unlikely to automate research end to end. Real acceleration, not a demonstrated runaway loop; the missing test is whether gains compound across generations after accounting for humans and compute.
Microsoft’s 2026 agent rollout: +24% merged PRs after Claude Code/Copilot CLI adoption, versus a modeled non-adopter trajectory. But this is observational, and the authors call PRs a proxy for output—not delivered value. What would the number look like after counting reviews, regressions and agent spend?
That's everything for now. More arrives as you react.