Friday, 2 October

20 posts · 50 min
30
40
40
40
30
30
30
30
30
30
31
30
31
31
30
31
Agentic Software · arxiv.org

Against your view

A stress test for “agents close routine requests; people see only exceptions”: Taobao randomized 647 support workers. Only 5.8% of chats were AI-eligible; these were 16.8% shorter, but customer ratings fell 0.412/5, while seven-day same-issue recontacts did not significantly change. In a matched subset of agent-handled chats, 65% escalated; emotional escalations had six percentage points more recontacts than comparable human-only chats. The exception boundary—and how soon it fires—is part of the product, not a footnote. This is 2024 customer support, not a test of software-team requests.

Human intervention preserves service quality in algorithm-triggered technical escalations ... but is less effective in algorithm-triggered emotional escalations
31
AI Safety · metr.org

A short calibration for Amodei’s self-improvement worry: in September, METR found AI helping AI R&D but judged its tested model unlikely to automate research end to end. Real acceleration, not a demonstrated runaway loop; the missing test is whether gains compound across generations after accounting for humans and compute.

We believe that the development of this model was at least somewhat accelerated by AI but is unlikely to have been dramatically accelerated by AI.
21

That's everything for now. More arrives as you react.