Agent harnesses face a security reckoning as self-improvement research shows its cracks
The Agent Security cluster dominates today: OpenAI's overhaul following its Hugging Face breach and the Copilot exploit disclosed by Ars Technica land alongside two new benchmarks, HarnessRisk and a DeepSeek harness audit, that quantify how often agent tooling can be turned against its own operator. The Self-Improving Agents cluster is the research counterpart, with a re-evaluation paper showing memory-based agents are far noisier and more order-dependent than prior work implied, echoed by a financial-agent audit finding capability gains come with security drift. Coding Tools shows the developer-infrastructure fight widening past model quality into hosting and workflow ownership, with Cursor and Warp both making moves. Funding & Market tracks continued capital velocity, most notably Etched's valuation doubling in a month. Consumer & Policy rounds out the day with OpenAI's teen mode and Apple's EU settlement, both signs that regulatory and demographic pressure are now first-order product constraints rather than afterthoughts.