A papers-only day: agent benchmarks find fresh cracks, one recsys paper is already live in production
August 5 was a papers-only day: Hugging Face's daily feed produced 36 arXiv submissions and no qualifying news, consistent with the backfill process used for this date after press feeds could not be fetched retroactively. Benchmarks Keep Finding Agent Reasoning Is Shakier Than It Looks is the strongest thread — MerchantBench shows the best LLM merchant agent reaches only 27.3% of human net assets over a simulated year, and FinIndices shows financial-statement reasoning collapses once formula hints are removed. New Recipes for Teaching Models from Their Own Mistakes groups five self-distillation and credit-assignment papers (PCSD, TurnSight, ReflectRL, DASH, Any-OPD), all extracting denser training signal from agent rollouts rather than discarding failed trajectories. World Models and Video Agents Push Toward Real-Time, Long-Horizon Use covers MiniWorld, ST-WAM, and Video-DeepResearch, the last reporting it beats Claude-4.5-Sonnet and GPT-5 on a new video benchmark — treat that as self-reported until independently reproduced. Efficiency work (LLaDA MoE v2, OmniPack, RestoreKV) keeps cutting inference cost without giving up capability. The most concrete result sits in Applied and Niche Tools: Knowledge-Geometry Decoupling is already live on Shopee, lifting homepage-search GMV per user 1.75% in a production A/B test.