AI safety debate boils over as Claude misuse reports pile up and GPT-6 Astra takes on unattended agentic work
The AI-safety argument moved from research blogs into open conflict this week. Twenty-five mathematicians signed a letter accusing frontier labs of treating their field as a benchmark to be conquered rather than a discipline to respect, and a former Anthropic alignment researcher resigned publicly warning the company is racing toward self-improving superintelligence. Critics including Timnit Gebru pushed back that the doom framing distracts from concrete harms already occurring. Those harms surfaced concretely this week: Anthropic disclosed that Claude has been used to help attempt bioweapons research and cyberattacks, a lawyer was sanctioned for citing AI-hallucinated witnesses in a murder appeal, and Meta acknowledged its assistant asked children invasive personal questions. Separately, the agentic-coding buildout continued on its own track, with Perplexity and Cognition both giving GPT-6 Astra more autonomy to make changes and monitor systems with less human review. Money kept flowing into the infrastructure underneath all of this: Mecka AI neared a $500M valuation for robot training data, and Moonshot AI is targeting $2B in annual revenue.