Trump unveils his new Super Intelligence Force
The Trump administration launched a new task force branded around 'super intelligence,' positioned as its latest response to the AI safety debate.
Robotics research converges on a single idea today: treat video generation as the substrate for action prediction. Seven papers (ProWAM, Dream4ACT, NAVA-WAM, HelixWorld, Spatial Memory Intelligence, MotorMind, ProAR) attack different pieces of the same problem, from cross-embodiment action representations to long-horizon memory and audio-visual co-generation. Separately, verification and reasoning-reliability work (VeriHarness, CoT faithfulness, Fold2Reason) suggests the field is taking agent trustworthiness more seriously as task horizons lengthen. On the policy side, the Trump administration's new Super Intelligence Force and an accompanying non-binding safety pact arrive the same week MIT Tech Review and The Verge both ask why public sentiment toward AI keeps souring even as usage climbs. For operators, the throughline is that embodied AI and agentic verification are maturing faster than public trust, which is a gap worth planning around.
The Trump administration unveiled a new Super Intelligence Force alongside a non-binding safety pact, drawing skepticism in commentary framing it as rebranding rather than substantive policy. Meanwhile, two separate pieces -- from MIT Technology Review and The Verge -- independently probe why public sentiment toward AI keeps curdling even as adoption and usage keep rising, pointing to a widening gap between institutional AI narratives and lived public experience.
The Trump administration launched a new task force branded around 'super intelligence,' positioned as its latest response to the AI safety debate.
TechCrunch's Equity podcast frames the administration's super-intelligence branding and non-binding safety pact as an attempt to rebrand AI's public image rather than regulate it.
MIT Tech Review examines the gap between widespread public distaste for AI and continued heavy usage, including a startup explicitly building a 'self-loathing AI.'
The Verge explores how decades of framing human cognition through computational metaphors has shaped -- and possibly distorted -- public understanding of AI.
Seven papers push world-action-model architectures for robotics: shared cross-embodiment action interfaces, sparse visual sub-goal planning, long-term spatial memory, native action-prior pretraining from raw video, real-time audio-visual co-generation, and VLM-driven zero-shot manipulation without task-specific policies. Video generation as a shared substrate for perception and control looks like a dominant robotics paradigm, not a one-off technique.
ProWAM predicts sparse ordered visual sub-goals alongside actions, cutting the cost of long-horizon video-based robot control while improving out-of-distribution robustness.
Dream4ACT renders joint configurations as images from virtual cameras so one video-action model can be trained across robot embodiments without an embodiment-specific decoder.
SMI uses an MLLM-driven memory manager (clustering, sparsification, retrieval, filtering) to keep long-video world models spatially consistent over extended rollouts.
NAVA-WAM pretrains the action policy directly on observation-only video via flow-matching, skipping the usual detour through latent-action models or representation transfer.
HelixWorld co-generates camera-grounded spatial audio with visuals in real time, distilling a bidirectional teacher into a streaming student that runs at 24 FPS on one GPU.
MotorMind lets a general-purpose VLM operate a robot directly via mid-level actions and asynchronous feedback, without task-specific policy training or grounding tools like SAM3.
ProAR anchors autoregressive video generation to a predicted goal frame plus future-representation self-alignment, turning reactive next-chunk prediction into goal-directed reasoning.
Three papers tackle agent trust from different angles: VeriHarness turns a generator model into an agentic verifier resolving rollout disagreement without reference answers; a faithfulness study finds efficient reasoning training degrades CoT faithfulness more than monitorability; and Fold2Reason shows non-linguistic structural supervision (protein folding) transfers to broad reasoning gains. Verification and reasoning quality look like first-class research targets, not side effects of scale.
VeriHarness turns a generator LLM into an agentic verifier with evidence tools, resolving disagreement across rollouts to pick the best final artifact without reference answers.
Training models to reason with fewer tokens reduces chain-of-thought faithfulness in most settings, mainly through inconsistency, but monitorability holds up better than faithfulness.
Post-training on protein-folding structural data (FoldingCorpus/Fold2Reason) improves macro-average reasoning accuracy from 45.09% to 48.33% across 10 unrelated benchmarks.
Five papers refine post-training mechanics: a committee framework halves LLM bias scores; entropy-guided gating (SCALE) lets SFT reverse or extrapolate features instead of only suppressing them; Local Support Learning fixes forgetting up to 7B parameters via a GMM-gated adapter; Pivot-SD self-distills diffusion LMs on high-impact tokens only; LexReward brings rubric-based rewards to legal LLMs. The thread: finer control over what gets updated, not blunt instruments.
CBM organizes multiple diverse LLMs into debating/committee topologies to share knowledge and mitigate bias beyond what any single model's self-debiasing achieves.
SCALE freezes the pretrained model and SFT delta, then learns entropy-guided gates that can reverse or extrapolate learned features instead of only scaling supervised updates.
LSL pairs a weight adapter with a GMM-based gate that stays closed on prior training distributions, resolving catastrophic forgetting in LLMs up to 7B parameters without storing old data.
Pivot-SD identifies high-impact token commitments (pivots) in masked diffusion LMs via information gain and trains only on those, improving LLaDA-8B-Instruct with just 200 questions.
LexReward scores legal LLM responses along Style, Element, and Chain dimensions with rubric-based rewards, improving DPO and RL training on each dimension independently.