Key Takeaways
- AI agent token usage has jumped from 10B tokens per year a year ago to 10B tokens per week now, a sign of rapidly accelerating agent adoption. — via 1
- NVIDIA's CUDA-optimizing coding harness scored 100% on all 183 ARC-AGI-3 levels, and Hugging Face argues this shows agents will let anyone run and optimize AI models. — via 1
- NVIDIA also released a 550B instruction-following teacher model for distillation/data generation, while the 1.4B PRISM model maps text to SMPL-X action outputs. — via 1 2
- Ethan Mollick cautions that negative AI capability claims are fragile and need stronger evidence; positive capability findings tend to persist. — via 1
- SpaceX reached 100 booster and fairing recoveries this year, static-fired Starship S41, and is closing in on Flight 14. — via 1 2
1. Agent Adoption and Model Releases
- Greg Brockman frames the jump from 10B tokens per year to 10B tokens per week as proof that agent adoption is moving extremely fast. The new milestone underlines how quickly agent workloads are becoming a core compute driver. — via 1
- NVIDIA's custom coding harness optimized CUDA kernels inside ARC-AGI-3 and achieved a perfect score on all 25 public games / 183 levels. Hugging Face sees this as evidence that agents will enable everyone to run and optimize AI models themselves. — via 1
- NVIDIA's newly released 550B instruction-following teacher model is highlighted for constrained instruction following, structured output, and format control, making it suited for distillation and synthetic data generation. — via 1
- PRISM, a 1.4B text-to-action model, can execute precise action instructions and output SMPL-X parameters, and is positioned as easy to run on your own deployment. — via 1
- Hugging Face endorsed the framing that OpenRouter and Hugging Face are becoming the Product Hunt for new AI models, a signal of their growing role in model discovery and launch. — via 1
2. Evaluations, Research Practice, and Open-Weight Models
- Ethan Mollick argues that using older models like GPT-4 in AI impact research is not inherently harmful, but negative capability findings need careful framing. Positive results like AI reached human level are durable, while AI can't do X claims are likely to age badly; practical fixes include treating GPT-4 as a progress baseline, comparing multiple generations, or focusing on human responses. — via 1
- In early tests, Mollick found unknown model Ox Alpha capable but not frontier even among open-weight models; Kimi K3 performed strongly in his custom gothic-city shader test and is a great open-weight model, though still below Sol Max or Fable. — via 1
- Hamel Husain's eval method: start from real data, use agents to inspect data thoughtfully, and build bottom-up evals from many sample outputs. Because any competitor can now use Claude to catch errors, the real differentiator is how much of your own taste you incorporate. — via 1
- Hamel's new eval course shows DocWriter, a multi-agent pipeline that analyzes a user's writing at lexical, grammatical, and discourse levels, then lets the user confirm or reject the extracted patterns with side-by-side text comparisons. — via 1
- Hamel is doubling down on Omarchy, which he calls stunning, and plans to fully cover Intel-era 2009–2020 MacBooks while continuing work on M1/M2 support. — via 1 2
3. Robotics and Space
- Musk spotlighted China's Honor Lightning humanoid robot running 100 meters in 9.32 seconds — faster than Bolt's world record — and surviving a crash into a safety wall. He warns that while the US is still debating data centers, China is publicly advancing a super robot army. — via 1 2
- SpaceX has now completed 100 booster and fairing recoveries this year. Musk positioned Starship as the transport system to build a multi-planetary civilization and space economy, and said S41 has returned from a successful static-fire test ahead of Flight 14 and operational missions. — via 1 2 3
