Key Takeaways
- Zhipu AI released GLM-5.2 under MIT license, supporting 1M context and rivaling Opus 4.8, though real-world tests reveal quality gaps. — via 1 2 3 4
- xAI launched Grok Imagine Video 1.5 for image-to-video, and Grok 4.3 is now available on Amazon Bedrock with leading hallucination and tool‑use performance. — via 1 2
- NVIDIA and Azure achieved a MLPerf training record: 8,192 Blackwell GPUs trained Llama 3.1 405B in 7.07 minutes. — via 1
- NVIDIA Research introduced SpatialClaw, a training‑free spatial reasoning agent that outperforms prior methods by 11.2 points across 20 benchmarks. — via 1
- OpenAI’s leaked financials show 40%+ gross margins on inference while training remains extremely expensive; automated AI research is proposed to improve efficiency. — via 1
- Ethan Mollick warns that open‑source coding models lag 8–12 months behind closed ones, leaving a 4–8 month window to harden systems. — via 1
1. Model and Platform Updates
- GLM-5.2 has been released by Zhipu AI under the MIT license, featuring a 1M context window and two inference effort levels. It is compared to Opus 4.8 in performance, but real-world tests by Ethan Mollick show errors on specific tasks (e.g., shader coding) and that Fable wove constraints better in a poem, highlighting benchmark limitations. — via 1 2 3 4
- xAI launched Grok Imagine Video 1.5, a new image-to-video model with better realism, physics, and speed. Additionally, Grok 4.3 is now available on Amazon Bedrock, offering state-of-the-art hallucination rates and tool‑use abilities. — via 1 2
- NVIDIA Research introduced SpatialClaw, a training‑free spatial reasoning agent that uses code as an action interface, composing perception modules and adjusting strategies across steps. It outperforms prior methods by an average of 11.2 points on 20 benchmarks and works consistently across six model backbones. — via 1
2. Infrastructure and Training Efficiency
- NVIDIA and Azure set a new MLPerf training record: 8,192 Blackwell GPUs (GB200 NVL72) trained Llama 3.1 405B to the target in 7.07 minutes. — via 1
- NVIDIA is collaborating with Google Cloud to provide a Vera Rubin NVL72 cluster for IneffableLabs, and with Coherent Corp to expand advanced photonics manufacturing in Texas to boost AI infrastructure. — via 1 2
- Leaked financial data suggests OpenAI operates with 40%+ gross margins on serving customers, but training costs remain extremely high. Ethan Mollick notes that automating AI research could be a path to improving training efficiency. — via 1
3. Evaluation and Real-World Performance
- Ethan Mollick compared GPT-5.2 and GLM-5.2 Deep Think Max on a shader task; GLM-5.2 had errors, while in a poem task, Fable better integrated constraints, showing that benchmarks may miss quality differences. He also criticized the GDPval-AA v2 benchmark for using AI to evaluate AI on public questions with unclear ELO. — via 1 2 3
- Ethan Mollick warns that assuming open coding models lag 8–12 months behind closed, the window to harden systems against Mythos‑class models is only 4–8 months, emphasizing the need for publicly available defensive Mythos‑class models today. — via 1
- OpenAI discussed evaluation methodology: using de‑identified real user requests to simulate model behavior before release, and highlighted the need for new benchmarks as current ones saturate. — via 1 2
- Anthropic released an economic research framework for tracking Claude Code extensions, examining user personas, task value changes, and domain expertise effects. — via 1
