Key Takeaways
- OpenAI's Greg Brockman says an internal version of the next major model Astra solved ten important math/theoretical CS problems with Lean certificates and chain-of-thought proofs at around $2,000 in Sol API compute.
- NVIDIA's Spatial-IQ benchmark shows a wide gap between humans (82.1%) and the best off-the-shelf multimodal model (17.7%) on 3D object counting, while sub-task training lifted Qwen2.5-VL-32B to 62.6%.
- SGLang now supports Thinky Machines' Inkling-Small on dual DGX Spark systems at 24 tok/s, and StudyFetch cut its largest inference workload cost by ~10x with NVIDIA Riva/Parakeet/NIM.
- Runway has added Grok Imagine Video 1.5 to its platform.
- Agentic-loop practitioners still mostly avoid /loop, but multi-agent handoffs, event triggers, and cron can form long-running autonomous loops; many expect human code review to fade within 1-2 years.
1. AI Models and Research
- OpenAI's Greg Brockman announced that an internal version of the next major model Astra solved ten important problems in mathematics and theoretical computer science at a compute cost of roughly $2,000 using the Sol API. The published proofs are accompanied by Lean certificates and chain-of-thought traces, pointing to stronger verifiable reasoning in large models. — via 1
- NVIDIA Research's Spatial-IQ benchmark breaks 3D object counting into nine perception and cognition subtasks and scores them separately. Humans reach 82.1% accuracy, while the best off-the-shelf multimodal models hit only 17.7%; targeted training on the subtasks improves Qwen2.5-VL-32B from 2.9% to 62.6%, giving spatial-reasoning systems a concrete loop for finding weaknesses, improving them, and validating combinations. — via 1
- Brockman also points to the Jevons paradox playing out in AI: as model prices drop, usage expands sharply, with large-scale research possible at extremely low cost, e.g., 143M tokens for about $60. He praises the cheap and efficient Luna model and says using ChatGPT for everyday automation is increasingly cost-effective. — via 1 2 3
2. Infrastructure, Tools, and Platform Updates
- SGLang now officially supports Thinky Machines' Inkling-Small model, running at 24 tok/s (concurrency 1, no MTP) on two DGX Spark systems. This enables capable local agent intelligence on a compact platform, with DSpark support expected soon. — via 1
- StudyFetch leveraged NVIDIA Riva, Parakeet ASR, and NIM microservices to lower the cost of its biggest AI inference workload by nearly 10x. The deployment underpins voice tutoring, real-time personalization, and a new agentic learning platform. — via 1
- Runway has made Grok Imagine Video 1.5 available on its platform, letting users try the video generation model directly in Runway. — via 1
3. Agentic Workflows and Developer Practice
- At an agentic-loop dinner, most attendees said they were not actively using /loop; swyx notes he is one of the few who still find it useful when controllability and autonomy are both needed or when pursuing open-ended "loops generating loops" goals. The group found that multi-agent handoffs, event triggers, and multiple cron jobs can build long-running autonomous loops, and almost all expect human code review to disappear in 1-2 years. — via 1
- swyx observes that the pejorative connotation of "vibe coding" has faded completely. The term is now used across the spectrum, from non-technical users to senior technical people. — via 1
