Key Takeaways
- Meta released open weights for Muse Glimmer, a 30B-parameter multimodal model with a 120K+ context window under Apache 2.0, while teasing Muse Spark 1.2 weights. — via 1 2 3
- Hugging Face shipped day-0 support in transformers and llama.cpp, with a DFlash path that is 2-4x faster and runs on 24GB VRAM. — via 1 2
- Ethan Mollick called it the best non-Chinese open-weight model in about a year, though still behind Chinese open models and closed frontier systems. — via 1
- NVIDIA says Qwen 3.5 now serves at 25K tokens/s/gpu on GB200 in vLLM, with Qwen 3.8 preparation underway. — via 1
- Demis Hassabis highlighted WeatherNext in Nature: cyclone forecasts gain an average of 24 extra hours of preparation. — via 1
- Hamel Husain built specspecs to expose wasted compute from rejected draft tokens in speculative decoding. — via 1
1. Meta Returns to Open Source: Muse Glimmer and Muse Spark 1.2
- Meta released weights for Muse Glimmer, a dense 30B-parameter multimodal model with a 120K+ context window designed for long-running agent workflows. It runs locally on consumer hardware with about 24GB VRAM, and NVIDIA AI reports ~20K tokens/sec on a single GPU; it is available under Apache 2.0. NVIDIA AI welcomed Meta's return to open models. — via 1 2 3
- Hugging Face provided day-0 integration in transformers and llama.cpp and demonstrated a DFlash path that accelerates inference 2-4x. It also ran tests, fine-tunes, and demos to confirm agentic reliability on local hardware. — via 1 2
- Yann LeCun confirmed Meta will also release Muse Spark 1.2, the latest base model, as open weights, reinforcing the company's commitment to open AI. — via 1
- Ethan Mollick called this major news: a good model and the best non-Chinese open-weight model in about a year, but still behind Chinese open-weight models and closed frontier systems. The open question is whether Meta can sustain a cadence of new open releases. — via 1
2. Faster Serving, Better Prediction, and Observability
- NVIDIA AI announced that Qwen 3.5 has been optimized in vLLM, hitting 25K tokens/s/gpu in production serving on GB200 systems. The work is explicitly preparation for Qwen 3.8, signaling close co-optimization between open models and new accelerator hardware. — via 1
- Demis Hassabis highlighted the WeatherNext model's Nature publication: it predicts cyclone tracks and intensity with high accuracy, giving communities an average of 24 extra hours of preparation. — via 1
- Hamel Husain warned that speculative decoding can waste compute on rejected draft tokens when run blindly. He released specspecs, a tool to observe and fix those hidden inference costs. — via 1
- Yann LeCun explained that the dial-up modem handshake is real-time channel estimation and adaptive equalization, sending known signals and analyzing echoes to adapt to the line. The idea dates to Bob Lucky's work at Bell Labs in the 1960s. — via 1
3. Agents, Real-World Gains, and Infrastructure Bets
- Greg Brockman shared a concrete Codex use case: reading contract fine print to uncover savings, with one user reportedly cutting about $6,000 per year from electricity costs. — via 1
- Supabase has been integrated into Perplexity Computer, letting users query production data directly inside Perplexity chat. Aravind Srinivas highlighted the integration as a sign of agents moving toward real data. — via 1
- Elon Musk predicted that AI agents will generate far more internet traffic than humans, and that Starlink V3/V4 could provide 50-250x today's global bandwidth over the next decade, connecting billions of unconnected people plus digital and physical robots. — via 1
- Elon Musk described Grok Build as becoming an all-in-one creation environment, with Grok Imagine generating images and video directly in the workflow. He added that Grok 4.5 leads in daily usage as a design tool and shows strong character consistency. — via 1 2 3
- Ethan Mollick criticized agent products like ChatGPT Work and Claude Cowork for hiding decisions from non-programmers. Instead, they should explain which tasks to delegate and how to generalize, the way good product managers do. — via 1
- swyx advised periodically deleting accumulated AI skills because they consume context and can interact unpredictably. It is a useful maintenance practice for agent-heavy workflows. — via 1
