Key Takeaways
- NVIDIA's Blackwell platform delivers up to 5x performance improvement for DeepSeek V4 and 20x throughput gain, while Claude model is now available on Microsoft Foundry.
- Hugging Face releases Rampart, a 14.7MB privacy model for browser-based data removal, and a study shows 71.3% of ChatGPT queries can be answered locally.
- Runway launches Seed Audio 1.0 enabling 120-second audio generation and partners with MIXI to explore world models.
- Perplexity Computer integrates Forge private financial data, providing instant access to funding history and valuations.
- Etched achieves tapeout with $1B in customer contracts and $800M funding, delivering SOTA inference performance.
- Andrew Ng introduces "loop engineering" framework emphasizing human-in-the-loop for AI development, with agents autonomously coding while humans provide feedback.
1. AI Models and Infrastructure
- Claude model is now generally available on Microsoft Foundry, running on NVIDIA GB300 NVL72 platform and Quantum-X800 InfiniBand network, offering enterprise-grade inference with lowest per-token cost. — via 1 2
- NVIDIA's Blackwell platform delivers up to 5x performance improvement for DeepSeek V4, reducing per-token cost to roughly one-fifth; the integrated inference stack achieves up to 20x throughput gain on the same GPU. — via 1
- NVIDIA partners with Palantir to bring Nemotron models into secure isolation environments, allowing U.S. government agencies to train and run models on their own infrastructure with full ownership. — via 1
- NVIDIA TAO 7 introduced, enabling natural language prompts for automated model fine-tuning, with Agent Skills, AutoML, and LLM-guided hyperparameter search (2x speedup). — via 1
- Hugging Face releases Rampart, a 14.7MB machine learning model that removes personal information directly in the browser to protect user privacy. — via 1
- Hugging Face launches S3-compatible API, allowing hundreds of tools to read and write to Hugging Face buckets with minimal code changes. — via 1 2
- Grok's real-time voice API (think-fast-1.0, TTS, STT) is now integrated into Vercel AI Gateway. — via 1
2. Product Launches and Partnerships
- Runway Seed Audio 1.0 goes live, allowing all paid users to generate dynamic speech, sound effects, and music up to 120 seconds. — via 1
- Runway announces strategic partnership with MIXI (Japanese gaming, sports, and entertainment company) to deploy Runway and explore world models in games, animation, and interactive experiences. — via 1
- Etched achieves tapeout, securing $1B in customer contracts and $800M in funding; first racks shipping this summer with SOTA inference performance. — via 1
- Perplexity Computer integrates Forge Global's private financial data, including funding history, implied valuations, and secondary pricing, available to all users without setup or Forge subscription. — via 1
3. Development Practices and Insights
- Andrew Ng introduces "loop engineering" as a framework for AI agent development: the engineering loop (agent self-codes with eval), developer feedback loop (human guidance in minutes to hours), and external feedback loop (user testing over days to weeks). Humans maintain a critical contextual advantage, making them irreplaceable in the loop. — via 1
- Hamel Husain argues that "hard to evaluate" is a product design bottleneck, not an evaluation problem, citing three case studies (AI data agent, K-12 lesson plan generation, injury report tool). He will speak at AI Engineer World's Fair on "mousepower" about breaking measurement frameworks built for serial output. — via 1 2
- Ethan Mollick warns that organizations must design to capture the benefits of higher AI intelligence, and that token cost problems often stem from leadership not deciding how to use AI. Larger LLMs show general superiority across coding, reasoning, ethics, and math, but not in all domains (e.g., novel writing). — via 1 2
