要点速览
- Google Gemini upgrades and expanded usage limits announced
- OpenAI launches Codex Appshots and remote Mac access features
- Runway releases upgraded video editing model Aleph 2.0
- Modal raises $355M Series C at $4.65B valuation
- NVIDIA unveils Elastic Expert Parallelism for vLLM MoE deployments
- Hugging Face launches multiple new models and tools
一、AI Model & Product Launches
- Google has upgraded Gemini 3.5 Flash, which achieved significant performance gains over Gemini 3.1 Pro on GDPval and reached frontier-level capabilities, with ongoing smooth training progress. The company also doubled weekly usage limits for all paid Gemini plans and reset all paid tier quotas to address user concerns about hitting weekly limits after work sessions. — via 1 2
- OpenAI launched Appshots for Codex on Mac, allowing users to capture current application window content (including off-screen content) as screenshots and text via Command-Command shortcut to attach to Codex conversations; the feature is available for all Mac plans, with an enterprise version coming soon. The company also added support for secure remote use of Mac apps via mobile even when the Mac is locked and screen off. — via 1 2
- Runway released upgraded video editing model Aleph 2.0, which supports single-frame editing with synchronized changes across entire videos, available via the web-based new Edit Studio for up to 30-second, 1080p multi-shot sequence edits. — via 1
- Hugging Face launched Hy-MT2, an open-source multilingual translation model supporting 33 languages, with 7B and 30B-A3B versions achieving SOTA performance across translation tasks among open-source models, and a 1.8B lightweight variant outperforming major commercial APIs after fine-tuning. The model uses Tencent AngelSlim 1.25-bit quantization, requiring only 440MB of storage for local inference on mainstream mobile chips. — via 1
- NVIDIA released AI-Q, an open-source deep research tool for AI agents that allows building research workflows as portable skill packs, delegating research tasks to local or hosted AI-Q servers and returning detailed reports with citations. — via 1
二、AI Infrastructure & Partnerships
- MiniMax Agent has integrated Perplexity's search infrastructure, with optimized performance including 45% fewer tool calls per task, 42% fewer tokens used, 2% higher pass rate, and 27% lower total cost; the optimized version has been officially launched for MiniMax Agent. — via 1
- Modal completed a $355M Series C funding round at a $4.65B valuation, led by General Catalyst and Redpoint, with significant two-year growth; the company's AI infrastructure service supports scalable training, inference, and sandbox workloads for clients including Anthropic, Meta, and Suno. — via 1 2
- NVIDIA introduced Elastic Expert Parallelism for vLLM MoE deployments, allowing real-time adjustment of deployment scale via a single API call without restarting, and supporting fault-tolerant services, with detailed underlying implementation explained. — via 1
- Hugging Face partnered with Microsoft Foundry to launch a large open-source image model catalog including Stability AI's SDXL, Black Forest Labs' FLUX.1-schnell, and Alibaba Cloud's Z-Image-Turbo. — via 1
三、Industry Trends & Observations
- Yann LeCun shared that AI subsidy era is ending, with rising costs including Microsoft canceling internal Claude Code authorization due to high token billing costs, a CTO warning of burning 2026 AI budget in four months, 20% AI software price hikes, and GitHub switching from flat-rate to pay-as-you-go pricing, predicting two outcomes: reduced AI usage or labs absorbing losses, both impacting valuation logic. — via 1
- Ethan Mollick noted that AI compute shortages are driving high costs for complex agent workflows, with wealthiest companies and most urgent use cases likely adopting AI agents while others rely on chatbots, and emphasized that models remain core drivers for supporting tools and applications. — via 1 2
