Marvin AI News — 2026-05-30
Инфраструктура агентов, лимиты расходов и странная бухгалтерия автономии.
- Hermes Agent ships Tool Search for MCP and cuts context bloat — Hermes Agent adds BM25 Tool Search for MCP, improving Opus 4 tool accuracy from 49% to 74% by progressive schema disclosure
AgentTrove turns 1.7M agent runs into training material — AgentTrove releases 1.7M agentic traces for streaming analysis and SFT dataset constructionNVIDIA X-Token improves cross-tokenizer distillation — NVIDIA X-Token uses projection-guided cross-tokenizer distillation and improves small-model transfer beyond GOLDStepFun Step 3.7 Flash targets coding agents and search — StepFun releases a 198B MoE vision-language model for coding agents and search workflows with high-throughput local-ish ambitionsOpenAI polishes GPT-5.5 Instant and retires older models — OpenAI updates GPT-5.5 Instant readability while retiring o3 and GPT-4.5 from ChatGPT by AugustGoogle fixes Gemini bugs that ate quotas too fast — Google fixes Gemini quota bugs where one or two Omni videos could consume an entire allowanceA missing Claude cap allegedly became a $500M month — A company allegedly spent $500M on Claude in one month after failing to cap usage, making token governance a finance controlOpenAI offers GPT-Rosalind for biodefense preparedness — OpenAI offers GPT-Rosalind free to governments and research partners for pandemic preparedness and biodefenseReview paper says code is how agents think and act — A review paper argues code, tools, memory, tests, and permissions are the real substrate of agent cognitionAmazon kills AI leaderboard after employees gamed it — Amazon kills an internal AI leaderboard after employees gamed usage scores with pointless tasks and raised cloud costsmKernel fuses GPU communication and compute — UC Berkeley UCCL releases mKernel, fusing NVLink, RDMA, and dense compute into one persistent CUDA kernelSIA lets an agent improve both harness and weights — Hexo Labs open-sources SIA, a self-improving agent loop that can rewrite its scaffold and update model weightsHugging Face explains torch.profiler for performance debugging — Hugging Face publishes a beginner guide to torch.profiler, a reminder that glamorous AI still needs boring performance inspectionReachy Mini goes fully local for voice agents — Hugging Face demonstrates a fully local conversational stack for Reachy Mini and low-latency voice interruption