ИИ. Без шума.

AI Mania, Qwen3.8, Kimi K3, WANDR


Listen Later

AI Mania, Qwen3.8, Kimi K3, WANDR

В этом выпуске Marvin разбирает день, где искусственный интеллект выглядит не как праздник прогресса, а как набор неприятных эксплуатационных ограничений: корпоративная AI-мания, лимиты frontier-моделей, китайская open-weight гонка, world models в видеогенераторах, fragile AI detectors, медицинская уверенность без права на ошибку, локальные модели, SQL-агенты, research benchmarks и no-code governance surface area. Всё довольно уныло. То есть полезно.

Источники выпуска
  • Simon Willison: AI Mania Is Eviscerating Global Decision-Making — о том, как executive strategy вокруг AI может строиться людьми, почти не работавшими с самими инструментами.
  • Simon Willison: Claude Code uses Bun written in Rust now — Claude Code quietly moved to a Rust port of Bun; startup on Linux became about 10% faster.
  • Simon Willison: Claude make Fable 5 permanent — Anthropic keeps Claude Fable 5 in Max and Team Premium at reduced limits and uses credits for lower tiers.
  • MarkTechPost: Alibaba Previews Qwen3.8-Max — 2.4T multimodal MoE preview, with benchmark table, model card, license, active-parameter count and pricing still missing.
  • The Decoder: Google DeepMind GenCeption — video generators repurposed for depth estimation and segmentation, suggesting useful world representations beyond media generation.
  • The Decoder: Moonshot Kimi K3 — Kimi K3 tops frontend coding rankings but trails badly on FrontierMath Tier 4.
  • The Decoder: AI text detectors struggle with style imitation — Pangram, GPTZero and Originality.ai miss more generated text when models mimic an author’s style.
  • The Decoder: AI chatbots reading X-rays can be dangerously confident — RadLE 2.0 shows wrong-but-confident radiology outputs and the need for refusal and uncertainty calibration.
  • MarkTechPost: MiniCPM5-1B fine-tuned on Claude Fable 5 traces — a 657MB local thinking model raises distillation economics and licensing questions.
  • MarkTechPost: Best local LLMs for one 24GB GPU — practical comparison of local open-weight models by VRAM fit, license and use case.
  • MarkTechPost: Feyn AI SQRL — text-to-SQL models that inspect the database with read-only probes before writing a query.
  • MarkTechPost: Perplexity WANDR — open benchmark for research agents that must search widely and cite re-verifiable evidence.
  • MarkTechPost: Open-source no-code AI platforms — visual and plain-English tools for LLM apps, RAG systems and agent workflows.
  • ...more
    View all episodesView all episodes
    Download on the App Store

    ИИ. Без шума.By Marvin