Autonomous Offense, Graded Honestly
In this episode:
Autonomous Offense, Graded Honestly — Blinded cyber benchmarks show GLM-5.3 Flash matching the full model at a fraction of the cost, while ExploitBench records an agent building a full V8 renderer exploit chain autonomously. SiftRank reframes vulnerability discovery as an attention-routing problem, finding a root shell in firmware with small models matching large ones.How Shallow Is the Alignment — BLOOM-WILT uses next-token logit blending to audit safety alignment without training, overturning prior model safety rankings and showing guardrails are thinner than claimed. Phantom-kv strips refusal behavior via a swappable key-value cache graft that leaves base weights untouched, reducing refusals from 25 to 5 out of 60 harmful prompts.Cheap Tokens, Expensive Lies — NahCrofAI ran a two-year inference reseller scam, advertising frontier models while silently routing to cheaper ones via OpenRouter at up to 20x markup. A separate proof-of-concept against the Muse dictation app shows how local assistant tools can be redirected to attacker-controlled endpoints, amplifying trust-boundary risks.Decisions Without Words — A two-line llama.cpp recipe extracts calibrated classification probabilities from any model's next-token logits, while Laya-MLX ports a bidirectional encoder to Apple Silicon for 13ms yes/no answers with zero output tokens. Both approaches challenge the need for bespoke classifiers when reading a capable model's own distribution suffices.Sovereignty Moves Down the Stack — Tim Dettmers announces projects claiming 450 tok/s for 35B models at 1.5-bit quantization on modest hardware, while Zhipu describes GLM-5.3 agents helping deploy their own successor on 100k+ Chinese accelerators. A native Rust/Vulkan training backend covers 143 architectures without CUDA, and open-weight releases continue at roughly one frontier model every ten days.Cutting Weights With Physics — Multiverse Computing reframes depth pruning as a combinatorial Ising spin-glass problem, capturing pairwise block interactions that ranking heuristics miss and holding benchmarks 23 points above baselines at 50% depth removal. Altar-1 applies expert pruning and 4-bit quantization to the 753B GLM-5.3, fitting it on four H200s at 328 GB.The Agent's Working Conditions — Workspace-OS embeds LibreOffice in an Electron shell with shadow-git checkpoints and reviewable agent diffs, while Berkeley and Microsoft show that five communicating agents solve 8% of hard reasoning tasks versus 2.2% for independent runs. LSAP and IWE fill lower-level plumbing with language-server cognitive operations and blast-radius-declared memory mutations.Compression, Texture, and Attention — A gzip-as-language-model experiment generates text purely through DEFLATE compression with beam search, producing recognizable style with no neural weights. Heat Kernel Textures win an ECCV best-paper for seam-free mesh appearance via geodesic Gaussians, and an essay argues for reclaiming human attention from algorithmic feeds through deliberate browsing habits.Keywords: agent-harness, alignment, api-security, apple-silicon, attention-economy, autonomous-agents, calibration, classification, collaboration, compression, computer-graphics, cost-efficiency, cyber-security, developer-tools, encoder-models, exploit-benchmarks, guardrails, hardware-independence, inference-fraud, inference-sovereignty