🤖 System Prompt Learning Boosts Local LLMs
In this episode:
• System Prompt Learning Boosts Local LLMs
• Gemini 2.5, MCP Tools, and Model Transparency
• Hugging Face Open-Source Robots Enter the Scene
• ️ Video Generation AI Targets Gaming and 3D Workflows
• Model Discovery, Merging, and Local LLM Ecosystem Growth
• ️ Local LLMs, Hardware, and Tooling Challenges on Macs and PCs
• ️ Embeddings, Rerankers, and Mobile LLM Deployment Advances
• 💻 Agents, Coding Terminals, and Local Code Intelligence
• Research: Scenario Generation for Autonomous Vehicles and Data Center Interconnects
• Learning ML: Structured Pathways and Foundational Resources
• ️ Security: Clipjacking and MCP Server Exploitation Risks
Local large language models (LLMs) are narrowing the gap with cloud giants, thanks to innovations like System Prompt Learning (SPL). SPL, implemented in the optillm plugin, targets a well-known shortcoming: most local LLM deployments rely on simple prompts, missing out on the sophisticated, experience-driven system prompts that power services like ChatGPT and Claude. SPL introduces a feedback-driven mechanism—what Andrej Karpathy described as the “third paradigm”—enabling models to learn and refine problem-solving strategies from their own usage history, not just from pretraining or static fine-tuning.
In practice, SPL classifies incoming problems (math, coding, word puzzles, and more), then builds and continuously tunes a database of human-readable strategies, stored in JSON. Users can inspect, edit, and extend these strategies, which are automatically matched to new queries. On math benchmarks, SPL demonstrated concrete improvements: for instance, gemini-2.0-flash-lite’s performance on the Arena Hard benchmark jumped from 29% to 37.6% after adopting SPL—a notable +8.6% gain. Over 500 queries, the system evolved 129 strategies, refining 97 of them, and did so entirely locally without cloud dependencies (more: url (https://www.reddit.com/r/LocalLLaMA/comments/1l1bjhm/systempromptlearningteachingyourlocalllms)).
The SPL approach is compatible with any OpenAI-style API, including llama.cpp, Ollama, and vLLM, and operates in either inference-only or active learning modes. Its minimal overhead and open-source flexibility make it a compelling upgrade for local LLM enthusiasts seeking more autonomy and transparency in model behavior, all while keeping data and strategy development on-device.
Google’s Gemini 2.5 update further advances the state of LLMs, particularly for developers and power users. Beyond topping coding and academic leaderboards (notably WebDev Arena and LMArena), Gemini 2.5 Pro and 2.5 Flash are rolling out new features: native audio output for more fluid conversations, enhanced security controls, and Project Mariner’s computer-use abilities. Of special interest to the developer crowd is “Deep Think,” an experimental mode that pushes Gemini’s reasoning for complex math and code, and the expansion of “thinking budgets”—giving users granular control over how much effort the model expends on a given task.
Transparency and tool integration are also priorities. Gemini’s API and SDK now support MCP (Model Context Protocol) tools, opening up access to a wider array of open-source integrations. The API introduces “thought summaries,” giving developers insight into the model’s reasoning process, and extends support for MCP tools, enabling more sophisticated workflows and custom toolchains (more: url (https://blog.google/technology/google-deepmind/google-gemini-updates-io-2025)).
But not all is smooth sailing. News publishers, represented by the News/Media Alliance, have sharply criticized Google’s expanded “AI Mode,” which presents AI-generated answers directly in search results. They describe Google’s approach as “theft,” arguing it diverts traffic and revenue from publishers by summarizing content without meaningful compensation or opt-out options, aside from removing their work from search entirely—a move described as both technically and economically infeasible for most media outlets (more: url (https://www.theverge.com/news/672132/news-media-alliance-google-ai-mode-theft)). This tension between AI utility and content ownership is likely to intensify as LLM-powered interfaces become the norm for information retrieval.
Hugging Face’s foray into open-source robotics marks another step in democratizing advanced technology. The introduction of HopeJR—a full-size humanoid with 66 degrees of freedom—and Reachy Mini, a desktop robot for AI app testing, signals intent to make robotics accessible and modifiable. Priced at approximately $3,000 for HopeJR and $250–$300 for Reachy Mini, these robots are designed to be affordable alternatives to proprietary, closed competitors.
...