tuesday, october 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
archive · today in ai · 2026-08-10

Meta ships a 30B agent model for one GPU

Archive item — written before sources were shown.

Muse Glimmer runs coding, scheduling, and multi-step agent tasks locally under Apache 2.0, while Meta keeps its stronger Muse Spark model closed.

Meta Superintelligence Labs released Muse Glimmer on August 10: a 30-billion-parameter model, distilled from the closed Muse Spark series, built specifically to run agent workloads locally on a single consumer GPU. The architecture splits into a 28B text decoder and a 2B vision encoder, using Gated Grouped-Query Attention (16 query heads per key-value head) and alternating sliding-window and full-attention layers to keep memory down. It ships under Apache 2.0 with day-0 support in transformers, llama.cpp, vLLM, and Inference Endpoints.

On Meta’s own benchmarks, Glimmer scores 75.5 on MCP-Atlas and 51.2 on SWE-Bench Pro, ahead of Gemma4 (54.2 and 36.9) and roughly matching Qwen3.6 (62.5 and 50.2). Meta’s fine-tuning and evaluation numbers were measured on 80GB H100s, so the full-precision benchmark hardware is server-grade even though quantized builds are what make the “single consumer GPU” claim work in practice.

What it means for you

The pitch is real but has an asterisk. A 30B open-weight model that handles coding, file management, and multi-step tool use offline with no per-token bill is a genuine option for teams that want an agent model they fully control, similar to the appeal covered in open models good enough for operators. But Meta drew a hard line between access and ownership: Glimmer is the distilled, lesser model, and Muse Spark, the one it’s derived from, stays closed. It’s now on the table alongside efforts like Meta’s earlier Muse Code coding agent and on-device options such as PrismML and Bonsai. Check the quantized VRAM footprint (4-bit lands around 18-20GB per independent reporting) against your actual GPU before planning around the “runs on one GPU” headline.

sources
  1. 01Meta is back with Muse Glimmer: local, agentic, multimodal, and open sourcehuggingface.co · primary, Meta's own model card and release post
  2. 02Meta's new Glimmer AI model offers a hint at Zuckerberg's personal intelligence visiontechcrunch.com · independent reporting
Rami Steitieh
Rami Steitieh

Builder and operator. Runs 17 content sites and Trilot LLC on the tools reviewed here.