⚡ Uncle Cat AI Radar
ModelsOpen SourceResearch

Meta's Muse Glimmer debuts mid-pack in Arena boards

Arena's first human-preference scores put Meta's new Apache-2.0 model at #77 in Code Arena WebDev and #97 in Text Arena, 26th and 24th among open models.

First independent numbers on Meta's open-weights return

Arena published its first human-preference scores for Muse Glimmer overnight, a day after Meta released the 30-billion-parameter model under Apache 2.0. In Code Arena's WebDev leaderboard the model entered at #77 overall with 1,359 points, ranking 26th among open models. In Text Arena it landed at #97 overall with 1,426 points, 24th among open models.

Muse Glimmer is Meta's most consequential open release in more than a year and the concrete expression of a strategy reset toward open weights. It is built for local, always-on agent work rather than leaderboard chat: roughly 30B parameters including a 1.8B vision encoder, multimodal input, a context window above 131,000 tokens, and a footprint small enough to run in about 18 GB on consumer hardware through LM Studio. Hosting providers including Baseten carried it on day zero.

That design context cuts both ways when reading the Arena placements. Human-preference boards reward conversational polish and are dominated by far larger frontier and near-frontier systems, so a 30B local model placing outside the top 70 is not by itself evidence of weakness on the tasks it was built for — tool use, function-schema adherence and multi-step task completion, which these leaderboards measure only indirectly. Developer accounts of early testing have been more favourable on agentic code exploration than on open-ended creative generation.

Why it matters

Still, the placements are the first outside data point against a release Meta positioned as a strategic reset, and they land at an awkward angle. Being 24th to 26th among open models means Glimmer debuts well behind the Chinese open-weight tier — Kimi, GLM, Qwen and DeepSeek derivatives — that has spent the past year defining what an open model at the frontier looks like. Meta's argument for the pivot rests on out-executing that cohort, not on trailing it in the boards it will inevitably be compared on.

Arena separately noted this week that the gap between proprietary and open-source performance has narrowed to its tightest point yet. Glimmer is not where that narrowing is happening, which is the more pointed finding for Meta.

Sources