⚡ Uncle Cat AI Radar
AIGCModels

Grok Imagine Image 2.0 debuts at No.2 in both arenas, then slips to third

SpaceXAI's new image model entered both Arena leaderboards at No.2 behind GPT-Image-2; by 11 August it had been pushed to third in text-to-image by Microsoft's MAI-Image-2.6, while holding second in image editing.

The result

Arena, the crowdsourced model-comparison platform, published leaderboard placements for the newest image model from SpaceXAI — the company formerly known as xAI, rebranded in July after its merger into SpaceX — on Saturday morning UTC. Grok Imagine Image 2.0, tested in its faster "Low" configuration, entered the Text-to-Image Arena in second place with 1,320 points, behind OpenAI's GPT-Image-2 at 1,380. In the Image Edit Arena it also landed second, scoring 1,439 against GPT-Image-2's 1,463 — a jump from seventh place in that arena for the previous generation, Grok Imagine Image Quality, which Arena ranks 14th in text-to-image.

One of those two placements has already changed. As of 11 August the model sits third in text-to-image with 1,316 points, displaced by the MAI-Image-2.6 preview from Microsoft AI, which entered on Monday at 1,336; GPT-Image-2 leads with 1,381. Ranked below the SpaceXAI entry are Reve AI's Reve 2.1 at 1,302, Meta's Muse-Image at 1,282, Reve AI's older Reve 2.0 at 1,270, Google's Gemini 3.1 Flash Image at 1,264, ByteDance's Seedream 5.0 Pro at 1,258 and Alibaba's Qwen-Image-3.0-Pro at 1,257. In the Image Edit Arena it still holds second, 1,439 against 1,463. Both scores remain provisional: roughly 2,700 text-to-image votes against about 70,000 for the leader, and about 5,900 image-edit votes against 184,000. Positions keep moving as voting accumulates.

What shipped

SpaceXAI released Image 2.0 on 7 August as the "Quality Mode" of Grok Imagine, available on grok.com/imagine and in the company's iOS and Android apps. First-party API access has not opened; the company says it is coming, without giving a date, a model ID or pricing. Developer access so far runs through a third party: a preview of the model has been made available on Vercel's AI Gateway. SpaceXAI describes the model as designed to follow detailed instructions precisely, plan typography and page layout the way a designer would, and hold small text legible inside dense composite visuals.

The release bundles an editing toolkit rather than a bare generator: a magic-wand tool that alters only the region a user points at, segmentation-based selection, background removal with transparent export, multi-reference editing that accepts up to five input images in a single generation, and aspect-ratio conversion via Smart Resize. Templates package workflows for photo retouching, product photography, marketing assets and game art. The company also highlights consistent characters, locations and props across separate generations — a capability it frames as groundwork for later video production.

Why it matters

Human-preference leaderboards have been led by OpenAI's image stack for much of this year, with Reve AI, Microsoft and Chinese entrants such as Qwen-Image-3.0-Pro pressing from below. SpaceXAI took second on both boards at launch using its cheaper inference tier — but within three days the text-to-image order had shifted again, leaving it about 65 Elo points off the top and behind Microsoft's newest preview, while its 24-point editing deficit still stands. The catch for professional users is distribution: with no first-party API and developer access limited to a Vercel gateway preview, the model lives mainly inside Grok's consumer apps, so unlike Qwen's and ByteDance's image and video models — already wired into many third-party platforms — it is only beginning to reach production creative pipelines. Whether the ranking translates into real usage depends on how quickly that API arrives.

Sources