Grok Imagine Image 2.0 debuts at No.2 in both arenas
xAI's new image model entered the Arena leaderboards at No.2 for both text-to-image and image editing, trailing only OpenAI's GPT-Image-2.
The result
Arena, the crowdsourced model-comparison platform, published leaderboard placements for xAI's newest image model on Saturday morning UTC. Grok Imagine Image 2.0, tested in its faster "Low" configuration, entered the Text-to-Image Arena in second place with 1,320 points, behind OpenAI's GPT-Image-2 at 1,380. In the Image Edit Arena it also landed second, scoring 1,439 against GPT-Image-2's 1,463 — a jump from seventh place for the previous generation, Grok Imagine Image Quality. Ranked below the new xAI entry in text-to-image were Reve AI's Reve 2.1 at 1,302, Meta's Muse-Image at 1,283, Google's Gemini 3.1 Flash Image at 1,263, Alibaba's Qwen-Image-3.0-Pro at 1,258 and ByteDance's Seedream 5.0 Pro at 1,257. The xAI score is still provisional, drawn from roughly 2,700 votes against nearly 70,000 for the leader.
What shipped
xAI released Image 2.0 the previous day as the "Quality Mode" of Grok Imagine, available on grok.com/imagine and in the company's iOS and Android apps. API access has not opened; xAI says it is coming. The company describes the model as designed to follow detailed instructions precisely, plan typography and page layout the way a designer would, and hold small text legible inside dense composite visuals.
The release bundles an editing toolkit rather than a bare generator: a magic-wand tool that alters only the region a user points at, segmentation-based selection, background removal with transparent export, multi-reference editing that accepts up to five input images in a single generation, and aspect-ratio conversion via Smart Resize. Templates package workflows for photo retouching, product photography, marketing assets and game art. xAI also highlights consistent characters, locations and props across separate generations — a capability it frames as groundwork for later video production.
Why it matters
Human-preference leaderboards have been led by OpenAI's image stack for much of this year, with Reve AI and Chinese entrants such as Qwen-Image-3.0-Pro pressing from below. A simultaneous No.2 placement in both generation and editing puts xAI within roughly 60 Elo points of the top in text-to-image and 24 in editing, using its cheaper inference tier. The catch for professional users is distribution: the model is confined to Grok's consumer apps with no API, so unlike Qwen's and ByteDance's image and video models — already wired into third-party platforms — it cannot yet enter production creative pipelines. Whether xAI's ranking translates into real usage depends on how quickly that API arrives.