Ideogram and Pruna ship $0.003-per-image model line
P-Image-Ideogram launches with four quality tiers, native 1K and 2K output and three-to-eight-second latency, aimed squarely at production image pipelines.
Ideogram introduced P-Image-Ideogram on Thursday, a family of text-to-image models co-developed with inference-optimisation startup Pruna AI and positioned around cost and speed rather than peak quality alone.
What shipped
The family offers four quality modes — very low, low, medium and high — letting developers trade fidelity against price within a single API surface. Generation runs natively at 1K and 2K resolution across a wide range of aspect ratios, with latency reported between three and eight seconds and pricing starting at $0.003 per image. The models support JSON-structured prompting and layout control, features aimed at programmatic pipelines rather than interactive creative sessions. Ideogram highlighted photorealism, text rendering and style range in its launch examples, with typography historically the company's strongest differentiator.
Ideogram's central claim is a Pareto argument: on DesignArena's blind head-to-head human-vote image leaderboard, P-Image-Ideogram advances the frontier on both quality-versus-cost and quality-versus-speed, meaning no listed competitor is simultaneously cheaper and better, or faster and better. The company did not claim the top absolute quality position.
Distribution went wide on day one. Beyond Ideogram's own API, the models are live on ComfyUI, Replicate, Runware, Cloudflare, Leonardo AI, Picsart and Magnific. Replicate and other partners posted availability within hours of the announcement. The launch follows Ideogram's recent partnership with AMD to serve its open-weight Ideogram 4.0 model on Instinct accelerators, suggesting a broader push to control its own inference economics.
Why it matters
Image generation has largely stopped competing on whether a model can render a scene and started competing on what a rendered scene costs at volume. At $0.003 per image with sub-ten-second latency, the calculus shifts for applications that generate thousands of assets per day — e-commerce catalogues, ad variant testing, game and app asset pipelines — where per-image cost, not marginal aesthetic quality, sets the ceiling on what can be built. The tie-up with Pruna also signals that specialist compression and inference vendors are becoming co-authors of frontier consumer models rather than downstream optimisers.