⚡ Uncle Cat AI Radar
IndustryOpen SourceAgents

Ollama Replaces Usage Windows With Per-Token Pricing

Ollama has replaced opaque GPU-time limits with published token rates and monthly credits across its paid cloud plans.

From time windows to measurable consumption

Ollama has introduced per-token pricing for new Pro, Max and Team subscriptions, replacing the GPU-time allowances and five-hour or weekly usage limits that made consumption difficult to predict. Existing subscribers may retain their current plans or opt into the new structure.

The $20 monthly Pro plan includes $60 in usage credits, while the $100 Max plan includes $300. A newly available Team plan costs $500 per month and provides a shared $1,000 credit pool for unlimited users. Credits refresh each month and do not roll over. Once the included pool is exhausted, customers may continue at the same published token rates rather than waiting for a usage window to reset.

Ollama also expanded its free plan. It now supplies a small monthly allowance for selected starter models, while users can purchase credits to reach the wider catalogue without subscribing. The company says it adds no service fee and displays the cost of each request in the user’s account.

Ollama becomes a more conventional cloud provider

The change pushes a product best known for running open models locally further into the competitive inference market. The cloud service supports coding agents including Claude Code and Codex, as well as direct API access. Ollama says requests run on dedicated infrastructure in the United States and Europe, with Singapore hosting available for a limited selection of Qwen models. It also promises zero data retention and says prompts are neither logged nor used for training.

This matters because long-running coding agents can generate highly variable workloads. Token metering makes model and provider costs easier to compare, budget and audit than an allowance expressed in GPU time. However, the headline credit multiples—three times the subscription price for Pro and Max, and twice for Team—only represent genuine value if Ollama’s published token rates remain competitive. The new system improves predictability; the model-by-model price table will determine whether it also improves economics.

Sources