Gimlet Labs Raises $300M for Distributed AI Inference
Gimlet’s large Series B backs software that pools heterogeneous accelerators as inference buyers seek alternatives to fixed GPU clusters.
A large bet on separating inference from servers
Gimlet Labs has raised a $300 million Series B led by Andreessen Horowitz, with participation from investors and strategic backers including Arm, Microsoft’s M12, Samsung Ventures, Sapphire Ventures and Tiger Global. The company is building a disaggregated inference platform intended to combine processors and memory across machines instead of requiring every model workload to fit within a conventional, tightly coupled GPU server.
Gimlet says it has added billions of dollars in contracted revenue since March, assembled a gigawatt-scale data-center pipeline and is moving toward hundreds of megawatts of managed capacity. It did not publish customer-by-customer commitments, valuation, recognized revenue or a detailed delivery schedule, so those figures describe forward contracts and planned infrastructure rather than operating scale already reached.
The platform’s premise is that inference workloads contain different computational stages and do not always need identical accelerators. Separating prefill, token generation and other operations can allow operators to use a mix of hardware, reclaim stranded capacity and avoid sizing every server for the most demanding portion of a request. Strategic participation from Arm and Samsung underscores the opportunity for non-Nvidia processors and memory suppliers if heterogeneous inference becomes practical.
Execution will be difficult. Moving data between disaggregated components can introduce latency and networking costs, while serving large models reliably requires scheduling software to respond to rapidly changing demand. Gimlet must also convert contracted projects into installed capacity without being overwhelmed by construction and hardware financing.
Why it matters
The round is unusually large for a Series B because inference infrastructure is becoming a capital-intensive market in its own right. If Gimlet can make mixed hardware behave like a coherent pool, model providers could buy useful compute rather than standardized GPU boxes, improving utilization and widening the supplier base. The unresolved issue is whether its software savings remain meaningful after networking, power and financing costs are included at hundreds-of-megawatts scale.