Advertisement

Arm (ARM) and Microsoft (MSFT) back Gimlet's $300 million bet on cheaper AI

Gimlet Labs logo on a black background

Key points

  • $300 million Series B at a $3 billion valuation
  • Arm and Microsoft's M12 join as new investors
  • Company reports billions in contracted revenue since March
  • Its own modeling showed 1.7x to 4x cost gains

Gimlet Labs has raised $300 million at a $3 billion valuation to expand software that divides AI inference work among different types of chips. Andreessen Horowitz led the Series B, announced Friday, September 4, with Arm (ARM) and Microsoft's (MSFT) venture arm, M12, joining as new investors.

The company says it has added billions of dollars in contracted revenue since March. Its last disclosed annualized revenue figure was about $10 million as of that month, according to a May interview with Chipstrat. The figures measure different things, but the gap puts attention on the contracts behind Gimlet's expansion and the economics of delivering them.

The financing brings Gimlet's total funding to $392 million, less than six months after its $80 million Series A on March 23. Sapphire Ventures also joined the new round, alongside returning investors Menlo Ventures and Factory. The full list in the announcement runs to 18 names, including Samsung Ventures, Tiger Global Management, Hudson River Trading and XTX Markets. Gimlet's four founders, chief executive Zain Asgar, Michelle Nguyen, Omid Azizi and Natalie Serrino, previously built Pixie, an observability tool for Kubernetes that New Relic bought in December 2020.

Gimlet calls the product a multi-silicon inference cloud: software that breaks the work of answering a request into stages and runs each stage on the chip best suited to it. Besides the contracted revenue, the company's own announcement claims "gigawatts of datacenter pipeline" and says it is "quickly scaling to hundreds of megawatts in managed capacity." Bloomberg's Dina Bass, who interviewed Asgar, reported that Gimlet is working with Arm to make its software run on Arm-based chips. Arm rose 3.9 percent Friday to $252.10. Microsoft fell 2.0 percent to $499.68.

Advertisement

Does splitting inference across chips pay?

Answering a prompt has two phases. Prefill reads the prompt, and it is heavy on raw compute. Decode writes the answer one token at a time, and generating each token means reading the model's weights, which makes decode sensitive to memory bandwidth. Gimlet's argument is that different chips may be better suited to each phase, and that running both on the same GPU leaves part of the chip idle in each phase.

In an analysis it published in October 2025, Gimlet benchmarked individual H100, B200 and Gaudi 3 accelerators and fed those measurements into performance models. It then modeled a configuration pairing an Nvidia (NVDA) B200 for prefill with an Intel (INTC) Gaudi 3 for decode, and compared it with running both phases on H100s and with running both on B200s.

The headline figure was a 1.7x improvement in modeled total cost of ownership. For a prompt-heavy workload of 4,096 input tokens and 512 output tokens, the paper reports a 3x cost benefit against the H100 baseline and says the mixed pair also beats the B200-only configuration. For an answer-heavy workload of 512 in and 4,096 out, it reports as high as 4x against H100s, while B200s alone modeled at about 2.5x on the same workload. Some of the gain against H100s reflects newer hardware, so the B200-only figure is the closer comparison, and the difference between 2.5x and 4x is what the model attributes to using Gaudi 3 for decode.

Those results come with qualifications. It's a model, not a deployed system, and the paper says its goal was to compare disaggregation strategies, so its results are relative multipliers rather than dollars per token. The split also has a cost of its own: after prefill finishes, the working memory called the KV cache has to move from one chip to the other. Gimlet's paper puts the format translation at 20 to 50 microseconds and the optimized transfer at 5 to 10 milliseconds, and it says "it's challenging to disaggregate the same workload across multivendor accelerator types" because the software stacks differ by vendor.

Some accelerators Gimlet is exploring for decode have limited onboard memory. In a March post on chips from Cerebras and d-Matrix, which keep model weights in fast on-chip SRAM instead of external memory, the company notes that even Cerebras's wafer-scale chip holds 44 gigabytes, "so weights for large models need to be distributed across multiple chips." Gimlet's answer is to use those chips for smaller draft models paired with a frontier model on GPUs. In the Chipstrat interview, the team said hardware amortization runs to about 70 percent of an inference provider's annual costs, and listed the hard parts as KV cache transfer between different systems, network latency between racks, and scheduling across several hardware platforms at once. The company lists Nvidia, AMD (AMD), Intel, Arm, Cerebras and d-Matrix as supported chip partners.

This is different from a router that sends whole requests to whichever machine is free, the design we described in Nvidia's PAIR tool last week. Gimlet's software divides one request between chips by phase and moves the request's state between them. The gain it claims comes from matching each phase to hardware built for it, not from concurrency.

Advertisement

What contracted revenue means

Contracted revenue generally refers to future business under customer agreements, but Gimlet hasn't disclosed its calculation. It cannot automatically be equated with accounting revenue, remaining performance obligations or another company's backlog. As an illustration only, a three-year contract worth $300 million would count as $300 million of contracted revenue on the day it is signed and, assuming service is delivered evenly over three years, as $100 million of revenue a year. We covered the same distinction when Nscale pitched $103 billion in contracted revenue ahead of its IPO.

Gimlet's own disclosures give the scale of the gap. When the company came out of stealth in October 2025, TechCrunch's Julie Bort reported it had eight-figure revenue, meaning at least $10 million, and 30 employees. In the May interview with Chipstrat, the team put annualized revenue at $10 million as of the March raise. The March release said the customer base had tripled and included "one of the top three frontier labs" and "one of the top three hyperscalers," neither named.

DateWhat Gimlet reportedSource
October 2025Eight-figure revenue at launchTechCrunch
March 2026$10 million annualized revenue; $80 million Series AChipstrat interview, company release
September 2026Billions in contracted revenue added since March; $300 million Series BCompany announcement

The contracted figure is at least 100 times Gimlet's last disclosed annualized revenue figure. That comparison illustrates scale rather than revenue growth: one is a yearly run rate, while the other spans undisclosed contract terms. Gimlet says the funding will support its cloud operations and team expansion.

Three things the announcement doesn't say: how long the contracts run, whether they carry cancellation rights, and how much of the contracted value depends on Gimlet delivering the hundreds of megawatts of capacity it says it is scaling toward.

Advertisement

Frequently asked questions

Who invested in Gimlet Labs' $300 million Series B?

Andreessen Horowitz led the round, announced September 4, 2026, at a $3 billion valuation. Arm (ARM), Microsoft's venture fund M12 and Sapphire Ventures invested for the first time, and Menlo Ventures and Factory returned. The full list runs to 18 investors including Samsung Ventures, Tiger Global and Hudson River Trading. Total funding is $392 million.

What does Gimlet Labs' software actually do?

It splits the work of answering an AI prompt into phases and runs each phase on a different kind of chip: the compute-heavy prefill phase on a GPU such as an Nvidia (NVDA) B200, and the memory-bound decode phase on another accelerator such as an Intel (INTC) Gaudi 3 or an SRAM-based chip from Cerebras or d-Matrix. The request's working memory, the KV cache, is moved between chips. Gimlet's own October 2025 analysis, which modeled configurations from benchmarks of individual chips rather than a deployed system, reported a 1.7x cost improvement against running both phases on Nvidia H100s, with up to 4x on answer-heavy workloads.

What does contracted revenue mean, and how much revenue does Gimlet Labs have?

Contracted revenue generally refers to future business under customer agreements, but Gimlet has not disclosed its calculation, so it cannot be equated with accounting revenue, remaining performance obligations or another company's backlog. Gimlet says it has added billions in contracted revenue since March 2026 but does not define the term or disclose contract lengths. Its last disclosed annualized revenue figure was about $10 million as of March 2026, according to a May interview with Chipstrat.

How did Arm (ARM) and Microsoft (MSFT) stock react?

Arm rose 3.9% to $252.10 on Friday, September 4, 2026, and Microsoft fell 2.0% to $499.68. Both companies' investments in Gimlet are venture stakes that are small relative to their size, and the moves reflect the broader market session rather than the announcement. This is general information, not investment advice.

More on AMD and ARM

Dennis Singleton
Dennis Singleton

Dennis Singleton has spent years following the markets, but what keeps his attention is how AI is built. He writes about the companies behind the technology, from semiconductor designers and advanced packaging to photonics, memory, networking, and the hardware powering modern AI. His approach starts with filings, earnings, and industry research, then translates the important details into clear, straightforward analysis without unnecessary hype.