LIVE
Etched ↑ $700M at $21B — valuation doubles in four weeks · Aug 18, 2026
Cerebras CS-4 + WSE-3 Turbo clock doubled to 2.8GHz, 250 PFLOPS · Aug 18, 2026
Groq ↓ $3.5B valuation $350M Series A recapitalisation · Aug 17, 2026
OLIX ↑ $312M at $3.3B — Europe's largest chip round · Aug 3, 2026
OpenAI Jalapeño custom inference ASIC with Broadcom · Jun 24, 2026
Nvidia Groq 3 LPU decode co-processor for Vera Rubin · GTC 2026
Cerebras Q2 core revenue $209.9M +103% · RPO $25.4B · Aug 12, 2026
Intel Crescent Island inference GPU, 160GB LPDDR5X, samples H2 · 2026
SambaNova ↑ $1B Series F at $11B — Intel buyout collapsed · Jul 8, 2026
d-Matrix Corsair in full production shipping in volume · Jun 2026
Cerebras IPO $5.55B raised at $185/share · May 14, 2026
Fractile ~$250M Anthropic chip deal reported; in talks at $6.5B · Aug 19, 2026
Rebellions ↑ $400M Pre-IPO round · March 2026
Nvidia licensed Groq inference tech for ~$20B · Dec 2025
Nvidia FY2026 revenue $215.9B · data centre $197.3B
Etched ↑ $700M at $21B — valuation doubles in four weeks · Aug 18, 2026
Cerebras CS-4 + WSE-3 Turbo clock doubled to 2.8GHz, 250 PFLOPS · Aug 18, 2026
Groq ↓ $3.5B valuation $350M Series A recapitalisation · Aug 17, 2026
OLIX ↑ $312M at $3.3B — Europe's largest chip round · Aug 3, 2026
OpenAI Jalapeño custom inference ASIC with Broadcom · Jun 24, 2026
Nvidia Groq 3 LPU decode co-processor for Vera Rubin · GTC 2026
Cerebras Q2 core revenue $209.9M +103% · RPO $25.4B · Aug 12, 2026
Intel Crescent Island inference GPU, 160GB LPDDR5X, samples H2 · 2026
SambaNova ↑ $1B Series F at $11B — Intel buyout collapsed · Jul 8, 2026
d-Matrix Corsair in full production shipping in volume · Jun 2026
Cerebras IPO $5.55B raised at $185/share · May 14, 2026
Fractile ~$250M Anthropic chip deal reported; in talks at $6.5B · Aug 19, 2026
Rebellions ↑ $400M Pre-IPO round · March 2026
Nvidia licensed Groq inference tech for ~$20B · Dec 2025
Nvidia FY2026 revenue $215.9B · data centre $197.3B

Latest News

All stories →
Funding

Fractile Raises $220M to Accelerate Next-Generation Inference Hardware

UK-based chip startup Fractile secured $220 million led by Accel, Factorial Funds and Founders Fund to bring its first purpose-built inference chips and systems to market. Founded on the premise that the world's most capable AI systems will be limited by output speed, Fractile is now among the best-capitalised inference hardware startups in Europe.

Funding

Groq Raises $650M to Pivot From Chip Maker to AI Inference Neocloud

Following Nvidia's $20B acquisition of its chip assets, Groq is repositioning as a managed inference cloud service provider under entirely new leadership.

IPO

Cerebras Prices IPO at $185/Share — $5.5B Raise Makes It 2026's Biggest Listing

Cerebras priced its Nasdaq debut at $185 per share, raising $5.55B, and the stock rose 68% on its first day — briefly lifting market capitalisation close to $100B. It is the first inference specialist to reach the public markets, and the only one whose numbers can now be checked against filings.

Funding

Rebellions Raises $400M Pre-IPO at $2.34B Valuation for Efficient Inference Chips

Seoul-based Rebellions — backed by South Korea's government 'K-Nvidia' initiative — is targeting 2026 IPO as energy-efficient inference chips attract sovereign capital.

2026 Funding Tracker

Full tracker →
Verified raises 2026 $5.3B+ Disclosed private rounds tracked below
Median round size $350M Up from $100M in 2024
Biggest round 2026 $1B Cerebras (Feb), SambaNova (Jul); Etched $700M (Aug)
Market size 2026 $117B Third-party projection → $312B by 2034
Etched
🇺🇸 United States
$700M
Growth

Led by Jane Street — also Etched's first customer — at a $21B valuation, doubling its July mark inside four weeks. Kleiner Perkins, Sequoia, a16z, Tiger Global and Bain Capital Ventures joined. Full profile →

18 August 2026
OLIX
🇬🇧 United Kingdom
$312M
Series B

Photonic inference chips at a $3.3B valuation — reported as Europe's largest semiconductor round. Arm, Hudson River Trading and the UK Sovereign AI fund participated. Full profile →

3 August 2026
Groq
🇺🇸 United States
$350M
Series A · recap

Led by Disruptive with planned Nvidia participation, at a $3.5B valuation — roughly half the $6.9B peak of September 2025. Framed as a reset for the post-licensing company. Full profile →

17 August 2026
Etched
🇺🇸 United States
$300M
Series C

Led by Sequoia at a $10.3B valuation, double its $5B December 2025 mark — and itself superseded four weeks later by the $700M round at $21B. Andreessen Horowitz, SK hynix and Jane Street participated. Full profile →

23 July 2026
SambaNova
🇺🇸 United States
$1B
Series F · first close

Led by General Atlantic at an $11B valuation, with T. Rowe Price and Capital Group. Follows a $350M Series E in February and the collapse of Intel's ~$1.6B acquisition approach. Full profile →

8 July 2026
Fractile
🇬🇧 United Kingdom
$220M
Series B

Led by Accel, Factorial Funds and Founders Fund at a ~$1B post-money valuation, with Conviction, Felicis, 8VC and existing backers Kindred Capital, the NATO Innovation Fund and Oxford Science Enterprises. In-memory compute in SRAM, with no HBM and no off-chip DRAM. Since reported: an initial ~$250M chip agreement with Anthropic, and advanced talks to raise ~$600M at a $6.5B pre-money valuation — neither confirmed by the company. Full profile →

28 May 2026 · updated 21 August 2026
Groq
🇺🇸 United States
$650M
Neocloud Pivot

Growth capital led by Disruptive and Infinitum, following Nvidia's ~$20B non-exclusive licence of Groq's LPU technology. Founder Jonathan Ross left for Nvidia in December 2025. Full profile →

22 June 2026
Rebellions
🇰🇷 South Korea
$400M
Pre-IPO

Mirae Asset and the Korea National Growth Fund at a $2.34B valuation, including ~$166M directly from South Korea's government. Cumulative funding ~$850M. KOSPI listing now targeted for H1 2027. Full profile →

March 2026
Cerebras
🇺🇸 United States
$1B
Pre-IPO Round

Then IPO'd at $185/share raising $5.5B — 2026's biggest public offering. The company markets its wafer-scale systems on single-stream token generation speed; published comparisons against GPU systems are its own.

Feb–May 2026
MatX
🇺🇸 United States
$500M
Growth Round

Transformer-native ASIC designed post-ChatGPT, built for LLM inference workloads from the ground up. One of three $500M rounds in 2026.

2026
D-Matrix
🇺🇸 United States
$275M
Series C

$2B valuation, $450M raised in total. Backed by Temasek, Qatar Investment Authority and Microsoft's M12. Corsair entered full production in June 2026. Full profile →

12 November 2025

Training builds the brain.
Inference makes it think.

AI training is a one-time event — a massive compute job that teaches a model everything it knows. But inference is the continuous, real-time process of that model answering your questions, writing your code, driving your car.

Every API call. Every agent action. Every token of output. That's inference. And it happens billions of times per day across every AI-powered product on earth.

The problem: the GPUs that won the training era weren't designed for this. They burn more power, cost more per token, and introduce more latency than the moment demands. That's why a new class of silicon — purpose-built inference chips — is attracting more capital than any hardware category in history.

Two-thirds of all AI compute in 2026 is inference. By 2027 it will be 80%. This isn't a niche. It's the entire delivery layer of AI.

Dimension
Training Chips
Inference Chips
Primary metric
FLOPS throughput
Tokens/second
Run frequency
Once per model
Billions/day
Latency need
Not critical
Sub-millisecond
Memory type
HBM (high bandwidth)
SRAM / in-memory
Power profile
700W+ per GPU
Optimised for efficiency
Revenue model
One-time capex
Recurring / per-call
2026 share
~33% of workloads
~66% of workloads

Market Landscape

Full map →

Purpose-Built ASICs

Chips designed from scratch for transformer inference. Etched, MatX and Fractile belong here — Fractile computes inside SRAM and uses no HBM at all — trading GPU flexibility for radical speed gains on LLM workloads. Etched has since broadened its design to a split prefill chip and cluster-scale decode memory, and now claims support beyond transformers.

LPUs & Tensor Processors

Language Processing Units (Groq's original architecture) and wafer-scale processors (Cerebras WSE-3) use massive on-chip SRAM to eliminate the memory bandwidth bottleneck that makes GPUs slow on sequential token generation. Groq's LPU technology was licensed to Nvidia in December 2025.

Inference Neoclouds

Companies like Groq (post-pivot) and SambaNova wrap proprietary silicon into fully-managed inference APIs. Customers pay per token, not per chip — the recurring-revenue model that investors find most compelling in 2026.

Photonic Computing

Two distinct bets. Ayar Labs uses light to move data between chips, raising $500M in March 2026 at a $3.75B valuation, led by Neuberger Berman with Nvidia, AMD, MediaTek and Alchip. OLIX goes further and computes optically, raising $312M at $3.3B in August. Nvidia put $4B into photonic networking the same month as the Ayar round — $2B each into Coherent and Lumentum — signalling interconnect as the next constraint.

Hyperscaler Custom Silicon

Google (TPU v7 "Ironwood"), AWS Trainium and Inferentia, Microsoft Azure Maia and Meta MTIA each build inference chips to cut cost and latency for their own workloads. Collectively they are the largest deployed inference silicon on earth — and increasingly they are specialising, with Google splitting its eighth generation into separate training and inference parts.

Edge Inference Chips

Running models locally on device — phones, vehicles, robotics — is a fast-growing tier. Qualcomm's AI Engine, Apple Neural Engine, and specialised automotive inference SoCs are embedding AI inference directly into billions of endpoints.

Company Profiles

All companies →

Sourced reference profiles covering chip families, inference role, leadership, disclosed funding, customers and the open questions on each company. Private-company revenue is recorded as undisclosed rather than estimated, and vendor performance claims are labelled as such. Each profile carries a review date and a corrections route.

Nvidia →

The default platform for inference at scale. FY2026 revenue of $215.9B with data centre revenue of $197.3B. Hopper, Blackwell and Rubin, plus the CUDA stack that most production workloads were written against.

AMD →

The only credible merchant second source. First-half 2026 data centre revenue of $12.5B, up 81%. MI400 series and Helios racks, with a 6 gigawatt OpenAI commitment beginning H2 2026.

Groq →

Built the most visible non-GPU inference architecture, then licensed it to Nvidia for ~$20B. Now an inference cloud and neocloud, valued at $3.5B — down from $6.9B.

Cerebras →

Wafer-scale processors, and the first inference specialist to reach the public markets. Launched the rack-scale CS-4 and an overclocked WSE-3 Turbo in August 2026. Q2 core revenue $209.9M, up 103%.

d-Matrix →

Digital in-memory compute, placing computation inside the memory holding the weights. Corsair reached full production in June 2026 — a stage most specialists have not.

Etched →

The most aggressive specialisation bet in the market, at a $21B valuation with over $1B in pre-booked orders — but still pre-production at scale, with every figure company-stated.

Intel →

The only player that designs accelerators and owns leading-edge fabs. Crescent Island bets on 160GB of LPDDR5X instead of HBM — cheaper, air-coolable, and outside the queue everyone else is standing in.

OpenAI →

Jalapeño, built with Broadcom on TSMC 3nm and taped out in nine months. Captive silicon that is not for sale — which is precisely why it matters to everyone who does sell inference chips.

Google TPU →

Seven shipped generations since 2015. Ironwood is generally available, the eighth generation splits into separate training and inference chips at 2nm, and Anthropic has committed to up to a million TPUs.

AWS →

Close to a million Trainium2 chips serve Anthropic's models, and Project Rainier runs 500,000 more on a 1,200-acre Indiana site. None of it is for sale — it exists to lower the cost of AWS.

Microsoft Maia →

Maia 200 is explicitly an inference chip: TSMC 3nm, 216GB of HBM3e, over 10 PFLOPS FP4 in 750W. Already serving OpenAI's production models inside Azure.

SambaNova →

Reconfigurable dataflow systems for enterprises and governments. Went from exploring a sale at a reported $1.6B to raising at $11B in eight months — the sharpest reversal in the market.

Rebellions →

South Korea's sovereign champion, backed by Samsung, SK hynix and the state. Rebel100 chiplets on Samsung 4nm, shipping as racks and pods, with a KOSPI listing targeted for H1 2027.

Fractile →

Computes inside SRAM with no HBM and no off-chip DRAM. Reportedly holds a ~$250M Anthropic supply agreement and is in talks at $6.5B — for silicon that does not ship until 2027.

OLIX →

London photonics, computing with light and skipping HBM altogether. $312M at $3.3B in Europe's largest semiconductor round, backed by Arm and the UK Sovereign AI fund.

Full company landscape → Glossary → Editorial methodology →

What's Next in Inference

M&A Wave

The Acquisition Sprint Is Not Over

Nvidia's ~$20B licence of Groq's inference technology — which also took Groq's founding leadership in-house — was the opening move, and it remains the defining transaction of the cycle. Intel's approach to SambaNova did not survive: the two signed a non-binding term sheet at a valuation of roughly $1.6B, the deal lapsed, and SambaNova subsequently raised $1B at an $11B valuation in July 2026, around seven times the mooted price. Intel retains a minority stake of roughly 9%, and its chief executive Lip-Bu Tan chairs SambaNova. The lesson buyers drew is that inference assets have repriced faster than acquirers moved, which makes the licence-and-hire structure Nvidia used look more repeatable than an outright purchase.

The other pressure on the M&A thesis is that the largest buyers are increasingly building instead. OpenAI's Jalapeño, announced with Broadcom in June 2026, joins Google, Amazon, Microsoft and Meta in captive inference silicon — removing from the market the very customers an acquirer would be buying access to.

IPO Pipeline

Rebellions, D-Matrix and More Are IPO-Ready

Cerebras' landmark IPO in May 2026 opened the door. Rebellions has explicitly targeted a late-2026 listing after closing its pre-IPO round. Multiple US-based inference startups are running IPO readiness processes. The public market appetite — with Cerebras 20× oversubscribed — has validated inference as a standalone investable category.

Architecture Shift

Photonics + In-Memory Compute Will Reshape the Stack

The next inflection point isn't a faster GPU — it's moving data with light (Ayar Labs, VSORA) and processing it where it lives (D-Matrix's in-memory compute). Nvidia's $4B photonics bet in early 2026 signals the incumbent sees it too. By 2027–28, hybrid optical-electrical inference racks could deliver 10× bandwidth improvements over current copper interconnects.

Sovereign AI

National Governments Are Building Inference Infrastructure

South Korea's $166M direct investment in Rebellions is the clearest signal yet. France, Germany, UAE, and Japan are all funding domestic inference chip programmes under sovereign AI strategies. The goal: reduce dependence on US-controlled Nvidia hardware for critical AI infrastructure. Expect $10B+ in government inference chip commitments globally by end of 2027.

Frequently Asked Questions

What is AI inference?

Inference is the step where a trained AI model generates output — answering a question, writing code, analysing an image. Unlike training (which happens once to build the model), inference runs continuously, billions of times a day, powering every AI product you use.

Why are inference chips different from standard GPUs?

GPUs were designed for graphics and later adapted for AI training — they excel at raw floating-point throughput across huge parallel workloads. Inference chips optimise for different goals: low latency (fast first token), high tokens-per-second, energy efficiency, and cost-per-token at massive scale. Different workload, different silicon.

What is an LPU?

A Language Processing Unit — coined by Groq — is a chip architecture built specifically for sequential token generation in large language models. Rather than the parallel batch processing of GPUs, LPUs optimise for the sequential, memory-bandwidth-limited nature of autoregressive LLM inference.

How much has the inference chip market raised in 2026?

InferenceChips.com tracks over $5.3 billion in individually verified disclosed private rounds between February and August 2026, across eleven rounds by inference chip companies. That is a floor rather than a market total: it counts only rounds we could confirm against a primary source, and excludes public offerings such as the Cerebras IPO. The median round size has grown from around $100M in 2024 to roughly $350M in 2026.

Is Nvidia still dominant in inference?

Nvidia remains the default platform, with FY2026 data centre revenue of $197.3B. It is also acting to close its main gap: in December 2025 it licensed Groq's inference technology in a deal reported at about $20B, and at GTC 2026 it unveiled the resulting Groq 3 LPU as a dedicated decode co-processor inside the Vera Rubin platform — an explicit acknowledgement that a GPU is not the ideal device for token generation. Specialists such as Cerebras publish single-stream speed comparisons showing far larger advantages, but those are vendor figures measured in the regime most favourable to them, and should be read alongside independent benchmarks.

What is inference-as-a-service?

Rather than selling chips, inference-as-a-service companies (Groq's new model, SambaNova, Baseten) run proprietary hardware in data centres and sell API access priced per token. Customers get faster inference without managing infrastructure. Investors like this model because it generates recurring revenue rather than lumpy hardware sales.

What does 'inference at the edge' mean?

Running AI models locally on end devices (phones, cars, robotics, wearables) rather than in a cloud data centre. Edge inference reduces latency to near-zero, works offline, and keeps sensitive data on-device. Qualcomm, Apple, MediaTek and a wave of automotive SoC startups are the key players here.

Why does inference efficiency matter for the environment?

AI inference now consumes a measurable fraction of global electricity — and that fraction is growing rapidly as AI usage scales. Purpose-built inference chips can deliver the same token output at 10–100× lower energy than general-purpose GPUs, making efficiency a financial, strategic and environmental priority simultaneously.

About InferenceChips.com

🎯

Editorial Independence

InferenceChips.com publishes independent analysis and news. We carry no affiliate links, sponsored placements, or advertiser relationships. Our coverage is shaped by newsworthiness, not commercial arrangements.

🔬

Primary Source Commitment

Every funding figure, valuation and technical claim on this site is sourced from primary disclosures, SEC filings, verified press releases, or named expert sources. We cite our sources and update when facts change.

Real-Time Coverage

The inference chip market moves fast. Our team monitors global funding announcements, regulatory filings, chip launches and analyst reports daily. The ticker and news sections are updated continuously.

🌐

Global Scope

From Silicon Valley to Seoul, London to Shanghai, the inference race is international. We track US, European, Asian and Middle Eastern actors with equal rigour — including sovereign AI programmes often missed by US-centric tech media.

Why InferenceChips.com? The AI conversation is dominated by training milestones — new model releases, benchmark records, parameter counts. But the real-world delivery of AI intelligence happens at inference, and the infrastructure being built right now to serve it will define the economics, geopolitics and capabilities of AI for the next decade. InferenceChips.com exists to cover that story with the depth and accuracy it deserves.