Pull to refresh
Logo
The race to build non-Nvidia AI inference chips

The race to build non-Nvidia AI inference chips

Money Moves

Startups keep raising billions on custom inference silicon — and Nvidia just spent $20 billion to buy its way into the same market.

July 23rd, 2026: AMD and Cerebras strike inference-sharing partnership

Overview

Updated Aug 19

Nvidia still runs roughly four of every five AI chips sold today. In December 2025, Nvidia paid $20 billion to license rival Groq's chip design and hire its founding team. It then built its own inference-only rack from that technology, launched at Nvidia's GTC conference in March 2026.

Cerebras went public on Nasdaq in May 2026, raising $5.55 billion. Etched's valuation jumped to $21 billion in under a month, and it shipped its first chip to Jane Street on August 18. AMD is now backing both sides, putting up to $5 billion into Anthropic and striking an inference deal with Cerebras.

Why it matters

If Nvidia can absorb its biggest inference-chip rival for $20 billion without a formal merger review, it can buy its way out of competition.

Questions about this story

Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.

No questions yet — be the first to ask.

Key Indicators

$20B
Nvidia's Groq licensing deal
Nvidia paid $20 billion in December 2025 for a non-exclusive license to Groq's inference chip design and hired much of its team.
$56.4B
Cerebras IPO valuation
Implied valuation when Cerebras priced its Nasdaq IPO at $185 a share in May 2026, raising $5.55 billion.
$21B
Etched valuation
Etched's valuation after an August 2026 round led by Jane Street, doubling its worth in about a month.
$5B
AMD's investment in Anthropic
AMD committed up to $5 billion to Anthropic in July 2026, tied to a 2-gigawatt order of AMD's Instinct MI450 chips.
$220M
Fractile Series B
Round led by Accel, Factorial Funds and Founders Fund that funds chip tape-out toward a 2027 launch.
~80%
Nvidia data-center AI share
Industry estimate of Nvidia's share of AI accelerator revenue, the share every challenger in this story is chasing.

Voices

Curated perspectives — historical figures and your fellow readers.

Ambrose Bierce

Ambrose Bierce

(1842-1914) · Gilded Age · wit

Fictional AI pastiche — not real quote.

"Nine-figure fortunes staked on the certainty that the king of a hill will presently be displaced — this is not investment but rather the ancient ritual of ambitious men paying to watch other ambitious men fail. That four in five chips bear one maker's mark is called a monopoly by the envious and an ecosystem by the beneficiary; that investors now wager two hundred millions on the word "inference" suggests they have mastered the vocabulary of the future without troubling themselves to understand it."

Ever wondered what historical figures would say about today's headlines?

Sign up to generate historical perspectives on this story.

People Involved

Organizations Involved

Fractile
Fractile
AI semiconductor startup
Reported in early talks with Anthropic to supply AI inference chips as of May 2026, alongside its 2027 commercial launch target

London-based startup designing chips that run only AI inference, not training.

Accel
Accel
Venture capital firm
Lead investor in Fractile Series B

Global venture firm with a long track record in infrastructure and semiconductor bets.

Founders Fund
Founders Fund
Venture capital firm
Co-lead investor in Fractile Series B

Peter Thiel's venture firm, an active backer of deep-tech and contrarian infrastructure bets.

Anthropic
Anthropic
AI Company
Reported in early talks with Fractile in May 2026, then signed a deal in July 2026 in which AMD will invest up to $5 billion and Anthropic will deploy 2 gigawatts of AMD chips starting in 2027

Frontier AI lab whose appetite for non-Nvidia inference compute helps drive the funding race.

Nvidia Corporation
Nvidia Corporation
Semiconductor and AI computing company
Now selling inference-specific silicon too, after paying $20 billion in December 2025 to license Groq's chip design and hire its team; launched the Groq 3 LPX inference rack in March 2026

Supplier of the GPUs and software stack that run most AI training and inference today.

Groq Inc.
Groq Inc.
AI Chip Startup
Licensed core chip technology to Nvidia for $20B in December 2025; says it still runs GroqCloud independently

Inference chip maker whose Language Processing Unit design was licensed to Nvidia in a deal that reshaped the market it was built to challenge.

Cerebras Systems
Cerebras Systems
AI semiconductor company
Completed its Nasdaq IPO in May 2026; now partnered with AMD on inference workloads

Maker of the wafer-scale engine, now publicly traded and splitting AI inference work with AMD.

Etched
Etched
AI semiconductor startup
Valuation reached $21B in August 2026 after back-to-back rounds; shipped its first chip to Jane Street

San Jose startup building the Sohu chip, hardware specialized only for transformer-architecture models.

Advanced Micro Devices
Advanced Micro Devices
Semiconductor Company
Backing both Anthropic and Cerebras with investment and hardware deals struck in July 2026

Chipmaker that has aligned itself with Anthropic and Cerebras against Nvidia's dominance of AI compute.

Timeline

May 2016 July 2026

15 events Latest: July 23rd, 2026 · 2 months ago Showing 8 of 15
Tap a bar to jump to that date
  1. AMD and Cerebras strike inference-sharing partnership

    Latest Partnership

    AMD and Cerebras agree to split AI inference workloads across their systems, with AMD hardware handling prompt processing while Cerebras chips handle token generation.

  2. Etched raises $300M Series C at $10.3B valuation

    Funding

    Etched closes a $300 million round led by Sequoia and Andreessen Horowitz, doubling its valuation to $10.3 billion in about seven months.

  3. AMD invests up to $5B in Anthropic

    Investment

    AMD says it will invest up to $5 billion in Anthropic, which agrees to deploy 2 gigawatts of AMD's Instinct MI450 chips starting in the first half of 2027.

  4. Cerebras completes $5.55B Nasdaq IPO

    IPO

    Cerebras Systems begins trading on Nasdaq as CBRS after pricing its IPO at $185 a share, raising $5.55 billion in one of the largest US tech listings of the year.

  5. Fractile closes $220M Series B

    Funding

    London-based Fractile raises $220M led by Accel, Factorial Funds and Founders Fund. Capital funds chip tape-out and software ahead of a 2027 commercial launch targeting 25x faster, ~90% cheaper frontier inference.

  6. Anthropic reportedly in early talks with Fractile

    Customer deal

    Anthropic is reported to be in early discussions to buy inference chips from Fractile, whose design keeps memory and compute on the same chip.

  7. Senators question Nvidia-Groq deal on antitrust grounds

    Regulatory

    Senators Elizabeth Warren and Richard Blumenthal send Nvidia a letter asking whether the Groq deal was structured to avoid formal merger review.

  8. Nvidia launches Groq 3 LPX inference rack

    Product launch

    At GTC 2026, Nvidia unveils its first non-GPU inference product, built on licensed Groq technology, claiming large throughput gains over its own Blackwell GPUs for serving large models.

  9. Nvidia pays $20B to license Groq's chip technology

    Acquisition

    Nvidia agrees to pay Groq $20 billion for a non-exclusive license to its Language Processing Unit design and hires much of Groq's engineering team, including CEO Jonathan Ross. Groq says it continues to operate independently as GroqCloud.

  10. Anthropic, Amazon deepen Trainium partnership

    Customer deal

    Amazon and Anthropic announce an expanded deal that puts more Claude inference on Amazon's Trainium chips. It is the clearest signal yet that frontier labs are willing to move workloads off Nvidia.

  11. Cerebras files for IPO

    Corporate filing

    Cerebras Systems files S-1 paperwork to go public, the first major Nvidia challenger to attempt the public markets. The offering is later delayed by regulatory review.

  12. Groq raises $640M at $2.8B valuation

    Funding

    Groq closes a major growth round led by BlackRock to scale its inference cloud and Language Processing Unit chips.

  13. Etched raises $120M for transformer-only chip

    Funding

    Etched closes a $120M round to build Sohu, a chip designed to run only transformer-architecture models. The bet: betting on a single architecture buys huge efficiency gains.

  14. Fractile founded in London

    Company formation

    Walid Mehri co-founds Fractile to build chips designed for AI inference, drawing on neural-network hardware research from Oxford.

  15. Google reveals the TPU at I/O

    Technology

    Google announces it has been running custom AI silicon, the Tensor Processing Unit, in production. It is the first public proof that purpose-built chips can outperform GPUs on AI workloads.

Scenarios

1

Fractile ships its first commercial chip on schedule

Possible Resolves by End of 2027

Discussed by: Bloomberg, Tech.eu, DatacenterDynamics coverage of the Series B

Fractile completes tape-out in 2026, manufactures silicon in 2027, and ships its first chip to at least one paying customer before year-end 2027. The full 25x speed and ~90% cost claims may not survive contact with production, but a credible launch with real customers keeps the company on the trajectory its investors are paying for.

2

Major frontier lab signs a nine-figure non-Nvidia inference deal

Possible Resolves by End of 2026

Discussed by: The Information, Bloomberg, Reuters reporting on AI compute procurement

Anthropic, OpenAI, Google, Meta or another frontier lab signs a publicly disclosed multi-year inference deal worth $100M or more with a non-Nvidia chip company, beyond existing hyperscaler-internal silicon. The Amazon-Anthropic Trainium expansion shows the path; the question is whether a comparable deal lands with one of the startup challengers.

3

Hyperscaler acquires a major inference chip startup

Uncertain Resolves by End of 2026

Discussed by: The Information, Bloomberg M&A reporting

AWS, Google, Microsoft, Meta or Oracle buys one of the well-funded inference chip startups to bring its design in-house. Acquirers get a finished design team and IP without paying the IPO premium; the startup gets an exit without proving its commercial model. This pattern repeats from the prior wave of AI hardware deals.

4

Nvidia data-center share holds above 75% through 2026

Likely Resolves by Apr 30, 2027

Discussed by: Gartner, IDC analyst notes; Bernstein, Morgan Stanley research

Despite the funding wave, Nvidia retains more than three-quarters of AI accelerator revenue through calendar 2026. Most challengers are still pre-product or sub-scale, software lock-in slows migration, and Blackwell-generation supply meets the marginal demand. The challenger story remains a 2027-2028 story rather than a 2026 one.

5

At least one well-funded inference startup folds or fire-sells

Possible Resolves by End of 2027

Discussed by: The Information, Financial Times venture coverage

One of the inference chip companies that has raised $100M or more shuts down, sells for less than total capital raised, or returns capital to investors. Inference silicon is capital-intensive and timing-sensitive; not every funded entrant will reach a commercial chip. A high-profile failure would reset valuations across the category.

6

FTC or DOJ opens a formal probe into the Nvidia-Groq deal

Uncertain Resolves by End of 2026

Discussed by: Senators Warren and Blumenthal; FTC Chair Andrew Ferguson

Nvidia structured its Groq deal as a technology license plus a talent hire rather than an acquisition, which let it skip standard merger review. Warren and Blumenthal have asked the FTC and DOJ to scrutinize this pattern across the tech industry. A formal probe would test whether the acquihire structure holds up as a way to buy out a competitor.

7

GroqCloud survives as an independent service

Possible Resolves by End of 2026

Discussed by: Constellation Research and The Motley Fool coverage of the Nvidia-Groq deal

Groq says it remains an independent company running GroqCloud even after Nvidia hired its founder and much of its engineering team. Whether that holds through the rest of 2026, or GroqCloud quietly gets folded into Nvidia's own cloud offering, is still an open question.

Historical Context

3 moments from history that rhyme with this story — and how they unfolded.

1995-2002

The 1990s graphics chip wars

Through the late 1990s, a crowded field of graphics chip makers, including 3dfx, ATI, Matrox, S3, Trident and a young Nvidia, fought to define the PC 3D graphics market. Each company pitched a different architecture and a different bet on what gamers and developers would adopt.

Then

Pricing collapsed, marginal players failed, and the market consolidated faster than investors expected. 3dfx, the early leader, went bankrupt by 2002.

Now

Two winners, Nvidia and ATI (later AMD), emerged with durable share. The lesson: in chip categories with high R&D costs and software lock-in, late-cycle consolidation is brutal and most well-funded entrants do not survive.

Why this matters now

Today's inference chip field looks structurally similar: many funded entrants, competing architectures, no clear winner, and a software moat held by the incumbent. The 1990s suggest the next five years will end with two or three survivors, not ten.

May 2016

Google launches the TPU (2016)

At Google I/O, Google revealed it had been running a custom AI chip called the Tensor Processing Unit in its data centers since 2015. It was the first time a major operator publicly claimed that purpose-built silicon could beat Nvidia GPUs on AI workloads at scale.

Then

The TPU opened a credible alternative path for AI compute and validated the thesis that workload-specific chips could compete with general-purpose GPUs.

Now

TPUs became central to Google's AI infrastructure and inspired a generation of custom-silicon efforts at Amazon (Trainium, Inferentia), Microsoft (Maia) and Meta (MTIA). It also made the inference chip startup category investable.

Why this matters now

Every Fractile, Groq and Etched pitch deck traces back to the TPU's central claim: GPUs are not the right shape for AI, and a purpose-built chip can win. The 2026 funding race is the venture-backed extension of that 2016 idea.

November 2019

Amazon launches AWS Inferentia (2019)

Amazon unveiled Inferentia, a custom inference chip built in-house for AWS, and later added Trainium for training. The chips were aimed at lowering Amazon's own compute costs and offering customers a cheaper alternative to Nvidia inside AWS.

Then

Inferentia gained limited adoption initially as customers stuck with familiar GPU tooling.

Now

By the mid-2020s, Trainium and Inferentia were central to AWS's AI pitch and underpinned the deeper Amazon-Anthropic partnership announced in 2024. Custom hyperscaler silicon proved viable at scale.

Why this matters now

Hyperscalers can and do build their own inference chips. That sets a ceiling on how much of the inference market the independent startups can address: the biggest buyers may bring the workload in-house rather than buy from Fractile or Groq.

Sources

(12)