Pull to refresh
Logo
IBM releases Granite Speech 5.0 with record transcription speed

IBM releases Granite Speech 5.0 with record transcription speed

New Capabilities

The 470M-parameter encoder-only models drop translation and keyword biasing for roughly 20x faster English transcription.

August 25th, 2026: IBM releases Granite Speech 5.0 TurboCTC

Overview

Updated Aug 26

IBM says two new models turn spoken English into text more than twelve thousand times faster than real time — about 3.5 hours of audio per second on one Nvidia H200 GPU. Each compact Granite Speech 5.0 model packs just 470 million parameters.

The speed comes from a blunt design choice. IBM removed the language-model decoder that gave earlier Granite Speech releases translation and keyword biasing, leaving an encoder-only pipeline that is over 20 times faster than before. The headline figures are vendor-reported and await independent leaderboard confirmation.

Why it matters

Real-time transcription on a laptop or phone, with no cloud round-trip and no per-minute fees — if the claimed speed survives independent testing.

Questions about this story

Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.

No questions yet — be the first to ask.

Key Indicators

12,600 RTFx
Claimed throughput on one NVIDIA H200
Batched inference; audio flows through the models 12,600x faster than real time.
470M
Parameters per model
Two variants: an Apache 2.0 licensed model and a noncommercial CC-BY-NC-SA one.
4.85% / 5.00%
Aggregate word error rate
Noncommercial model 4.85%, Apache-2.0 model 5.00% on OpenASR public test sets.
20x
Speedup over prior Granite Speech models
Throughput gain from removing the LM decoder; prior 4.1 NAR variant ran at roughly 1,820 RTFx.

Voices

Curated perspectives — historical figures and your fellow readers.

Ever wondered what historical figures would say about today's headlines?

Sign up to generate historical perspectives on this story.

People Involved

Organizations Involved

Timeline

1 event Latest: August 25th, 2026 · 3 weeks ago
  1. IBM releases Granite Speech 5.0 TurboCTC

    Latest Product release

    IBM launched two 470M-parameter English ASR models claiming 12,600 RTFx on an H200. Official OpenASR Leaderboard rankings were still pending at publication.

Scenarios

1

Granite Speech 5.0 tops its own leaderboard with claims confirmed

Likely Resolves by End of 2026

Discussed by: Unite.AI noted the official OpenASR rankings had not yet absorbed the new models at launch.

IBM runs the OpenASR Leaderboard tooling used to score the models and says the WER and RTFx figures are expected to match official results. If the leaderboard absorbs the models with word error rates near the claimed 4.85–5.00% and throughput above 10,000 RTFx, the vendor-reported numbers become official.

2

IBM restores translation and keyword biasing in a follow-up

Possible Resolves by Q2 2027

Discussed by: IBM's release notes and RuntimeWire coverage of the decoder removal.

The encoder-only design sacrifices capabilities the 4.1 models had: speech translation and keyword-prompted recognition. IBM may ship a 5.1 or hybrid variant that restores those features at higher speed than before, widening the models beyond pure transcription.

3

Device-level throughput lags the datacenter claim

Uncertain Resolves by Q2 2027

Discussed by: Unite.AI cautioned that the 12,600 RTFx figure is batched inference on an H200, not a laptop WebGPU demo.

The streaming WebGPU demo on consumer laptops will run slower than the headline number. If independent tests show on-device throughput far below the claim, the marketing value drops even if the datacenter figure holds.

Historical Context

3 moments from history that rhyme with this story — and how they unfolded.

2017–2019

Mozilla DeepSpeech (2017)

Mozilla shipped DeepSpeech, an open-source, connectionist-temporal-classification (CTC) speech recognizer designed to run on phones and low-power hardware without a cloud connection.

Then

Became a reference for on-device, self-contained speech recognition.

Now

Demonstrated that encoder-only CTC models could deliver practical, offline transcription.

Why this matters now

IBM's encoder-only design revives that same philosophy with token-level output and claims of far higher throughput.

November 2023

Distil-Whisper (2023)

Hugging Face distilled OpenAI's Whisper large-v2 into a 756-million-parameter model that ran 6x faster and was 51% smaller while staying within 1% word error rate on most benchmarks.

Then

Showed that distillation could shrink a massive speech model while keeping accuracy nearly intact.

Now

Set the expectation that speed and size gains did not have to cost much accuracy.

Why this matters now

IBM's 5.0 release pushes the same speed-versus-accuracy trade-off far harder — a 20x throughput jump instead of 6x.

Late 2024

NVIDIA Parakeet (2024)

NVIDIA released Parakeet TDT models through its OpenASR toolkit — non-autoregressive, edge-focused speech recognition models that climbed to the top of open ASR leaderboards.

Then

Gave developers a fast, compact open alternative to Whisper-class models.

Now

Made the OpenASR and FFASR leaderboards the de facto benchmark arena for compact ASR.

Why this matters now

Granite Speech 5.0 enters that same arena, ranking among the fastest two models on the FFASR Leaderboard at launch.

Sources

(5)