September 24, 2026 · 6 min read

Multiverse Computing, HPE and Intel run enterprise speech-to-text on CPUs alone

Our CompactifAI-optimized version of Whisper doubles throughput and transcribes about 17,000 hours of audio a day on a single standard server, with no discrete accelerators and no significant loss in quality

Multiverse Computing, together with HPE and Intel, has run enterprise-grade speech-to-text entirely on standard server CPUs, transcribing about 17,000 hours of audio a day on a single machine with no GPUs and no loss in quality.

Transcription is one of the highest-volume AI workloads inside a large organisation, and for years it has been treated as a job that needs dedicated accelerators. This work shows it does not. Our optimized version of Whisper carries that workload on the general-purpose servers companies already run, on Intel Xeon 6 processors inside an HPE ProLiant server, using the CPUs alone, while the audio never leaves the customer's own infrastructure.

See it in action

What we built

We took Whisper Large V3 Turbo, one of the most widely deployed speech-to-text models, and produced our own optimized version of it with CompactifAI, our proprietary AI model optimization technology. The result is a Multiverse Computing model with half the parameters and half the memory of the original, 0.4 billion parameters against 0.8 billion and 0.75 GB against 1.5 GB, and no meaningful loss in quality.

Because the model needs far less infrastructure, it does not need specialised silicon to run well. We deployed it on a single HPE ProLiant DL380 server with two Intel Xeon 6 processors, using Intel Advanced Matrix Extensions and served through vLLM. Inference runs on the CPUs alone, with no discrete accelerators anywhere in the system. It is a standard enterprise server doing a job the industry assumed required a GPU.

The results

We benchmarked our optimized model against the uncompressed original on the very same server, so the comparison isolates the model rather than the hardware.

  • Throughput roughly doubled, reaching 1,413 tokens per second against 708 for the original at 128 concurrent requests.
  • Time to first token was cut roughly in half, from 2.11 seconds to 1.11 seconds.
  • A real-time factor of 712 against 356, which means one server transcribes about 17,000 hours of audio a day.
  • Accuracy held: word error rate of 2.75% against 2.14% for the original on LibriSpeech in English and Spanish, both comfortably below 3%.
  • In a like-for-like batch, 40 customer service calls were transcribed in 15.60 seconds against 28.91 seconds for the original, roughly twice as fast.

The pattern holds across the whole test rather than appearing only at the peak: at every level of load the optimized model keeps its advantage, and that advantage compounds as volume grows.

Why it matters

Contact centres, banks, insurers, healthcare providers and public administrations generate millions of hours of recorded audio every year, and most of it is never transcribed because the economics do not hold. An enterprise contact center handling 20 million calls a year produces around 2 million hours of audio, and at standard cloud transcription rates that single workload can exceed one million dollars a year, running on expensive GPUs.

Running the same workload on a standard server changes both sides of that equation. It costs a fraction of what teams pay today, and because the model is small enough and fast enough to run on infrastructure the organisation already operates, the recordings never have to leave the building. Data sovereignty stops being a trade-off against capability.

That is what unlocks the use cases the current price tag rules out: transcribing 100% of recorded calls for quality and compliance instead of sampling a handful, building searchable archives of every meeting and every support conversation, captioning entire media catalogues, and generating the volume of clean transcript that downstream AI systems depend on. When transcription stops being expensive, it stops being rationed.

This is the first result of an ongoing collaboration with HPE and Intel, and it points to a broader shift: efficient, specialised AI models that let organisations run advanced workloads securely, on their own terms, on the hardware they already have.

Want to know more?

Reach out to us at business@multiversecomputing.com