Skip to main content

Audio to text

Ten hours of audio, done in thirty minutes

Whisper takes tens of minutes per audio hour on a CPU and tens of seconds on a GPU. The bottleneck in transcription was never the model — it was whether you had a card.

Transcription per hour, from
$0.193
Audio in about 30 minutes
10h
Languages supported
99+

01 — When it applies

When to run transcription yourself

Commercial transcription APIs bill per audio minute, which is convenient for occasional files. The moment you have hundreds of hours of backlog — podcast archives, meeting recordings, course video, support calls — a per-minute bill turns ugly. And that work is exactly the one-off batch job that suits renting a machine, finishing, and leaving.

The second reason is sensitivity. Meeting recordings, medical consultations, legal calls and support conversations routinely contain material that cannot leave your control. Running transcription yourself means the audio only ever exists on the machine you rented, and destroying the instance erases it.

The cost comparison is stark. Whisper large-v3 on one consumer GPU transcribes ten hours of audio in about half an hour. At $0.193/hr for an entry card that is a few cents. The same volume through a commercial API is usually double-digit dollars.

02 — Workloads

Common transcription jobs

What they share: high volume, one-off, sensitive content.

  • Archive transcription

    Podcasts, courses, interviews, stream recordings — turn hundreds of hours of back catalogue into text as a base for search, summarisation or repurposing.

  • Meeting notes and minutes

    Pair with speaker diarisation to turn multi-party recordings into speaker-labelled text, then feed an LLM to draft minutes. The entire chain runs on one machine.

  • Subtitles and timeline alignment

    WhisperX produces word-level timestamps, so subtitle alignment is far tighter than base Whisper's sentence-level output and exports directly to usable SRT or VTT.

  • Speech dataset cleaning

    Preparing TTS or speech-model training data means transcribing large volumes and filtering by quality. It is bulk repetitive work where GPU batching is the only realistic option.

03 — Stack

The speech toolchain, preinstalled

Three implementations covering different accuracy and speed trade-offs.

  • faster-whisper — speed first

    A CTranslate2 reimplementation that runs several times faster than base Whisper at equal accuracy, with lower VRAM use. The default choice for batch work.

    faster-whisper · CTranslate2 · Low VRAM

  • WhisperX — accuracy and alignment

    Adds forced alignment and speaker diarisation on top of Whisper, producing word-level timestamps. The correct answer for subtitles and meeting minutes.

    WhisperX · Word timestamps · Diarisation

  • Base Whisper and post-processing

    OpenAI's own implementation as an accuracy baseline. Chain an LLM afterwards for error correction, punctuation restoration and summarisation on the same instance.

    Whisper · Punctuation · Summarisation

04 — Choosing a GPU

Sizing a card for transcription

Whisper's VRAM appetite is small, which makes this one of the clearest cases for a cheap card. Do not rent an H100 to transcribe audio.

Job scaleVRAM neededCheapest availableNotes
small / medium models6GBGTX 1660 S$0.083/hrFine for clean everyday recordings and extremely fast. An entry card makes cost essentially negligible.
large-v3, single stream10GBRTX 3060$0.100/hrThe standard highest-accuracy setup, noticeably steadier on noisy audio and mixed languages.
large-v3 with diarisation16GBRTX 4060 Ti$0.126/hrWhisperX's alignment and diarisation models need extra VRAM; budget for this tier on meeting work.
Parallel batch processing24GBTesla V100$0.188/hrFor hundreds of hours of backlog, run multiple processes — more VRAM means more concurrent streams.

"Cheapest available" is derived from live inventory — the lowest per-GPU rate among models that clear the VRAM bar — and moves with the market. Multi-GPU nodes rent whole.

05 — Getting started

Transcribing a batch

  • 01

    Stage the audio somewhere pullable

    Object storage or any publicly reachable location. Pulling from inside the instance uses the host's bandwidth, far faster than uploading — which matters a lot across hundreds of files.

  • 02

    Pick a cheap card

    Transcription is not VRAM-hungry, so an entry card saturates fine. Choose a tier from the table and size the disk for the audio plus output text.

  • 03

    Run the batch, then destroy

    Loop over the files, download the transcripts, destroy the instance. The whole job usually finishes within an hour or two.

06 — FAQ

Speech to text

What does transcribing ten hours of audio cost?

faster-whisper running large-v3 on a consumer card needs about thirty minutes. At $0.193/hr for an entry card that is under a dime; even on an RTX 4090 it is around half of $0.540. At this scale the real cost is the time you spend organising the audio.

How accurate is it on non-English audio?

Whisper large-v3 handles major languages well, with high accuracy on clear speech. Dialects, specialised terminology and noisy environments degrade noticeably — for those, chain an LLM for post-transcription correction, which helps substantially and runs on the same instance.

Can it separate speakers?

Yes, through WhisperX's diarisation, which labels each segment with a speaker index — essential for meeting minutes and interview cleanup. Note that the diarisation model consumes extra VRAM, so choose 16GB or more per the table above.

Which audio formats are supported?

The image includes ffmpeg, so mp3, wav, m4a and flac all work, as does extracting the audio track straight from an mp4. No pre-conversion needed.

Can it produce subtitles directly?

Yes. WhisperX emits word-level timestamps and exports SRT or VTT with alignment tight enough to use as-is. Base Whisper only gives sentence-level timestamps, which produces perceptible drift in subtitles.

Start building on NexGPU

Enterprise R&D team or solo developer — either way, your first job can be running in minutes.

Sign up to browse live network pricing. No payment method required.