Audio to text
Ten hours of audio, done in thirty minutes
Whisper takes tens of minutes per audio hour on a CPU and tens of seconds on a GPU. The bottleneck in transcription was never the model — it was whether you had a card.
- Transcription per hour, from
- $0.193
- Audio in about 30 minutes
- 10h
- Languages supported
- 99+
01 — When it applies
When to run transcription yourself
Commercial transcription APIs bill per audio minute, which is convenient for occasional files. The moment you have hundreds of hours of backlog — podcast archives, meeting recordings, course video, support calls — a per-minute bill turns ugly. And that work is exactly the one-off batch job that suits renting a machine, finishing, and leaving.
The second reason is sensitivity. Meeting recordings, medical consultations, legal calls and support conversations routinely contain material that cannot leave your control. Running transcription yourself means the audio only ever exists on the machine you rented, and destroying the instance erases it.
The cost comparison is stark. Whisper large-v3 on one consumer GPU transcribes ten hours of audio in about half an hour. At $0.193/hr for an entry card that is a few cents. The same volume through a commercial API is usually double-digit dollars.
02 — Workloads
Common transcription jobs
What they share: high volume, one-off, sensitive content.
Archive transcription
Podcasts, courses, interviews, stream recordings — turn hundreds of hours of back catalogue into text as a base for search, summarisation or repurposing.
Meeting notes and minutes
Pair with speaker diarisation to turn multi-party recordings into speaker-labelled text, then feed an LLM to draft minutes. The entire chain runs on one machine.
Subtitles and timeline alignment
WhisperX produces word-level timestamps, so subtitle alignment is far tighter than base Whisper's sentence-level output and exports directly to usable SRT or VTT.
Speech dataset cleaning
Preparing TTS or speech-model training data means transcribing large volumes and filtering by quality. It is bulk repetitive work where GPU batching is the only realistic option.
03 — Stack
The speech toolchain, preinstalled
Three implementations covering different accuracy and speed trade-offs.
faster-whisper — speed first
A CTranslate2 reimplementation that runs several times faster than base Whisper at equal accuracy, with lower VRAM use. The default choice for batch work.
faster-whisper · CTranslate2 · Low VRAM
WhisperX — accuracy and alignment
Adds forced alignment and speaker diarisation on top of Whisper, producing word-level timestamps. The correct answer for subtitles and meeting minutes.
WhisperX · Word timestamps · Diarisation
Base Whisper and post-processing
OpenAI's own implementation as an accuracy baseline. Chain an LLM afterwards for error correction, punctuation restoration and summarisation on the same instance.
Whisper · Punctuation · Summarisation
04 — Choosing a GPU
Sizing a card for transcription
Whisper's VRAM appetite is small, which makes this one of the clearest cases for a cheap card. Do not rent an H100 to transcribe audio.
| Job scale | VRAM needed | Cheapest available | Notes |
|---|---|---|---|
| small / medium models | 6GB | GTX 1660 S$0.083/hr | Fine for clean everyday recordings and extremely fast. An entry card makes cost essentially negligible. |
| large-v3, single stream | 10GB | RTX 3060$0.100/hr | The standard highest-accuracy setup, noticeably steadier on noisy audio and mixed languages. |
| large-v3 with diarisation | 16GB | RTX 4060 Ti$0.126/hr | WhisperX's alignment and diarisation models need extra VRAM; budget for this tier on meeting work. |
| Parallel batch processing | 24GB | Tesla V100$0.188/hr | For hundreds of hours of backlog, run multiple processes — more VRAM means more concurrent streams. |
"Cheapest available" is derived from live inventory — the lowest per-GPU rate among models that clear the VRAM bar — and moves with the market. Multi-GPU nodes rent whole.
05 — Getting started
Transcribing a batch
01
Stage the audio somewhere pullable
Object storage or any publicly reachable location. Pulling from inside the instance uses the host's bandwidth, far faster than uploading — which matters a lot across hundreds of files.
02
Pick a cheap card
Transcription is not VRAM-hungry, so an entry card saturates fine. Choose a tier from the table and size the disk for the audio plus output text.
03
Run the batch, then destroy
Loop over the files, download the transcripts, destroy the instance. The whole job usually finishes within an hour or two.
06 — FAQ
Speech to text
What does transcribing ten hours of audio cost?
How accurate is it on non-English audio?
Can it separate speakers?
Which audio formats are supported?
Can it produce subtitles directly?
Related solutions
Other ways to use it
Same compute network — swap the image and it becomes a different production line.
AI Agents
Deploy and scale agents on LangChain, CrewAI
Private LLM Deployment
Open weights, running on your own machine
AI Fine-tuning
LoRA, QLoRA and full fine-tuning, on demand
AI Image & Video
Stable Diffusion, FLUX and ComfyUI, ready to run
AI Text Generation
vLLM, TGI and Ollama, live in minutes
AI/ML Frameworks
Native PyTorch, TensorFlow and JAX
Start building on NexGPU
Enterprise R&D team or solo developer — either way, your first job can be running in minutes.
Sign up to browse live network pricing. No payment method required.
