Whisper

Convert speech recordings to text with an open model.

Whisper: model memory requirements

tiny1 GB~1 GB
base1 GB~1 GB
small2 GB~2 GB
medium5 GB~5 GB
large10 GB~10 GB
turbo6 GB~6 GB
Relative speed
ModelRelative speed
tiny~10×
base~7×
small~4×
medium~2×
large~1×
turbo~8×
Approximate VRAM requirements from official Whisper documentation, checked 3 October 2026. Not measured by Findbest. Speed: published English transcription comparison on NVIDIA A100; large = 1×. Your hardware may perform differently. Data and test conditions

What is free?

Code and model weights use the MIT licence. Local transcription needs compute; hosted APIs may charge fees.

Check provider pricing

What can you use it for?

For technical users with their own recordings.

Getting started

Follow the official setup including audio dependencies. Start with a short recording and compare transcript and audio.

What the tool does

Convert speech recordings to text with an open model.

Who it is for

For technical users with their own recordings.

Things to consider

Recognition errors and invented passages are possible.

This is not a hands-on benchmark. Our analysis uses the linked provider information and the stated use case.

Related alternatives

DeepL

Languages & translation

Translate text between languages, then review the result.

Not yet verifiedView tool