What you get
A Python script that transcribes a local audio file to text with Whisper running on your own machine, with no API key and no upload of your audio.
Prerequisites
- Python 3.8 or later. The README says the codebase is expected to be compatible with Python 3.8-3.11 and recent PyTorch versions.
pip.
- A package manager that can install ffmpeg: apt, pacman, Homebrew, Chocolatey or Scoop.
- An audio file, for example
audio.mp3. The README examples use .mp3, .flac and .wav.
- For the
turbo model, about 6 GB of VRAM (README figure).
Steps
- Install ffmpeg. The README lists it as a required command-line tool. Run the one line for your system:
# Ubuntu or Debian
sudo apt update && sudo apt install ffmpeg
# Arch Linux
sudo pacman -S ffmpeg
# macOS (Homebrew)
brew install ffmpeg
# Windows (Chocolatey)
choco install ffmpeg
# Windows (Scoop)
scoop install ffmpeg
- Install the official Python package, pinned to the current release on PyPI:
pip install openai-whisper==20250625
If the install fails with a setuptools_rust error, the README says to run pip install setuptools-rust and try again.
- Create
transcribe.py:
import sys
import whisper
model = whisper.load_model("turbo")
result = model.transcribe(sys.argv[1])
print(result["text"])
- Pick a model size if
turbo does not fit your machine. The README lists tiny, base, small, medium, large and turbo: smaller models are faster and need less memory (tiny and base about 1 GB VRAM, large about 10 GB at 1x speed), and turbo is an optimized large-v3 at about 6 GB and about 8x speed with a small loss in accuracy. To switch, change the name in load_model:
model = whisper.load_model("base")
For English-only audio the README also lists tiny.en, base.en, small.en and medium.en. The turbo model is not trained for translation.
Test it
Put an audio file next to the script and run:
python transcribe.py audio.mp3
You should see the transcript of the file printed as plain text in the terminal. The first run takes longer because the model weights have to be fetched before transcription starts; your audio file is read locally and is not uploaded. Long files work too: transcribe() reads the whole file and processes it in 30-second windows.