🎙️
AI Tools/Whisper Audio Transcription

Whisper Audio Transcription

Transcribe audio and video files to text locally. Export SRT subtitles or plain text. 100% browser-based AI — no upload.

AI · 100% LOCALWhisper AI localSRT + TXT export2 model sizesNo upload
Saved Transcripts· Auto-save on
▾
The CalculatorPage 1
Whisper Tiny (Fast)
~144MB · Faster, lower accuracy — good for quick drafts
📥 Downloads on first use
💡 Model will auto-download when you click "Transcribe Audio" (~144MB, cached for offline use)
Private audio transcription — your recordings never leave your device
Upload interviews, podcasts, meeting recordings, or voice notes. Whisper AI transcribes them locally in your browser. Export as SRT subtitles, plain text, or CSV. No cloud upload, no API key, no subscription.
🎙️
Drop audio/video here or click to upload
MP3, WAV, M4A, FLAC, OGG, WebM · Max 50MB · 100% local
Data Source & Legal Disclaimer
Effective: 2026
Sources: Xenova/whisper-tiny.en (Transformers.js ONNX) · Xenova/whisper-base.en (Transformers.js ONNX) · GTCRN Speech Enhancement (WASM SIMD) · ONNX Runtime Web (WASM)

Audio transcription uses OpenAI Whisper models (tiny ~144MB or small ~278MB, fp32) from Xenova, running locally via Transformers.js / ONNX Runtime Web in your browser. Multi-CDN fallback ensures reliable downloads worldwide. Optional GTCRN speech enhancement (WASM SIMD) pre-processes audio to reduce background noise. All models download on first use and are cached afterward. All processing is 100% local — your audio never leaves your device.

See all data sources & update policy →
How it worksPage 2

How local audio transcription works — illustrated

This tool runs OpenAI's Whisper model in your browser to turn speech into text. If you enable noise reduction, the audio is first passed through the GTCRN speech-enhancement network (WASM SIMD) to suppress background noise. The cleaned audio is then decoded into mel spectrogram frames and processed by Whisper in ~30-second chunks with a 5-second stride, so each chunk overlaps the last and no words are dropped at boundaries. Whisper's encoder turns each chunk into a representation of the spoken audio, and its decoder emits the text tokens plus start and end timestamps. The output chunks are normalized into readable sentences and can be exported as plain text, CSV, or SRT subtitles. Your recording, the denoiser, and the model all stay on your device.

From a recording to timestamped subtitles
Local Whisper transcriptionAudio fileoptional GTCRN denoisespeech waveformWhisper model30s chunks · 5s strideEncoderDecodermel spectrogram framestext + timestampsoverlap keeps words safeTranscript + SRTsegments with timestamps00:00:03 → 00:00:08Welcome to the interview…00:00:08 → 00:00:14So let's start with how…export TXT · SRT · CSV

Audio is optionally denoised with GTCRN, split into overlapping 30-second windows, transcribed by the local Whisper encoder–decoder, and exported as text or SRT.

Worked example — Priya's podcast interview

Priya records a 12-minute interview with a guest in a noisy café. She uploads the M4A, leaves noise reduction on, picks the Small model for accuracy, and starts transcription.

  1. Denoise:GTCRN processes the clip first, suppressing the café background so the voice is clearer before Whisper sees it.
  2. Chunk:Whisper splits the 12 minutes into 30-second windows with 5-second overlap, so words spoken across a boundary are still transcribed once.
  3. Transcribe:The encoder turns each window into mel spectrogram features and the decoder emits text with timestamps — 18 segments in total.
  4. Export:Priya downloads subtitles.srt and drops it into her podcast editor, using the per-segment timestamps to line up captions automatically.
FAQ & detailsPage 3

Transcribe audio and video files to text locally using Whisper AI. Everything runs locally; nothing is uploaded.

FreeToolHub Whisper Audio Transcription is a free browser-based tool — transcribe audio and video files to text locally using Whisper AI. No signup, no upload; everything runs locally in your browser.

About this tool

What is this tool?

Transcribe audio and video files to text locally using Whisper AI. Export SRT subtitles or plain text. No upload, no API, free, private.

Whisper AI localSRT + TXT export2 model sizesNo upload

What Is the Whisper Audio Transcription Tool?

The Whisper Audio Transcription tool is a free, browser-based AI tool that converts audio and video files to text. It uses OpenAI's Whisper model running locally via Transformers.js with WebGPU acceleration. Upload interviews, podcasts, meeting recordings, or voice notes — the tool transcribes them entirely in your browser. Export as SRT subtitles for video, or plain text for notes. No cloud upload, no API key, no subscription. Your recordings never leave your device.

Who Should Use This Tool?

Journalists transcribing interviews without uploading to cloud services. Podcasters generating show notes and subtitles. Students recording lectures for searchable text. Researchers transcribing focus groups and oral histories. Business professionals who need meeting transcripts but can't use cloud transcription due to confidentiality. Anyone who wants free, private, local audio transcription.

How to Use the Whisper Transcription Tool

Upload an audio or video file (MP3, WAV, M4A, FLAC, OGG, WebM, MP4 — up to 500MB). Choose a model: Tiny (~40MB, faster) or Small (~240MB, more accurate). Click Transcribe Audio. The model downloads on first use and runs locally. The transcript appears with timestamps for each segment. Export as SRT subtitles, plain text, or copy to clipboard. All processing happens in your browser.

Frequently Asked Questions

How long can my audio file be?

The tool supports files up to 500MB. For long recordings (over 30 minutes), transcription may take several minutes depending on your hardware. The tool processes audio in 30-second chunks with 5-second overlap for accuracy.

Which model should I choose?

Start with Tiny (~40MB) for quick drafts and testing. Switch to Small (~240MB) for important audio where accuracy matters. The Small model handles accents, technical jargon, and overlapping speech significantly better than Tiny.

How accurate is Whisper transcription?

On clear audio — podcasts, interviews, voice memos recorded close to the mic — Whisper is highly accurate, including punctuation and capitalization. Accuracy degrades with heavy background noise, overlapping speakers, heavy accents, and specialist jargon. The transcript preserves hesitations and filler words in some tiers; for publication-ready text, a light manual edit pass is still expected.

What model size should I choose?

The tiers trade speed against accuracy: tiny is fastest and fine for clear speech and quick drafts; base and small add robustness against noise and accents; medium handles the hardest audio but takes noticeably longer to download and run. If you transcribe regularly, save your tier preference — the model caches locally after the first download, so subsequent runs start instantly at your chosen size.

Which audio and video formats are supported?

Common audio containers (mp3, wav, m4a, ogg, flac) and video files (mp4, webm, mov) all work — the audio track is extracted and decoded locally before transcription. For long recordings, processing time scales with duration; a one-hour file takes several minutes even on fast tiers. The transcript is produced as timestamped chunks you can review and copy.

Is my audio uploaded anywhere?

No. The Whisper model runs entirely in your browser via WebAssembly/WebGPU — audio is decoded and transcribed on your device, and nothing is sent to a server. This makes the tool safe for confidential meetings, medical or legal recordings, and any audio you would not upload to a cloud service. Your model-tier preference is the only thing stored locally.

Other names for this tool

This tool is also known by these tasks — each link opens the same tool with a focused guide:

Related tools

Joke of the Day
Sep 29

Why did the cookie go to the doctor?

Free core, no account

Keep the Free Edition Free

No signups, no data sold. Every tool is free to use — the free tier allows 5 downloads or saves per day, and the optional Pro plan ($7/mo, $59/yr) adds unlimited downloads, batch processing, white-label exports and an ad-free experience.

☕Support me on Ko-fi— support the free tier

100% of proceeds go towards hosting & building more free tools.