Transcribe audio and video files to text locally. Export SRT subtitles or plain text. 100% browser-based AI — no upload.
Audio transcription uses OpenAI Whisper models (tiny ~144MB or small ~278MB, fp32) from Xenova, running locally via Transformers.js / ONNX Runtime Web in your browser. Multi-CDN fallback ensures reliable downloads worldwide. Optional GTCRN speech enhancement (WASM SIMD) pre-processes audio to reduce background noise. All models download on first use and are cached afterward. All processing is 100% local — your audio never leaves your device.
This tool runs OpenAI's Whisper model in your browser to turn speech into text. If you enable noise reduction, the audio is first passed through the GTCRN speech-enhancement network (WASM SIMD) to suppress background noise. The cleaned audio is then decoded into mel spectrogram frames and processed by Whisper in ~30-second chunks with a 5-second stride, so each chunk overlaps the last and no words are dropped at boundaries. Whisper's encoder turns each chunk into a representation of the spoken audio, and its decoder emits the text tokens plus start and end timestamps. The output chunks are normalized into readable sentences and can be exported as plain text, CSV, or SRT subtitles. Your recording, the denoiser, and the model all stay on your device.
Audio is optionally denoised with GTCRN, split into overlapping 30-second windows, transcribed by the local Whisper encoder–decoder, and exported as text or SRT.
Priya records a 12-minute interview with a guest in a noisy café. She uploads the M4A, leaves noise reduction on, picks the Small model for accuracy, and starts transcription.
Transcribe audio and video files to text locally using Whisper AI. Export SRT subtitles or plain text. No upload, no API, free, private.
The Whisper Audio Transcription tool is a free, browser-based AI tool that converts audio and video files to text. It uses OpenAI's Whisper model running locally via Transformers.js with WebGPU acceleration. Upload interviews, podcasts, meeting recordings, or voice notes — the tool transcribes them entirely in your browser. Export as SRT subtitles for video, or plain text for notes. No cloud upload, no API key, no subscription. Your recordings never leave your device.
Journalists transcribing interviews without uploading to cloud services. Podcasters generating show notes and subtitles. Students recording lectures for searchable text. Researchers transcribing focus groups and oral histories. Business professionals who need meeting transcripts but can't use cloud transcription due to confidentiality. Anyone who wants free, private, local audio transcription.
Upload an audio or video file (MP3, WAV, M4A, FLAC, OGG, WebM, MP4 — up to 500MB). Choose a model: Tiny (~40MB, faster) or Small (~240MB, more accurate). Click Transcribe Audio. The model downloads on first use and runs locally. The transcript appears with timestamps for each segment. Export as SRT subtitles, plain text, or copy to clipboard. All processing happens in your browser.
The tool supports files up to 500MB. For long recordings (over 30 minutes), transcription may take several minutes depending on your hardware. The tool processes audio in 30-second chunks with 5-second overlap for accuracy.
Start with Tiny (~40MB) for quick drafts and testing. Switch to Small (~240MB) for important audio where accuracy matters. The Small model handles accents, technical jargon, and overlapping speech significantly better than Tiny.
This tool is also known by these tasks — each link opens the same tool with a focused guide:
What do you call a crab that plays baseball?
No paywalls, no signups, no data sold. Built by a solo developer who believes useful tools should be accessible to everyone.
☕Support me on Ko-fi— keep tools free100% of proceeds go towards hosting & building more free tools.