Speech to Text

Transcribe Japanese or English audio files on your device without uploading recordings.

Your entries, name, address, photos and files stay in this page’s memory. They are not uploaded and are cleared when you reload.

Upload an audio file (up to 3 minutes / 50 MB). Whisper runs locally in Japanese or English. No microphone access or recording is used. The first run needs a connection to download the model.

No file selected

How to use this tool

How to use

Choose an audio file up to three minutes and 50 MB, select the spoken language, and start. Review the editable transcript, then copy it or save a text file. Split longer recordings before uploading them to this local tool.

When it helps

Useful for short interviews, study notes and excerpts from meetings. Noise, overlapping speakers and proper names can cause mistakes. This lightweight model does not provide speaker identification or guaranteed accuracy.

How it works

A multilingual Whisper model runs inside a Web Worker. This file-based tool does not use remote Web Speech services or request microphone access. Initial model downloads need a connection; inference stays on your device.

Frequently asked questions

Q. Can I use the microphone?
A. This tool accepts audio files only and never requests microphone permission. Unsupported browsers show a helpful message.