Recordings made with the Voice Memos app on an iPhone, or with a handheld IC recorder, are saved with the ".m4a" extension. Some transcription services cannot read that format and tell you to convert the file to MP3 first, but whether a conversion is needed depends on the formats the tool supports, not on m4a being an awkward format.
This article walks through transcribing an m4a file without converting it, starting from moving the recording onto a PC and going all the way to separating speakers and exporting a subtitle file.
What You'll Learn from This Article
- Why an m4a file can be transcribed as it is
- How to move the recording to a PC
- How to transcribe the m4a file in the browser
- How to separate speakers, and how to export a subtitle file
- What to do when it does not work
All you need is a PC, Chrome or Edge, and a ".m4a" recording. There is no account, no app install, no file upload and no format conversion.
Why m4a can be transcribed as it is
"m4a" is the container used when audio is compressed with AAC, and it is the default save format for the iPhone Voice Memos app, for several Android recording apps, and for many IC recorders. As audio data it behaves the same way MP3 does, so as long as the tool supports it, no conversion is necessary.
The request to convert most likely means that the library the service uses internally cannot read m4a. The tool I built, "Transcriber", runs "ffmpeg", a program for converting audio and video, inside the browser, so it pulls the audio out of an m4a or an MP4 the moment the file is loaded.
There are several ways to transcribe a recording, and the choice changes both how large a file you can handle and whether your audio is sent to a server, so it is worth comparing them first.
Method | Conversion | Length limit | Speaker separation | Sent to a server |
|---|---|---|---|---|
Built into the iPhone | Not needed | None | Not available | No |
Cloud services | Depends on the service | Capped on free plans | Some offer it on free plans | Yes (the provider's servers) |
Processed in the browser | Not needed | None | Available | No |
Desktop software | Not needed | None | Configurable | No |
If you are on an iPhone and reading the text on that same device is enough, the built-in feature is the quickest route. On an iPhone 12 or later running iOS 18 or later, Voice Memos can transcribe recordings in ten languages, including English and Japanese, both while recording and afterwards (not available in some countries or regions).
The built-in feature cannot split the text by who was speaking, however, and it has no way to export a subtitle file, so for a meeting or an interview with several participants the later work goes more smoothly if you process the recording on a PC.
What to prepare
- A PC. Windows, macOS and Linux all work
- The latest Chrome or Edge. Safari and Firefox work too, but depending on the browser and device they can be slower for the reason described below
- A ".m4a" recording. Files up to 2 GB can be loaded
Step 1: Move the recording to a PC
First copy the file from the recording device onto the PC. The steps differ by device, so follow the one that matches your setup.
- On an iPhone, tap the recording in Voice Memos, tap the More button (…), then tap Share and choose AirDrop or "Save to Files". To get it onto a Mac, use AirDrop or iCloud sync; on Windows, save it to iCloud Drive and open it with iCloud for Windows or iCloud.com.
- On Android, share the file to Google Drive from the recording app, or connect over USB and open the recordings folder directly.
- With an IC recorder, connect it over USB and it appears as external storage, so you can copy the file straight across.
Step 2: Transcribe it in the browser
Once the file is on the PC, open a transcription tool in the browser. I am using "Transcriber", which I built myself, but the shape of the procedure is the same for any tool that reads m4a directly.

Transcribe audio and video files entirely in your browser. No account and no upload — your audio never leaves your device. Free with no time or usage limits. Speaker diarization and SRT subtitle export powered by Whisper AI, supporting MP3, WAV, M4A, MP4 and more.
- Open "Transcriber" in the browser.
- Set "Model", "Language" and "Speaker diarization" before doing anything else.
- Drop the m4a file onto the area marked "Drop an audio or video file, or click to choose".
- Wait for processing to finish.
Choosing a model
Six models are available. Larger ones transcribe more accurately, but they also use more memory.
Below are the numbers I measured earlier on a 4 minute 08 second video (51.9 MB). The PC had an RTX 2060 SUPER (8 GB), the browser was Chrome, and WebGPU, which lets the browser use the GPU, was enabled.
Model | Processing time | Speed relative to real time |
|---|---|---|
Small | 0 min 35 s | about 7.1x |
Medium | 1 min 22 s | about 3.0x |
Large v3 Turbo | 0 min 37 s | about 6.7x |
Large v3 | 0 min 52 s | about 4.8x |
In these results the processing time did not follow model size, and the slowest of the four was the mid-sized Medium. Large v3 Turbo finishes almost as fast as Small because the part that analyses the audio is the same size as in Large v3, and only the part that assembles the text was made smaller.
I use "Large v3 Turbo" myself. It picks up Japanese proper nouns and technical terms clearly more reliably than Small, while taking roughly the same amount of time.
Step 3: Separate the speakers and export
When the recording has more than one person in it, turn "Speaker diarization" on before starting. The exported text is then split line by line according to who said what.
You can normally leave "Speakers (0 = auto)" on Auto. Set a number between two and six only if the result clearly has the wrong number of speakers.
When a Whisper or Qwen3-ASR model runs on a PC where WebGPU is available (except in Firefox), with "Speakers (0 = auto)" left on Auto, the audio is processed by NVIDIA's speaker diarization model "Nemotron 3 Diarization". With SenseVoice Small, the default for Japanese, Chinese and Korean (it runs on the CPU), on a PC without WebGPU, on a smartphone, in Firefox, or when you set the number of speakers, the earlier method is used: a browser port of the "pyannote.audio" speaker diarization pipeline, which slices the audio into ten-second chunks, extracts the features of each speaker and groups the segments that belong to the same person. Very short interjections of about a second are sometimes missed, so those are worth fixing on screen after the export.
Choosing between text and subtitles
Two export formats are available, "TXT" and "SRT (subtitles)", and which one to pick depends on what happens next.
- Use TXT for meeting notes or a draft of an article
- Use SRT for adding subtitles to a video
When it does not work
When processing stops partway through, the failure reports that reach me point almost entirely at one cause, which is running out of memory.
- Smartphones have less memory available, so long recordings stop partway through more often
- Where WebGPU is unavailable, the work falls back to the CPU and becomes far slower
- For recordings over an hour, start with a small model and move up only if the accuracy is not good enough
If the warning "WebGPU is not available in this environment, so this model will be very slow" appears, reopen the page in Chrome or Edge, or pick a smaller model if that is not an option.
Frequently asked questions
Is it really free? Are there limits on length or on how many times I can use it?
It is free, and there is no limit on length or on the number of runs. The transcription runs on the visitor's own PC, which means no server costs land on my side.
Does accuracy improve if I convert the m4a to MP3 first?
It does not. Converting decodes the audio and then compresses it again, so the quality drops rather than improves. Load the m4a as it is.
Is there any chance the recording is sent to a server?
There is not. The transcription runs entirely inside the browser, so the audio file never leaves the device. Apart from downloading the models, Transcriber sends statistics used to fix problems (such as device specs and how long processing took) to the site's server, but these do not include the audio, the transcript, file names or IP addresses. You can turn this off with the switch in the "Help us improve the service" section near the bottom of the page.
Do I need an account or a login?
You do not. No email address is asked for either.
Can I do all of this on an iPhone or an Android phone?
You can, but I would not recommend it. Long recordings often stop partway through on a phone, so recording on the phone and transcribing on a PC is the dependable combination.
How long a recording can it handle?
There is no limit on length as long as the file stays under 2 GB. Longer recordings take longer to process, so the tab has to stay open until it finishes.
Summary
An m4a file can be transcribed as it is. Even if a service tells you to convert it to MP3 first, there is nothing wrong with the m4a format itself.
- Move the recording to a PC
- Set the model, the language and speaker diarization first
- Load the file (processing starts the moment it is loaded)
- Export as TXT or SRT
Related articles and tools
A step-by-step guide to turning a recorded X (formerly Twitter) Space into text for free. Covers downloading the recording, separating speakers and exporting subtitle files, all in the browser without an account or an app install.
No install, no upload. Transcribe a video, fix the SRT, and burn the subtitles into the picture — all inside your browser, free and without an account.

Convert audio to MP3, WAV, M4A, AAC, OGG, OPUS, FLAC or AIFF, or extract audio from MP4 and MOV videos. Free, in your browser, nothing uploaded, no length limit.