How to Transcribe an m4a File | No Conversion, No Upload

Created on:September 20, 2026 at 7:17 AMReading time: 8 min
thumbnail Image

Recordings made with the Voice Memos app on an iPhone, or with a handheld IC recorder, are saved with the ".m4a" extension. Some transcription services cannot read that format and tell you to convert the file to MP3 first, but whether a conversion is needed depends on the formats the tool supports, not on m4a being an awkward format.

This article walks through transcribing an m4a file without converting it, starting from moving the recording onto a PC and going all the way to separating speakers and exporting a subtitle file.

What You'll Learn from This Article 

  • Why an m4a file can be transcribed as it is
  • How to move the recording to a PC
  • How to transcribe the m4a file in the browser
  • How to separate speakers, and how to export a subtitle file
  • What to do when it does not work

All you need is a PC, Chrome or Edge, and a ".m4a" recording. There is no account, no app install, no file upload and no format conversion.

Why m4a can be transcribed as it is 

"m4a" is the container used when audio is compressed with AAC, and it is the default save format for the iPhone Voice Memos app, for several Android recording apps, and for many IC recorders. As audio data it behaves the same way MP3 does, so as long as the tool supports it, no conversion is necessary.

The request to convert most likely means that the library the service uses internally cannot read m4a. The tool I built, "Transcriber", runs "ffmpeg", a program for converting audio and video, inside the browser, so it pulls the audio out of an m4a or an MP4 the moment the file is loaded.

There are several ways to transcribe a recording, and the choice changes both how large a file you can handle and whether your audio is sent to a server, so it is worth comparing them first.

Method

Conversion

Length limit

Speaker separation

Sent to a server

Built into the iPhone

Not needed

None

Not available

No

Cloud services

Depends on the service

Capped on free plans

Some offer it on free plans

Yes (the provider's servers)

Processed in the browser

Not needed

None

Available

No

Desktop software

Not needed

None

Configurable

No

If you are on an iPhone and reading the text on that same device is enough, the built-in feature is the quickest route. On an iPhone 12 or later running iOS 18 or later, Voice Memos can transcribe recordings in ten languages, including English and Japanese, both while recording and afterwards (not available in some countries or regions).

The built-in feature cannot split the text by who was speaking, however, and it has no way to export a subtitle file, so for a meeting or an interview with several participants the later work goes more smoothly if you process the recording on a PC.

What to prepare 

  • A PC. Windows, macOS and Linux all work
  • The latest Chrome or Edge. Safari and Firefox work too, but depending on the browser and device they can be slower for the reason described below
  • A ".m4a" recording. Files up to 2 GB can be loaded

Step 1: Move the recording to a PC 

First copy the file from the recording device onto the PC. The steps differ by device, so follow the one that matches your setup.

  1. On an iPhone, tap the recording in Voice Memos, tap the More button (…), then tap Share and choose AirDrop or "Save to Files". To get it onto a Mac, use AirDrop or iCloud sync; on Windows, save it to iCloud Drive and open it with iCloud for Windows or iCloud.com.
  2. On Android, share the file to Google Drive from the recording app, or connect over USB and open the recordings folder directly.
  3. With an IC recorder, connect it over USB and it appears as external storage, so you can copy the file straight across.

Step 2: Transcribe it in the browser 

Once the file is on the PC, open a transcription tool in the browser. I am using "Transcriber", which I built myself, but the shape of the procedure is the same for any tool that reads m4a directly.

Free Unlimited AI Transcription — No Sign-Up, No Upload - Transcriber

Transcribe audio and video files entirely in your browser. No account and no upload — your audio never leaves your device. Free with no time or usage limits. Speaker diarization and SRT subtitle export powered by Whisper AI, supporting MP3, WAV, M4A, MP4 and more.

favicontranscriber.tools.ryusei.io
  1. Open "Transcriber" in the browser.
  2. Set "Model", "Language" and "Speaker diarization" before doing anything else.
  3. Drop the m4a file onto the area marked "Drop an audio or video file, or click to choose".
  4. Wait for processing to finish.

Choosing a model 

Six models are available. Larger ones transcribe more accurately, but they also use more memory.

Below are the numbers I measured earlier on a 4 minute 08 second video (51.9 MB). The PC had an RTX 2060 SUPER (8 GB), the browser was Chrome, and WebGPU, which lets the browser use the GPU, was enabled.

Model

Processing time

Speed relative to real time

Small

0 min 35 s

about 7.1x

Medium

1 min 22 s

about 3.0x

Large v3 Turbo

0 min 37 s

about 6.7x

Large v3

0 min 52 s

about 4.8x

In these results the processing time did not follow model size, and the slowest of the four was the mid-sized Medium. Large v3 Turbo finishes almost as fast as Small because the part that analyses the audio is the same size as in Large v3, and only the part that assembles the text was made smaller.

I use "Large v3 Turbo" myself. It picks up Japanese proper nouns and technical terms clearly more reliably than Small, while taking roughly the same amount of time.

Step 3: Separate the speakers and export 

When the recording has more than one person in it, turn "Speaker diarization" on before starting. The exported text is then split line by line according to who said what.

You can normally leave "Speakers (0 = auto)" on Auto. Set a number between two and six only if the result clearly has the wrong number of speakers.

When a Whisper or Qwen3-ASR model runs on a PC where WebGPU is available (except in Firefox), with "Speakers (0 = auto)" left on Auto, the audio is processed by NVIDIA's speaker diarization model "Nemotron 3 Diarization". With SenseVoice Small, the default for Japanese, Chinese and Korean (it runs on the CPU), on a PC without WebGPU, on a smartphone, in Firefox, or when you set the number of speakers, the earlier method is used: a browser port of the "pyannote.audio" speaker diarization pipeline, which slices the audio into ten-second chunks, extracts the features of each speaker and groups the segments that belong to the same person. Very short interjections of about a second are sometimes missed, so those are worth fixing on screen after the export.

Choosing between text and subtitles 

Two export formats are available, "TXT" and "SRT (subtitles)", and which one to pick depends on what happens next.

  • Use TXT for meeting notes or a draft of an article
  • Use SRT for adding subtitles to a video

When it does not work 

When processing stops partway through, the failure reports that reach me point almost entirely at one cause, which is running out of memory.

  • Smartphones have less memory available, so long recordings stop partway through more often
  • Where WebGPU is unavailable, the work falls back to the CPU and becomes far slower
  • For recordings over an hour, start with a small model and move up only if the accuracy is not good enough

If the warning "WebGPU is not available in this environment, so this model will be very slow" appears, reopen the page in Chrome or Edge, or pick a smaller model if that is not an option.

Frequently asked questions 

Is it really free? Are there limits on length or on how many times I can use it?

It is free, and there is no limit on length or on the number of runs. The transcription runs on the visitor's own PC, which means no server costs land on my side.

Does accuracy improve if I convert the m4a to MP3 first?

It does not. Converting decodes the audio and then compresses it again, so the quality drops rather than improves. Load the m4a as it is.

Is there any chance the recording is sent to a server?

There is not. The transcription runs entirely inside the browser, so the audio file never leaves the device. Apart from downloading the models, Transcriber sends statistics used to fix problems (such as device specs and how long processing took) to the site's server, but these do not include the audio, the transcript, file names or IP addresses. You can turn this off with the switch in the "Help us improve the service" section near the bottom of the page.

Do I need an account or a login?

You do not. No email address is asked for either.

Can I do all of this on an iPhone or an Android phone?

You can, but I would not recommend it. Long recordings often stop partway through on a phone, so recording on the phone and transcribing on a PC is the dependable combination.

How long a recording can it handle?

There is no limit on length as long as the file stays under 2 GB. Longer recordings take longer to process, so the tab has to stay open until it finishes.

Summary 

An m4a file can be transcribed as it is. Even if a service tells you to convert it to MP3 first, there is nothing wrong with the m4a format itself.

  • Move the recording to a PC
  • Set the model, the language and speaker diarization first
  • Load the file (processing starts the moment it is loaded)
  • Export as TXT or SRT

Related articles and tools 

How to Transcribe X Space Recordings | Free, No Sign-Up, No Upload | Ryusei.IO

A step-by-step guide to turning a recorded X (formerly Twitter) Space into text for free. Covers downloading the recording, separating speakers and exporting subtitle files, all in the browser without an account or an app install.

faviconryusei.io
How to Add Subtitles to a Video in Your Browser | Free, No Install, No Upload | Ryusei.IO

No install, no upload. Transcribe a video, fix the SRT, and burn the subtitles into the picture — all inside your browser, free and without an account.

faviconryusei.io
Audio Converter — MP3, WAV, M4A, AAC, OGG, OPUS, FLAC, AIFF | Media Tools

Convert audio to MP3, WAV, M4A, AAC, OGG, OPUS, FLAC or AIFF, or extract audio from MP4 and MOV videos. Free, in your browser, nothing uploaded, no length limit.

faviconmedia.tools.ryusei.io

Latest Tips