You have an hour of interview audio and something important to make from it — a dissertation chapter, an article, a research report. Search for how to transcribe an interview and the advice splits into two camps: purists who tell you to type every word yourself, and a wall of AI tools that want the recording uploaded to their servers. Both are half right. Manual transcription keeps you close to the data but costs four to six hours per hour of audio. Cloud AI is fast but sends your participant’s voice to a vendor — a real problem for confidential interviews. And there is a third option most guides skip: AI transcription that runs locally on your own computer, which is fast and private.
This guide compares all three methods, shows the transcription styles with samples, walks through a manual workflow step by step, and covers the formatting, consent, and IRB questions that often decide which method you are allowed to use in the first place.
The three ways to transcribe an interview
| Method | Speed | Typical cost | Where the audio goes |
|---|---|---|---|
| Manual (yourself) | 4–6 hrs per audio hour | Your time | Stays with you |
| Human service | Typically a day or more | ~$1–2 per audio minute | Vendor + human transcribers |
| Cloud AI service | Minutes | ~$8–17/month | Vendor servers |
| Local AI | Minutes | One-time (€49–59 typical) | Never leaves your machine |
Prices checked July 2026.
Manual transcription is slow on purpose. Many qualitative researchers treat typing the transcript as the first pass of analysis — you hear every hesitation, every change of direction. Expect four to six hours per hour of audio if you type well, more with difficult recordings. It is the right call for small studies, conversation analysis, and anywhere the exact texture of speech matters.
Human transcription services charge per audio minute — roughly $1–2 as of mid-2026 — and return a formatted transcript, usually within a day or a few. Good for difficult audio and certified transcripts, but you are still handing the recording to a third party, and the per-interview cost adds up fast across a twenty-participant study.
Cloud AI services such as Otter and Fireflies transcribe an hour of audio in minutes for a monthly subscription. The catch is in the name: your audio is uploaded, processed, and often retained on the vendor’s servers. Otter is facing ongoing privacy litigation and Fireflies is facing biometric-privacy litigation, which is exactly the kind of thing an ethics board notices. For non-sensitive interviews at volume they work well; we compared the field in 8 Otter.ai alternatives.
Local AI means the same class of speech models, running on your own laptop instead of a server. Modern consumer machines handle it comfortably, so you get cloud-like speed with manual-level privacy, and the tools tend to be one-time purchases rather than subscriptions. Full disclosure: we build Eidetic, one of the local tools covered below — the comparison stays honest anyway.
Pick a transcription style first (with samples)
Before you transcribe anything, decide how much of the mess of real speech to keep. There are three common styles, and the choice changes both the effort and what the transcript is good for. Here is the same example of a transcribed interview exchange in each style.
Verbatim keeps everything: fillers, false starts, laughter, pauses. Use it for conversation analysis, discourse analysis, or legal contexts where the exact wording matters.
Intelligent verbatim removes fillers and false starts but keeps the speaker’s own words and meaning intact. This is the default for most thematic analysis, journalism, and user research.
Edited (clean read) rewrites for readability: full sentences, no timestamps unless you want them. Use it for published Q&As and internal summaries — never for coding data, because the editing is interpretation.
How to transcribe an interview manually, step by step
If you go manual, a little setup saves hours:
- Set up a player with hotkeys. Use any audio player that supports global shortcuts (VLC works, dedicated transcription players are better) and map pause/resume and “skip back 5 seconds” so your hands never leave the keyboard. A foot pedal is worth it past ten hours of audio.
- Slow playback to 75–85%. Fast enough to keep rhythm, slow enough that you stop rewinding constantly.
- Build the document template before you start. Speaker labels, timestamp interval, header with participant code, date, and duration.
- First pass: keep moving. Type roughly and mark unclear passages as [inaudible 00:14:32] instead of looping the same three seconds ten times.
- Second pass: resolve the marks. Fix the inaudibles, check names and technical terms, tidy speaker attributions.
- Final proofread against the audio at full speed. This is where dropped negations (“did” vs “didn’t”) get caught.
Budget four to six hours per hour of clean two-speaker audio. Group interviews, heavy accents, or bad room acoustics push that well past six.
Formatting transcripts for qualitative research
Transcription in qualitative research has one extra requirement: the transcript is data, so it needs to be consistent enough to code. A few conventions cover most protocols:
- Speaker IDs, not names. Use codes like I: for interviewer and P01, P02 for participants — anonymization starts in the document itself.
- Timestamps at a fixed interval (every minute, or every speaker turn) so you can jump back to the audio while coding.
- Line numbers if you cite lines or code on paper — easy to add in Word or LibreOffice via layout settings.
- Plain formatting. Fancy styles and tables break imports. One file per interview, with a naming convention like
P01_2026-07-02.docx.
Analysis software like NVivo and ATLAS.ti works on transcripts, not audio: both import text or Word files, and NVivo can link transcripts to their media files with synced timestamps. NVivo has also offered its own transcription add-ons over the years — those are cloud-based and paid, so check the current docs and your protocol before relying on them.
Consent, IRB rules, and anonymization
Before any recording starts: consent. Recording laws vary by country and by US state — some require all parties to agree, not just you — so always tell participants they are being recorded and capture that consent on the recording or the consent form. This applies to every method on this page.
For academic work, your IRB or ethics board may also constrain the tool choice. Several university IRBs restrict uploading identifiable research audio to consumer cloud transcription services, and some recommend local, on-device transcription or institution-approved vendors instead. If your protocol promised participants that recordings stay on encrypted university equipment, a cloud AI subscription can put you in breach of your own consent form. Check your protocol before you check the pricing pages.
Anonymization basics: replace names with participant codes during transcription, remove or generalize identifying details (employers, small towns, unusual job titles), keep the re-identification key in a separate encrypted file, and delete the audio when your protocol says to — not when you get around to it.
Transcribing interviews locally on a Mac
This is the workflow we built Eidetic for, so read this section knowing it’s our app. Eidetic is a menu-bar app for Apple silicon Macs that records in-person interviews through the microphone, or both sides of a remote call — Zoom, Meet, Teams, or a plain browser call — by capturing system audio directly. Nothing joins the meeting, which matters when a visible bot would change how a participant talks; we wrote about that in how no-bot note takers work.
During the interview you get a live transcript beside your own notes, so you can mark a strong quote the moment it happens instead of hunting for it later. Transcription, editable summaries, and a built-in chat all run on-device: nothing is uploaded, and it works offline after the initial model download. Afterwards you can ask the chat to pull quotes on a theme across one interview or your whole library, then export to Obsidian if that is where your research notes live.
Honest limits: it is Apple silicon only, it is single-user (no shared team library), and the local transcription works best when the interview stays in one language — mixed-language sentences are a known weak spot. Pricing is €49 one-time after 6 free meetings, with 3 device slots — unlike every cloud meeting service, there is no subscription.
FAQ
How long does it take to transcribe an interview?
Manually, plan on four to six hours per hour of audio — a 45-minute interview is an afternoon. Human services typically return files in a day or a few. AI transcription, cloud or local, takes minutes, but budget another 15–30 minutes per interview to fix speaker labels and check names and terms.
How much does it cost to transcribe an interview?
Human services run roughly $1–2 per audio minute, so $60–120 for a one-hour interview. Cloud AI is subscription-based: Otter is $8.33/month billed annually ($16.99 monthly), Fireflies around $10/month. Local tools are one-time purchases — MacWhisper is €59, Eidetic €49. Manual transcription is free in money and expensive in time. Prices checked July 2026.
How do you transcribe an interview for a dissertation or qualitative research?
Start from your ethics protocol, not the tool: if your IRB restricts cloud services for identifiable audio, choose manual or local AI transcription. Use intelligent verbatim unless your method requires full verbatim, apply consistent speaker codes and timestamps, anonymize as you go, and keep one plain-formatted file per interview so NVivo or ATLAS.ti imports cleanly.
What does a transcribed interview look like?
A header with the participant code, date, and duration; then alternating speaker labels with timestamps, like the samples earlier in this guide. Verbatim keeps every um and false start; intelligent verbatim reads cleanly while preserving the speaker’s words; an edited transcript reads like published prose.
How do journalists transcribe interviews?
Most record (with the source’s knowledge — recording consent laws still apply), run an AI first pass, and then verify every quote they plan to print against the audio. For sensitive sources, many avoid uploading the recording anywhere at all, which is where local transcription or old-fashioned manual typing comes back in.

