Transcription
ChatGPT how-to 23 min read

Can ChatGPT Transcribe Audio? (2026): Files, Video, Record Mode and Limits

Yes, on paid plans: upload audio up to 512 MB. Free can't. Video, Record mode, the API's 25 MB limit, speaker labels and privacy, checked on OpenAI's pages.

Updated October 7, 2026

Yes, ChatGPT can transcribe audio: on any paid plan you can upload a recording (MP3, WAV, M4A and other audio formats, up to 512 MB) and ask for a transcript, and the microphone icon (Dictation) turns your own speech into text. The Free plan can't upload audio files, ChatGPT accepts video files but OpenAI says it may not interpret their audio accurately, and developers can send files up to 25 MB to OpenAI's transcription API.

On a Mac, ChatGPT's Record mode also transcribes meetings and voice notes, up to 4 hours a session, on Plus and higher plans. This page covers each way, with the file types, limits, plans, speaker labels and privacy rules for each.

Scribbl, an AI note taker for Google Meet, Zoom and Microsoft Teams, makes this page. If the audio you want transcribed is a meeting, Scribbl does it for every call, on Windows or Mac, with the video kept. Every ChatGPT and OpenAI fact below comes from OpenAI's own help center, API docs and pricing pages, checked on October 7, 2026, and listed under Sources.

512 MB
Largest audio file you can upload to ChatGPT (paid plans only)
4 hours
Longest Record mode session in the ChatGPT app for Mac
25 MB
Largest file OpenAI's transcription API accepts
Limits from OpenAI's help center and API docs, checked October 7, 2026.

Can ChatGPT transcribe audio?

Yes. ChatGPT has several ways to turn speech into text, and OpenAI sells the same kind of transcription to developers through its API. Which one you use depends on whether the audio is a file, your own voice, or a meeting happening now.

WayWhat it transcribesWherePlansLimitSpeaker labels
Upload an audio fileA recording you attach to a chatChatGPT; OpenAI says availability can vary by app version and regionPaid plans; not Free512 MB per file"May be unreliable," says OpenAI
Upload a video fileChatGPT may analyze a video you attachChatGPT, depending on platform and upload methodFree and paidCounts toward your plan's upload limitsNot stated; it "may not" interpret the audio accurately
DictationYour speech, as a message you can editThe microphone iconNot stated on OpenAI's Dictation pageNot statedOne speaker: you
VoiceA spoken conversation with ChatGPT; a transcript is added to the chat afterwardWeb, iOS and AndroidFree (limited) and paidDaily Voice limits by planNot built for several speakers; transcripts are "not verbatim"
Record modeMeetings, brainstorms and voice notesChatGPT desktop app for Mac onlyPlus, Pro, Business, Enterprise, Edu4 hours per sessionYes; unknown voices show as "Speaker 1" until you rename them
OpenAI's APIFiles you send from your own code or appAnywhere you can run codeAn API account, billed separately from ChatGPT25 MB per fileYes, with the gpt-4o-transcribe-diarize model
Whisper (open source)Files on your own computerMac, Windows or Linux, with PythonFree (MIT license)None stated; larger models need more video memoryNot mentioned in its README

Sources for each row are under Sources. The upload is the closest thing to a regular transcription service: you give ChatGPT a finished recording and get the words back. The others either work only on your own voice (Dictation, Voice), only live on a Mac (Record mode), or need some technical setup (the API and Whisper).

Can ChatGPT transcribe an audio recording?

Yes, on any paid plan. OpenAI's help page on uploads now lists audio next to documents and spreadsheets, and says you can "transcribe and discuss an audio recording." It also says: "Audio uploads are not available on the Free plan at this time."

OpenAI's pricing page lists Go, Plus, Pro, Business and Enterprise as paid plans; Go is the lowest-priced, at $8 a month, and Plus is $20 a month.

Supported audio files: WAV, MP3/MPEG, OGG/OGA, audio-only WebM, PCM, FLAC, AAC, M4A and audio-only MP4, up to 512 MB each. The file must contain audio ChatGPT can decode. MP4 and WebM files that are identified as video aren't accepted as audio uploads (see the video section).

How to transcribe a recording in ChatGPT:

  1. Sign in on a paid plan and open a new chat.
  2. Attach the file: select the + icon in the prompt area and choose Add photos & files, or drag the file into the chat.
  3. Ask for a transcript. Paste the prompt below and fill in the brackets.
  4. Ask follow-up questions in the same chat: a summary, the action items, a translation, or "What did they say about the deadline?"
  5. Check the transcript against the recording before you quote it or send it. OpenAI says: "Transcripts may contain errors, and speaker identification may be unreliable."
Transcribe this recording word for word, in its original language. Start a new paragraph each time the speaker changes and label speakers [names, or Speaker 1, Speaker 2]. Don't summarize, shorten or fix grammar. Where you can't make out a word, write [inaudible]. Words you'll hear: [names, product names, acronyms].

The last line matters: listing the names and terms in the recording is the same idea OpenAI recommends for its API, where you can pass expected keywords to improve how domain terms are transcribed.

Long recordings. OpenAI says longer recordings may be processed in smaller sections when Data Analysis is available, and that "very long recordings may time out." If a long file fails, split it into parts (one per hour, for example) and upload them in order in the same chat.

Can ChatGPT transcribe video to text?

Not reliably from the video file itself. OpenAI's Image Inputs FAQ says ChatGPT can accept video files as attachments, including on the Free plan, but adds that ChatGPT "may not analyze the entire video or accurately interpret its audio." Video attachments also count toward your plan's file-upload limits.

The route OpenAI documents for transcription is the audio upload, and it doesn't take MP4 or WebM files that are identified as video. So for a transcript you can trust, give ChatGPT the video's audio track instead.

What you give itFree planPaid plansDo you get a dependable transcript?
The video file, as isAcceptedAcceptedNot reliably: OpenAI says it may not interpret the audio accurately
The video's audio track (M4A, MP3 or another audio format)Not availableAccepted, up to 512 MBYes: this is OpenAI's documented transcription route
The video file sent to OpenAI's API (MP4 or WebM)Needs an API accountNeeds an API accountYes, for files up to 25 MB

Live camera video in ChatGPT Voice is a different feature: on iOS and Android, with the Advanced Voice option, you can share your camera during a conversation. It doesn't transcribe a video file.

Can I use ChatGPT to transcribe a video?

Yes, if you give it the audio. It takes two steps.

1. Save the video's audio as its own file.

  • On a Mac: open the video in QuickTime Player and choose File › Export As › Audio Only. Apple says this saves an MPEG-4 audio file with AAC audio, which is a format ChatGPT accepts.
  • On any computer: the free, open-source tool ffmpeg does it in one line. The -vn option leaves out the video:
ffmpeg -i video.mp4 -vn audio.m4a

2. Upload the audio file to ChatGPT on a paid plan and use the transcription prompt above.

YouTube videos. OpenAI's help pages don't describe transcribing a video from a YouTube link. If a YouTube video has a transcript, YouTube shows it: click Show transcript in the video's description (YouTube Help). For a video you own, export the audio from your copy and upload that.

Can ChatGPT do video transcription?

For a plain transcript of what's said, yes, through the audio track as above. For subtitles, it depends on where you do it:

  • In ChatGPT: OpenAI's upload help doesn't describe timestamps or subtitle files for transcripts. If you ask ChatGPT to add timestamps to a plain transcript, check them against the video before you publish them.
  • In OpenAI's API: the whisper-1 model returns word or segment timestamps, and can return the transcript as an SRT or VTT subtitle file. Files can be up to 25 MB, and it costs $0.006 a minute. The API section has the details.

For a long video, compress the audio first so it fits the 25 MB limit, or split it into parts (the next section shows how many minutes fit).

How long will ChatGPT transcribe audio?

It depends on the feature. OpenAI states most limits as file sizes, not minutes:

FeatureLimit OpenAI statesWhat happens at the limit
Audio upload in ChatGPT512 MB per fileLong recordings may be processed in sections when Data Analysis is available; very long ones may time out
Record mode (Mac app)4 hours (240 minutes) per sessionThe session stops on its own and the notes are saved as a private canvas
Voice (Live option), per rolling 24 hoursFree: limited. Go: 3 hours (GPT-Live-1 mini). Plus: 3 hours. Pro ($100 a month): 15 hours. Pro ($200 a month): unlimitedChatGPT tells you when you reach the limit
DictationNot stated on OpenAI's Dictation page
OpenAI's transcription API25 MB per fileCompress the audio or split it into files of 25 MB or less

How many minutes fit depends on how the audio is saved. This is our own arithmetic (file size divided by bitrate), not an OpenAI figure:

Audio formatSize per minuteFits in 25 MB (API)Fits in 512 MB (ChatGPT upload)
MP3 or M4A at 64 kbps (fine for speech)About 0.48 MBAbout 52 minutesAbout 17 hours
MP3 or M4A at 128 kbpsAbout 0.96 MBAbout 26 minutesAbout 9 hours
WAV, uncompressed (16-bit, 44.1 kHz, stereo)About 10.6 MBAbout 2.4 minutesAbout 48 minutes

So a 2-hour meeting saved as uncompressed WAV is too big for even the 512 MB upload, while the same meeting as a 64 kbps M4A or MP3 is under 60 MB. Converting to a compressed format first is the simplest way to fit more minutes under either limit.

Transcribing audio with the OpenAI API

If you have many files, or you want transcription inside your own app, OpenAI's API does it directly. You send a file to the /v1/audio/transcriptions endpoint and get text back. API usage is billed separately from any ChatGPT plan.

Files: up to 25 MB, in mp3, mp4, mpeg, mpga, m4a, wav or webm (OpenAI's API reference also lists flac and ogg).

ModelUse it forPrice
gpt-transcribeOpenAI's recommended model for recorded speech in its original language. Takes a prompt, expected keywords and language hints$0.0045 a minute
gpt-4o-transcribeEarlier general model; OpenAI says it improved word error rate over its original Whisper modelsAbout $0.006 a minute
gpt-4o-mini-transcribeLower-cost transcriptionAbout $0.003 a minute
gpt-4o-transcribe-diarizeSpeaker labels: who spoke when, with start and end times. Can match up to 4 known speakers from 2 to 10 second sample clips$2.50 per 1M audio input tokens, $10 per 1M output tokens
whisper-1Word or segment timestamps, SRT and VTT subtitle files, and translation into English. Powered by OpenAI's open-source Whisper V2 model$0.006 a minute
gpt-live-transcribeAudio that's still arriving, such as a live call or stream (Realtime API)About $0.017 a minute

A minimal Python request, following OpenAI's file transcription guide:

from openai import OpenAI

client = OpenAI()
with open("recording.m4a", "rb") as audio_file:
    transcript = client.audio.transcriptions.create(model="gpt-transcribe", file=audio_file)
print(transcript.text)

Speaker labels need gpt-4o-transcribe-diarize with response_format="diarized_json", and for audio longer than 30 seconds, chunking_strategy="auto". That model doesn't accept a prompt. Timestamps (the timestamp_granularities[] option) only work with whisper-1, which supports 98 languages. Translation into English uses whisper-1 at the /v1/audio/translations endpoint.

Whisper, free on your own computer. OpenAI also publishes Whisper as open-source software under the MIT license. It runs locally with Python and the ffmpeg tool, so the audio never leaves your computer. The README lists model sizes from about 1 GB of video memory (tiny) up to about 10 GB (large), and doesn't mention speaker labels. It's the free option for people comfortable with a command line.

Can ChatGPT transcribe my meeting?

Yes, if you're on a Mac with a Plus, Pro, Business, Enterprise or Edu plan. ChatGPT's Record mode is "only available for the macOS desktop app": click Record at the bottom of any chat, and ChatGPT transcribes as people talk, then writes notes you can turn into an email or a plan. Sessions stop at 4 hours. It hears the call through your Mac's microphone and system audio, so it works with a Zoom, Teams or Google Meet call on that Mac.

Things to know before you rely on it:

  • No recording is kept. OpenAI deletes the audio after transcription, so there's nothing to check a line against later.
  • Speaker names: voices ChatGPT can't identify show as "Speaker 1" and so on until you rename them.
  • Calendar auto-join works for Google Meet links only. For Zoom, Teams, Webex and other links, you start Record by hand.
  • Pro and Business plans can also use OpenAI's newer Meetings plugin, which OpenAI is still testing. It takes notes on a Mac for online calls or in-person conversations, without a bot joining the call. OpenAI lists iOS, Android and Windows support as "coming soon."
  • On Windows, the web or a phone, ChatGPT can't listen to a meeting. It can only work from a recording or a transcript you give it.

Whatever you use, tell everyone you're recording and get their OK first. OpenAI's Meetings help page says its consent reminder is visible only to you and doesn't ask the other people for you. For step-by-step Record mode instructions, prompts for minutes and a comparison with other ways to get notes, see ChatGPT meeting notes. For the rules on consent, see do you have to tell people you're recording.

How accurate is ChatGPT transcription?

OpenAI doesn't publish an accuracy figure for ChatGPT's transcripts. What it does say, on its own pages:

  • ChatGPT "may make mistakes, including in its transcriptions" (Record mode page).
  • "Transcripts may contain errors, and speaker identification may be unreliable" (uploads page).
  • Voice transcripts "are not verbatim records and may not exactly match what was said" (Voice page).
  • Record mode "works best in English today," and audio upload performance "may vary across languages."
  • For the API, OpenAI says gpt-4o-transcribe improved word error rate over its original Whisper models, without giving a number on the model page.

To get a cleaner transcript, OpenAI's Record mode troubleshooting suggests using a headset and lowering background noise. List the names and terms you expect in your prompt, and check every name, number and quote against the recording before you use it.

Speaker labels, feature by feature: Record mode tells speakers apart and lets you rename them. Uploaded files may get speaker labels, but OpenAI calls them unreliable. Voice is "not yet optimized for conversations with multiple speakers." In the API, gpt-4o-transcribe-diarize is the model built for labeling speakers.

Is ChatGPT transcription private?

It depends on the feature, and on your plan's data settings:

FeatureWhat happens to the audioUsed to train OpenAI's models?
Audio uploadThe file follows your chat and Library retention. Deleting a chat doesn't delete a copy saved in LibraryDepends on the service and your data settings; OpenAI says it doesn't use content from its business offerings, such as the API and ChatGPT Enterprise, to improve its models
DictationKept as long as the chat is in your history; deleted within 30 days after you delete the chat, unless needed for security or legal reasonsOnly if you chose to share audio ("Include your audio recordings")
VoiceLive and Advanced audio clips are kept with the transcript for 30 days. Standard deletes the audio after transcriptionClips only if you chose to share them; transcripts may be used if Improve the model for everyone is on
Record modeDeleted after transcription. Transcripts and notes follow your chat retentionAudio never. Transcripts and notes may be, on Free, Plus and Pro, if Improve the model for everyone is on
Meetings pluginDeleted from your Mac and OpenAI's servers once the notes are readyNot stated on the plugin's help page
OpenAI API (/v1/audio/transcriptions)No abuse-monitoring or application-state retentionNo

Business, Enterprise and Edu workspaces are excluded from model training by default. On Free, Go, Plus and Pro, you can turn off Improve the model for everyone in Settings › Data Controls before you upload anything sensitive. If the recording includes other people, get their consent before you record and before you upload it.

Which way should you use?

  • A recording you already have, and a paid ChatGPT plan: upload it and use the prompt above.
  • A video: export the audio first, then upload that.
  • On the Free plan: Whisper on your own computer, or one of the tools in the best free transcription software.
  • Hundreds of files, subtitles or speaker labels: OpenAI's API.
  • Your Google Meet, Zoom or Teams calls, on Windows or Mac: Scribbl (next section). You get the video, a transcript with speaker names, AI notes and action items for every meeting, free for 10 meetings a month, with no paid ChatGPT plan needed.
  • A single meeting on a Mac with Plus or higher: Record mode, though it keeps no recording to check against.
  • An interview: see how to transcribe an interview for formats and checking steps.

The easier way for meetings: Scribbl

This section is about our own product. If the audio you want in writing is a meeting, you don't need to record it, export it and upload it to ChatGPT. Scribbl is an AI note taker for Google Meet, Zoom and Microsoft Teams that does it for every call. Nothing joins the call: Google Meet runs from a Chrome extension, and Zoom and Teams run from Scribbl Desktop for Mac (Apple silicon) and Windows. After each meeting you get the video recording, a transcript with speaker names, AI notes and action items.

  • Ask ChatGPT or Claude about every meeting. On Pro, connect Scribbl once and ChatGPT or Claude can search the transcripts and notes of all your meetings, so you can ask what a client said about the deadline last month without uploading a file.
  • The follow-up done for you. Automations draft the follow-up email, post a summary to Slack or Teams, log notes in HubSpot or Pipedrive, or save them to Notion. They ask before changing anything by default.
  • Meetings sorted by themselves. Smart Collections put each new meeting into the groups you describe, so you can ask about just one group, such as "What were the main objections in my sales calls this week?"
Scribbl's sample meeting Atlas launch planning in the Scribbl web app: a summary with five bullet points, a card for 5 action items, the Notes, Transcript, Action items and AI Chat tabs, a Copy transcript button, the meeting video player and a timeline of how long each of the two speakers talked
A meeting in the Scribbl web app (Scribbl's sample meeting, Atlas launch planning, on Google Meet): the summary and a card for 5 action items on the left, with tabs for Notes, Transcript, Action items and AI Chat and a Copy transcript button at the top; the meeting video and a timeline of who spoke when on the right.

Why it's easier than ChatGPT for meetings: ChatGPT's Record mode needs a Mac and a Plus or higher plan, and keeps no recording. Uploading a recording needs a paid plan and a file you've already made. Scribbl works on Windows or Mac, keeps the video, and writes the transcript and notes for every Google Meet, Zoom and Teams call on its own. The free plan includes Google Meet, Zoom and Microsoft Teams, with 10 meetings a month. Pro is $13 per user per month billed annually, or $20 monthly (pricing). Recordings are private by default, and Scribbl doesn't use your meetings to train AI for anyone other than you. As with any recorder, tell everyone you're recording.

Get Scribbl for Google Meet

FAQ

Can ChatGPT transcribe audio?

Yes. On a paid ChatGPT plan you can upload an audio file (WAV, MP3, M4A, FLAC, AAC, OGG and other formats, up to 512 MB) and ask for a transcript. The microphone icon (Dictation) turns your own speech into text, and the ChatGPT app for Mac can transcribe meetings and voice notes with Record mode on Plus and higher plans. The Free plan can't upload audio files. Developers can use OpenAI's transcription API, which takes files up to 25 MB.

Can ChatGPT transcribe an audio recording?

Yes, on any paid plan. Attach the recording to a chat and ask for a word-for-word transcript. OpenAI lists WAV, MP3/MPEG, OGG/OGA, audio-only WebM, PCM, FLAC, AAC, M4A and audio-only MP4 files up to 512 MB. Very long recordings may time out, and OpenAI says transcripts may contain errors, so check names and numbers against the recording.

Can ChatGPT transcribe video to text?

Not reliably from the video file itself. ChatGPT accepts video attachments, even on the Free plan, but OpenAI says it may not analyze the entire video or accurately interpret its audio. For a dependable transcript, save the video's audio as an M4A or MP3 file and upload that on a paid plan, or send the file to OpenAI's API, which accepts MP4 and WebM files up to 25 MB.

Can I use ChatGPT to transcribe a video?

Yes, if you give it the audio. On a Mac, open the video in QuickTime Player and choose File, then Export As, then Audio Only. On any computer, the free ffmpeg tool does the same with: ffmpeg -i video.mp4 -vn audio.m4a. Then attach the audio file to a ChatGPT chat on a paid plan and ask for a transcript.

Can ChatGPT do video transcription?

For a plain transcript of what's said, yes, through the video's audio track. For subtitles, OpenAI's help pages don't describe timestamps or subtitle files for ChatGPT uploads. OpenAI's API does: the whisper-1 model returns word or segment timestamps and SRT or VTT subtitle files, for files up to 25 MB, at $0.006 a minute.

Can ChatGPT transcribe a YouTube video to text?

OpenAI's help pages don't describe transcribing a video from a YouTube link. If the video has a transcript, YouTube shows it: click Show transcript in the video's description. For a video you own, save its audio and upload that file to ChatGPT on a paid plan.

Can ChatGPT transcribe my meeting?

Yes, if you use the ChatGPT desktop app for Mac on a Plus, Pro, Business, Enterprise or Edu plan. Record mode transcribes and summarizes up to 4 hours per session, labels voices it can't name as Speaker 1 and so on, and deletes the audio afterward. On Windows, the web or a phone, ChatGPT can only work from a recording or transcript you give it. For every Google Meet, Zoom or Teams call on Windows or Mac, with the video kept, Scribbl (which we make) does it on its free plan. Tell everyone you're recording first.

Can ChatGPT transcribe audio for free?

Not from a file. OpenAI says audio uploads are not available on the Free plan. Free does include Voice with limited access, but Voice transcripts aren't word for word. For free transcription of files, OpenAI's open-source Whisper model runs on your own computer under an MIT license, though it takes some setup.

How long will ChatGPT transcribe audio?

It depends on the feature. Uploaded audio files can be up to 512 MB, which is about 9 hours of 128 kbps MP3, but OpenAI says very long recordings may time out. Record mode stops at 4 hours per session. OpenAI's API takes up to 25 MB per file, about 26 minutes of 128 kbps MP3, so longer recordings have to be compressed or split.

Can ChatGPT do voice to text?

Yes. Press the microphone icon in ChatGPT (Dictation), speak, and ChatGPT returns your words as text you can edit before you send it. Voice is different: it's a spoken conversation with ChatGPT, and OpenAI says Voice transcripts are not verbatim records.

Can ChatGPT identify different speakers?

Sometimes. Record mode in the Mac app tells speakers apart and labels anyone it can't name as Speaker 1, Speaker 2 and so on, which you can rename. For uploaded files, OpenAI says speaker identification may be unreliable. OpenAI's API has a separate model, gpt-4o-transcribe-diarize, that labels who speaks when and can match up to four known speakers from short sample clips.

Is ChatGPT accurate at transcription?

OpenAI publishes no accuracy figure for ChatGPT's transcripts and says they may contain errors. Record mode works best in English, and results for uploads vary by language. Use a headset or a quiet room, tell ChatGPT the names and terms to expect, and check names, numbers and quotes against the recording.

Does OpenAI keep or train on my audio?

It depends on the feature. Record mode deletes the audio after transcription and doesn't train on it, but its transcripts may be used for training on Free, Plus and Pro if Improve the model for everyone is on. Dictation audio stays with the chat until you delete the chat, then is deleted within 30 days. Business, Enterprise and Edu content isn't used for training by default, and OpenAI's transcription API doesn't train on your audio or keep it for abuse monitoring.

What is the difference between ChatGPT and Whisper?

ChatGPT is OpenAI's chat app. Whisper is OpenAI's speech recognition model, available as whisper-1 in OpenAI's API at $0.006 a minute and as free open-source software you run on your own computer. For most API transcription, OpenAI now recommends its newer gpt-transcribe model and keeps whisper-1 for timestamps, subtitles and translation into English.

Sources

Every product fact above was checked on these pages on October 7, 2026. OpenAI changes these features often, so check the linked page if a detail matters to you.

Try Scribbl

Let your meetings take their own notes.

Scribbl records, transcribes, and summarizes your Google Meet calls from your browser. No bot joins the call. Free forever for individuals.

Add to Chrome · It's free

Record meetings with Scribbl

No bot. No host permission.

Start free