Skip to main content

Distinguish manual vs auto-generated transcripts in metadata

I want a way to distinguish between manually uploaded transcripts and auto-generated ones in the metadata (specifically in the transcriptLanguages array). Currently, I have to iterate through all available languages using mode=native to find a manual one and avoid 206 errors or unwanted AI generation costs from mode=generate. It is difficult to implement this iteration when there are 50+ languages, and the first language in the list isn't always the native one. Adding a 'type' or 'isManual' flag to each language in the metadata would allow me to target the correct native transcript directly.
Status: Planned4 comments

Log in to comment and vote

Comments4

  • atjackiejohns

    •

    Aug 4

    This would fix a very annoying bug as well if it would include all the transcripts for a given language with metadata.

    When a YouTube video has multiple caption tracks in the same language (e.g. Spanish auto-generated + Spanish CC1/DTVCC1), /v1/transcript with mode=native&lang=es returns only one track and availableLangs collapses them to ["es"]. There is no way to choose which track.

    For news/broadcast uploads, the “manual” tracks are often CEA-608/708 closed captions that are out of sync with the YouTube audio, while the auto-generated (asr) track is correctly timed. Today the API appears to prefer the official/manual track, which makes native Spanish unusable for timed learning/playback on those videos.

    You can check this CNN video as an example:

    https://www.youtube.com/watch?v=UCHG5pJnLxI

    The captions are from CC1 that the API returns and these are totally out of sync plus they’re in caps lock. Ideally I should be able to ignore any CC[#] or DTVCC[#] captions and only get the ones I choose.

    Here’s an snippet of what it returned:

    {  "lang": "es",  "availableLangs": [  "es"  ],  "content": [  {  "lang": "es",  "text": "INFORMACIÓN Y GRACIAS POR ESTAR",  "offset": 1067,  "duration": 2002  },  {  "lang": "es",  "text": "AQUÍ CON NOSOTROS EN CNN ",  "offset": 2135,  "duration": 2001  },
  • martin

    •

    Jan 6

    •

    Merged request

    •

    1 vote

    Distinguish manual vs auto-generated transcripts in metadata

    I want the `transcriptLanguages` array in the metadata response to include a 'type' or 'isGenerated' flag for each language. Currently, it only provides language codes, making it impossible to know which transcripts are manually uploaded versus auto-generated/translated. This forces me to iterate through dozens of languages using `mode=native` to find a valid one without triggering expensive AI generation costs. Having a way to distinguish these in the metadata would allow me to target only manual transcripts efficiently.
  • martin

    •

    Jan 6

    •

    Merged request

    •

    1 vote

    Distinguish manual vs auto-generated transcripts and identify original language in metadata

    I want the transcript metadata (transcriptLanguages array) to include a flag or property that distinguishes between manually uploaded transcripts and auto-generated/auto-translated ones (e.g., "type": "manual" vs "type": "asr"). Additionally, it would be very helpful to have an "originalLanguage" or "isOriginal" flag to identify the video's source language. This is important because currently, using mode=native often results in 206 errors for auto-generated tracks, and without knowing which language is the original, I have to iterate through dozens of languages to find a valid native transcript, which is inefficient and difficult to implement.
  • martin

    •

    Dec 19, 2025

    Or maybe add another array to make the change less invasive where is nativeTranscriptLanguages, so is a new field, and does not impact current api clients