What Percentage of YouTube Videos Have No Captions? We Measured 245,703 of Them
Why nobody actually knows this number
Ask how many YouTube videos have captions and you will find confident-sounding answers ranging from “almost all of them” to “less than half.” None of them cite a measurement, because measuring it is awkward: YouTube’s API tells you whether a caption track is listed, not whether it can actually be retrieved, and there is no public dataset of caption availability.
We are in an unusual position to measure it. youtube-transcript.ai fetches caption tracks for whatever video a visitor pastes in. Every attempt produces a definitive answer: either a usable track came back, or it did not, and we know why.
This post reports what those attempts look like across 245,703 unique videos over 18 weeks (April 20 – August 17, 2026).
The headline number: 13.0%
Of 245,703 unique videos:
| Outcome | Unique videos | Share |
|---|---|---|
| Caption track retrieved | 211,468 | 86.1% |
| No usable captions | 32,051 | 13.0% |
| Other failures (network, rate limits) | 2,184 | 0.9% |
The rate is remarkably stable. Across 18 weekly buckets it never left the 9.7%–15.0% band:
Apr 20 12.9% Jun 08 14.1% Jul 20 13.9% Apr 27 11.8% Jun 15 12.6% Jul 27 13.1% May 04 9.7% Jun 22 13.6% Aug 03 11.7% May 11 10.3% Jun 29 13.5% Aug 10 11.0% May 18 10.3% Jul 06 13.3% Aug 17 12.1% May 25 12.7% Jul 13 13.3% Jun 01 15.0%
Read this as “roughly one in eight.” The week-to-week wobble is real, so quoting 13.04% to two decimals would imply a precision this measurement does not have.
What this number is, and what it is not
This matters more than the number itself.
This is not a random sample of YouTube. It is a sample of videos that someone deliberately brought to a transcript tool. That population is skewed, and the skew almost certainly runs in one direction: if a video already has captions, many people just use YouTube’s own transcript panel and never look for a third-party tool. The people who show up here are disproportionately the ones for whom the easy path already failed.
So 13% is best read as: among videos people actively try to get a transcript for, about one in eight has no usable caption track. The true figure across all of YouTube is probably lower.
Two things do support the number being in a sensible range, though. First, it is stable across 18 weeks and a 10x change in traffic volume. Second, the sample is broad — hundreds of thousands of distinct videos arriving from search, direct visits, and AI assistants, not one niche.
One methodological note we learned the hard way: we also have a downstream workspace page, and its caption-failure rate reads around 64%. That number is an artifact — the homepage routes videos it already knows have no captions to that page, so the downstream sample is pre-filtered. Measure at the point where every video enters, not after any branch has sorted them. We nearly published the wrong figure.
Why videos have no captions
Breaking down the 36,617 videos where we could not deliver a transcript:
| Reason | Unique videos | Share of failures |
|---|---|---|
| No caption track exists at all | 31,976 | 87.3% |
| Other / unclassified | 2,571 | 7.0% |
| Requires sign-in (age-restricted, etc.) | 2,218 | 6.1% |
| Video published <24h ago | 131 | 0.36% |
| Video unplayable in our region | 103 | 0.28% |
| Members-only content | 101 | 0.28% |
The “it’s too new” theory is wrong
The most common explanation you will hear is that YouTube simply has not generated automatic captions yet — give it a few hours. We tag every caption-less video published within the last 24 hours specifically to test this.
It accounts for 0.36% of failures. Out of 36,617 videos we could not transcribe, 131 were plausibly just waiting on YouTube’s caption pipeline.
Automatic caption generation is fast now. If a video has no captions, waiting almost never fixes it.
So what does explain it?
87.3% of failures are videos where no caption track exists at all — not delayed, not restricted, simply absent. From inspecting samples, these cluster into a few recognisable groups:
- Music and performance videos. No speech to caption. YouTube’s automatic captioning does not attempt lyrics.
- Non-speech content. Ambient footage, gameplay without commentary, ASMR, timelapses.
- Languages with weak ASR coverage. YouTube’s automatic captions do not cover every language equally; videos in less-supported languages frequently have nothing.
- Uploads with captions explicitly disabled. A creator setting, and one that is invisible from the outside.
A smaller but distinct group is sign-in required (6.1%) — age-restricted content and similar. These videos often do have captions; they are simply not reachable without an authenticated session.
Can anything be done about it?
Three approaches, in increasing order of effort.
1Check whether captions appeared later
Rarely useful, but not never. In our own testing we picked a caption-less video as a fixture and had to replace it twice because YouTube grew captions for it after the fact. If a video is important and recent, re-checking in a few days costs nothing. Given the 0.36% figure above, do not expect much.
2Check other caption sources
A video without captions in your language may still have them in another. Auto-translated tracks are also frequently available even when a native track is not — quality varies, but it beats nothing.
One warning from experience: YouTube exposes machine-translated variants of every caption track, which can mean hundreds of pseudo-tracks for a single video. Requesting them aggressively is a good way to get rate-limited. Pick the source track, then translate on your side.
3Run speech recognition on the audio
This is the only approach that works on a video with genuinely no caption data. Feed the audio to an ASR model (Whisper and its descendants) and generate the transcript yourself.
We ran this on caption-less videos and it succeeded on 197 of 240 — about 82%.
Modern ASR is also fast enough that length is nearly irrelevant. A 65-minute talk transcribes in about a minute of wall-clock time on GPU — roughly 170x real time — so a two-hour video is not meaningfully harder than a ten-minute one.
What ASR cannot fix — the remaining ~18%
The failures are instructive, because they are not bugs:
- A 3-minute video that was almost entirely music. Voice activity detection kept 23 seconds of usable audio out of 180, and the model misidentified the language from that fragment. The transcript would have been meaningless.
- A video with no detectable speech whatsoever. Nothing to transcribe.
A real slice of the caption-less pool is genuinely untranscribable. It has no captions for the same reason it cannot be transcribed: there is no speech in it. Any tool promising to transcribe any video is either not counting these or not telling you.
What we would want measured next
Two limits of this study are worth naming, because they are the obvious next questions:
- We cannot break the rate down by video language or category. Videos with no caption track also carry no language metadata from the caption endpoint, so the “which languages are worst served” question — probably the most useful one for accessibility — stays open.
- The sample is tool users, not YouTube. A genuinely unbiased measurement would need a random sample of video IDs, which is a different and harder study.
If you are working on video accessibility and want to compare notes on methodology, we would like to hear from you.
Need the transcript of a video right now? Paste the URL — free, no sign-up.
Get a YouTube transcriptMethod: 245,703 unique video IDs, April 20 – August 17, 2026. A video counts once regardless of how many times it was requested; counting requests instead of videos inflates the failure rate, since failed videos get retried more often. Measured at the entry point, before any routing branch. A small number of videos (~1%) appear in both the success and failure sets — typically a transient failure followed by a successful retry — which makes 13.0% a slight overstatement; this is one reason we quote it as “about one in eight” rather than to two decimals. Only aggregate counts are reported here; no individual video or visitor is identified.