How we score video difficulty
Every number on this site is computed from a video’s actual subtitle track. Nothing is a rating, a vote, or an editor’s impression — which also means every number can be argued with, and this page is where you find the grounds.
What the score means
One number from 0 to 100 that orders content from easiest to hardest to follow. Two cuts divide it into the three shelves used across the site.
The score and the band measure different things
Alongside the score, every video carries a band — A2, N4, HSK4. The band answers “how much vocabulary does this assume?”. The score answers “how hard is this to keep up with?”. A band is never derived from a score, or the reverse.
Each language uses the framework its learners actually think in: CEFR for English, French and Spanish, JLPT for Japanese, HSK for Chinese. Translating a JLPT level into CEFR would lose more than it explains, so we don’t.
The two ends of the band are two different questions. The left one is the level at which you would already know 80% of the words — enough to follow along, with the picture and the context filling the gaps. The right one is where you would know 95%, which is where watching stops being work. Those two numbers are not ours: they come from the first study to measure lexical coverage on video rather than on reading (Durbahn, Rodgers, Macis & Peters, 2024), which found viewers reach adequate comprehension at lower coverage than readers do, because the imagery carries part of the meaning. Note that coverage is not comprehension — knowing 80% of the words bought about 71% comprehension in that study, not 80%.
How the score is computed
We read the subtitle track, measure two things about it, and weight them. That is the whole pipeline — there is no model judging content quality and no human in the loop.
Where the numbers stop being trustworthy
A measurement whose limits are published can be argued with; one whose limits are hidden can only be believed or dismissed. These are ours.
- Scores compare within a language, not across languages. Word rarity is measured against each language’s own corpus, so a 40 in Spanish is not the same experience as a 40 in Japanese.
- Japanese bands read about a level too hard. Deciding which JLPT level a word belongs to needs thresholds on a Japanese word-frequency list, and ours are borrowed from the English calibration, which measures roughly two-thirds of a level harsh. That is why N1 is the most common label here, on videos plainly made for intermediate learners. Trust the score and the ordering; read the Japanese band as one notch pessimistic.
- The two ends of the band mean different things in different languages. English, Japanese and Chinese have a published word list that grades vocabulary by teaching level — CEFR-J, JLPT and HSK 3.0 — so their bands use the 80%/95% pair above. The rest (Spanish, French, Italian, Portuguese, Russian, Korean, German) have no such list we can license, so their levels are inferred from how common a word is, on thresholds that were themselves fitted at 95%. Moving those to 80% without redoing that work produced bands as wide as A1–C2, which says nothing at all, so those languages keep the older, stricter 95%/98% pair until their thresholds are refitted. Compare bands within a language, never across two.
- The band is a vocabulary reading, and it will not sort videos for you. We checked it against the level creators themselves put in their video titles — about 3,300 videos across five languages — and the band barely moves with it, in any language and at any threshold we tried. The reason is that it answers “what is the hardest thing you will run into” while a creator means “what is this mostly made of”. In one video labelled HSK1 by its own author, 91% of what is said is HSK1–3 — and the band is set by the remaining 9%. The difficulty score is the number that orders content; read the band as what it is.
- Automatic subtitles carry the machine’s errors into the score. Videos whose track was auto-generated rather than written are marked as such wherever they are listed, so you can discount them yourself.
- The sample is a curated set of channels, not a random slice of YouTube. Everything here was measured, but not everything on YouTube was measured.