Method

How we score video difficulty

Every number on this site is computed from a video’s actual subtitle track. Nothing is a rating, a vote, or an editor’s impression — which also means every number can be argued with, and this page is where you find the grounds.

What the score means

One number from 0 to 100 that orders content from easiest to hardest to follow. Two cuts divide it into the three shelves used across the site.

BeginnerIntermediateAdvanced02143100
One scale, three shelves, the same cuts in every language. Beginner takes only the first fifth of the axis — that is where the calibration put the cut, not a design choice, and it is a fair warning that most of what exists on YouTube is made for people who already speak the language.

The score and the band measure different things

Alongside the score, every video carries a band — A2, N4, HSK4. The band answers “how much vocabulary does this assume?”. The score answers “how hard is this to keep up with?”. A band is never derived from a score, or the reverse.

SPEAKING PACE, SLOWER → FASTERVOCABULARYSlow lecture, rare wordsband C1 · score 24Fast vlog, everyday wordsband A2 · score 64
This is why a band and a score can disagree. The lecture uses harder words, so its band is higher. The vlog is spoken far faster, so it is the one you will struggle to follow — and the score, which is mostly pace, ranks it harder. Both numbers are true; they answer different questions.

Each language uses the framework its learners actually think in: CEFR for English, French and Spanish, JLPT for Japanese, HSK for Chinese. Translating a JLPT level into CEFR would lose more than it explains, so we don’t.

The two ends of the band are two different questions. The left one is the level at which you would already know 80% of the words — enough to follow along, with the picture and the context filling the gaps. The right one is where you would know 95%, which is where watching stops being work. Those two numbers are not ours: they come from the first study to measure lexical coverage on video rather than on reading (Durbahn, Rodgers, Macis & Peters, 2024), which found viewers reach adequate comprehension at lower coverage than readers do, because the imagery carries part of the meaning. Note that coverage is not comprehension — knowing 80% of the words bought about 71% comprehension in that study, not 80%.

How the score is computed

We read the subtitle track, measure two things about it, and weight them. That is the whole pipeline — there is no model judging content quality and no human in the loop.

The video’ssubtitle trackSpeaking pacewords per minuteVocabulary loadhow rare each word is× 0.8× 0.2Difficulty score0 – 100
Four parts pace to one part vocabulary. Those weights are not a guess. They were fitted against 843 videos whose own titles declare a level — a channel calling its episode “beginner” or “advanced” — across Spanish, Japanese and French, and pace predicted that label far better than vocabulary did in all three. Pace is normalised per language, which is what lets one scale cover writing systems as different as Spanish and Japanese.

Where the numbers stop being trustworthy

A measurement whose limits are published can be argued with; one whose limits are hidden can only be believed or dismissed. These are ours.

  • Scores compare within a language, not across languages. Word rarity is measured against each language’s own corpus, so a 40 in Spanish is not the same experience as a 40 in Japanese.
  • Japanese bands read about a level too hard. Deciding which JLPT level a word belongs to needs thresholds on a Japanese word-frequency list, and ours are borrowed from the English calibration, which measures roughly two-thirds of a level harsh. That is why N1 is the most common label here, on videos plainly made for intermediate learners. Trust the score and the ordering; read the Japanese band as one notch pessimistic.
  • The two ends of the band mean different things in different languages. English, Japanese and Chinese have a published word list that grades vocabulary by teaching level — CEFR-J, JLPT and HSK 3.0 — so their bands use the 80%/95% pair above. The rest (Spanish, French, Italian, Portuguese, Russian, Korean, German) have no such list we can license, so their levels are inferred from how common a word is, on thresholds that were themselves fitted at 95%. Moving those to 80% without redoing that work produced bands as wide as A1–C2, which says nothing at all, so those languages keep the older, stricter 95%/98% pair until their thresholds are refitted. Compare bands within a language, never across two.
  • The band is a vocabulary reading, and it will not sort videos for you. We checked it against the level creators themselves put in their video titles — about 3,300 videos across five languages — and the band barely moves with it, in any language and at any threshold we tried. The reason is that it answers “what is the hardest thing you will run into” while a creator means “what is this mostly made of”. In one video labelled HSK1 by its own author, 91% of what is said is HSK1–3 — and the band is set by the remaining 9%. The difficulty score is the number that orders content; read the band as what it is.
  • Automatic subtitles carry the machine’s errors into the score. Videos whose track was auto-generated rather than written are marked as such wherever they are listed, so you can discount them yourself.
  • The sample is a curated set of channels, not a random slice of YouTube. Everything here was measured, but not everything on YouTube was measured.

← Back to the languages