Methodology

How the score is calculated

A thumbnail test is only useful if you can trust the score. This page explains exactly what each number measures, where it comes from, and what it does not tell you.

The five cover dimensions

The tool reads each cover image locally in your browser and turns what it sees into five scores from 0 to 100. No single number is the answer; together they show how each cover maps to the tool's fixed pixel rules.

  1. Salience
    How strongly the cover contrasts and how saturated its palette is.
    Measured byCombines global contrast (how spread the pixel luminance is across the frame) with saturation. A flat or washed-out cover scores low; a punchy, contrasty one scores high.
  2. Emotional pull
    The skin-tone and warm-color shares used by this heuristic.
    Measured byEstimates the share of skin-tone pixels as a rough proxy for a visible person, combined with the share of warm colors. It does not recognize faces or emotions.
  3. Color quality
    Whether the palette is balanced and the brightness sits in a viewable range.
    Measured byCombines a harmony score (how far saturation and warmth sit from a neutral baseline) with a brightness score. Oversaturated and near-monochrome covers both drift lower than a deliberate two-tone palette.
  4. Focus and composition
    The share of pixels that fall inside the scorer's skin-tone range.
    Measured byUses skin-tone pixel share as a rough subject proxy. It does not analyze spatial layout, detect the rule of thirds, or understand non-human focal subjects.
  5. Readability
    Whether brightness and contrast remain distinguishable at a small feed-like size.
    Measured byDownscales the cover to 320 pixels wide and combines overall brightness and contrast. It does not run OCR or test whether text can actually be read.

The output

What the five dimensions produce

Each cover leaves the scoring pass with five 0 to 100 dimension scores and one combined total. That total, paired with the title total, is what the preference model ranks. The card below is the artifact the tool returns for every variant.

How to read itThe micro-bars are the five cover dimensions and the big number is their weighted total. The model-share pill comes from the deterministic preference model further down the page.

Layout example only. These cover values are illustrative. The reproducible numeric example appears in the title section below.

Sample YouTube thumbnail, option B: a celebrating editor with fists raised and a bold 300%! triumph overlay, the shock-style alternate variant used in the A/B test demo.

Example cover

78/100
B
  • Salience82
  • Emotional pull71
  • Color quality66
  • Focus and composition74
  • Readability58

Lift cover contrast to sharpen the focal subject.

The seven title angles

A title is scored across seven angles, because the title and the cover are read as one unit in the feed. The scorer also detects the title's style (listicle, question, how-to, versus, number-led). These are rule-based pattern matches, not a semantic reading of audience intent.

Opening hook
Whether the first few words create curiosity or promise a payoff.
Emotion
Whether the title triggers a feeling such as surprise, outrage, relief, or FOMO.
Conflict
Whether there is tension, a versus, or an unexpected contrast.
Length
Whether the title is short enough to read fully in the feed. Long titles get truncated.
Structure
Whether the title follows a pattern recognized by the scorer, such as listicle, question, how-to, or versus.
Keywords
Whether the title contains explicit topical or numeric cues recognized by the scorer.
Readability
Whether the title is easy to parse at a glance, or convoluted.

The engine room

How the cover and title totals combine

Every cover and every title is reduced to one total with a fixed set of weights. The two totals are averaged into the 0 to 100 combined preflight score used for ranking. No audience or channel data enters this calculation.

Cover total 5 dimensions

  • Salience30%
  • Emotional pull25%
  • Color quality20%
  • Focus and composition15%
  • Readability10%

Plus a deterministic fingerprint of plus or minus 5, derived from the cover's feature values, so near-identical covers do not collide on the same total.

Title total 7 angles

  • Opening hook28%
  • Emotion22%
  • Conflict15%
  • Length10%
  • Structure10%
  • Keywords8%
  • Readability7%

No fingerprint. The title total is the plain weighted sum of the seven angle scores.

Combined preflight score

( cover total + title total ) ÷ 2

This is a unitless comparison score. A value such as 81 means 81 out of 100 under these fixed rules, not an 81% chance of success.

Worked example

One title, scored by the live rules

The input below is the first built-in sample in the comparison tool. During the build, this page passes it to the same title scorer and multiplies each dimension by its published weight.

Shocking: How an Everyday Editor Tripled Output in 30 Days
Detected angle
Shock-led
Title grade
B
Title total
74.5 / 100

The contribution column is dimension score × weight. The seven contributions sum to 74.47, which the scorer rounds to 74.5. It is a title heuristic score, not a predicted click rate.

Reproducible title-score calculation
DimensionScoreWeightContribution
Opening hook6328%17.64
Emotion8522%18.7
Conflict6515%9.75
Length10010%10
Structure7010%7
Keywords678%5.36
Readability867%6.02
Rounded title total74.5

The 10,000 deterministic model iterations

Each cover and title pair enters a preference model built only from the combined scores. A seeded PRNG runs 10,000 repeatable allocations across the options. The same inputs produce the same output.

The output is a model share. It is useful for relative ordering inside this set, but the iterations are not people, impressions, or clicks. They restate the scores at a finer resolution and add no new evidence about performance.

What the preflight can and cannot tell you

These seven questions define the score's limits and the decisions it can responsibly support.

Does the preflight score predict my YouTube CTR?

No. The score compares concepts under fixed image and title rules. It has not been calibrated against your channel's click-through data and is not a probability or CTR forecast.

How is the combined preflight score calculated?

Each cover is scored across 5 visual dimensions, weighted: salience 30%, emotional pull 25%, color quality 20%, composition 15%, readability 10%. Each title is scored across 7 angles, weighted: hook 28%, emotion 22%, conflict 15%, length 10%, structure 10%, keywords 8%, readability 7%. The cover and title totals are averaged for the combined score. A small deterministic fingerprint of plus or minus 5 keeps near-identical covers from colliding.

Does the title scorer work equally for English and Chinese titles?

Both languages use the same 7 dimensions and weights, with separate pattern sets. English recognizes how-to, suspense, surprise, question, comparison, list, number, and conflict cues. Chinese recognizes its own question, surprise, suspense, number, and comparison markers. Both remain rule-based heuristics, not semantic audience research.

How accurate is the ranking?

We do not claim an accuracy percentage because the ranking has not been validated against your channel's real data. What it guarantees is repeatability: identical inputs produce identical outputs. That makes revision comparisons consistent, but it does not establish how people will respond.

How is this different from YouTube's own Test & Compare?

YouTube's native test uses real watch-time data after publication with up to three options. This tool compares up to six concepts before publication using fixed heuristics. Use the preflight to build a shortlist, then use YouTube to validate it with real behavior.

What do the 10,000 iterations represent?

They are deterministic model allocations derived from the combined scores. They are not people, impressions, clicks, or a research sample. The count gives the model enough resolution to display a stable relative share.

Can the score be gamed?

Yes, to a point. A concept engineered to maximize these fixed signals can still fail if the topic is weak or the title is misleading. The score reports only what its rules can observe. A high score is not evidence that a video will perform.

Run the score on your own covers

Upload two to six thumbnails and titles to see 5 visual signals, 7 title signals, the relative ranking, and a first fix. Free, no login required for the core test.

Test your thumbnails