Turn a folder of ideas into a defensible shortlist

Compare 2 to 6 cover and title concepts, see which signals hold each one back, and decide what deserves another revision. The score is a deterministic preflight, not audience evidence.

Run a free preflight

The result card

What a scored cover looks like

Every variant comes back as one card: a thumbnail preview, an overall 0 to 100 score, five visual dimensions, a deterministic model share, and a focused first fix. The top candidate is flagged in green.

  • Overall score: five visual dimensions and seven title angles, combined into one number.
  • Five bars: the cover's visual dimensions, read straight from your pixels.
  • Model-share pill: a deterministic allocation across this set, not an audience percentage.
  • Top fix: the weakest dimension, named in priority order.
Sample YouTube thumbnail, option B: a celebrating editor with fists raised and a bold 300%! overlay, shown inside an example preflight score card.

Example result

81/100
ATop candidate
  • Saliency86
  • Emotional pull74
  • Color quality69
  • Composition62
  • Readability55

Lift cover contrast so the focal subject still reads at feed size.

Example card. The numbers illustrate the shape of the result, not a claim about any real cover.

Ready to see the same breakdown on your own concepts?

Compare my concepts

Five-dimension cover scoring

The tool reads saturation, contrast, warmth, luminance, and the share of the frame taken up by skin tones straight from your image on your device, then turns those readings into five scores: salience, emotion, color, composition, and readability.

See scoring details
What it measures
Five weighted scores from your pixels: salience (30%, global contrast and saturation), emotional pull (25%, skin-tone and warm-color share), color quality (20%, palette balance and brightness), composition (15%, share of the frame taken by skin tones), and readability (10%, brightness and contrast after the cover is downscaled to 320 pixels). A small plus or minus 5 fingerprint keeps near-identical covers from colliding on the same total.
Low score means
The lowest of the five drags the total down. A busy, low-contrast cover loses readability; a flat, washed-out one loses salience; a frame with no skin tones loses both composition and emotional pull.
How to raise it
Fix the lowest dimension first. Raise contrast and saturation for salience, adjust visible skin-tone share on people-led covers for composition, and pull brightness into a mid-range so the 320-pixel version retains contrast.

Seven-angle title scoring

Your title gets read across seven angles: hook, emotion, conflict, length, structure, keywords, and readability. A front-loaded curiosity hook and a plain noun phrase do different work, and both get measured.

See scoring details
What it measures
Seven weighted angles: hook (28%), emotion (22%), conflict (15%), length (10%), structure (10%), keywords (8%), and readability (7%). The scorer has separate English and Chinese pattern sets. English recognizes how-to, suspense, surprise, question, comparison, list, number, and conflict cues; Chinese recognizes its own language-specific cues.
Low score means
The title is long, structureless, or front-loaded with low-signal words. A title that runs past the feed's truncation point, or restates the topic without tension, lands low.
How to raise it
Front-load a curiosity hook or a number, keep the title under the truncation length, pick a recognizable pattern (listicle, question, how-to, versus), and include the term that names the topic clearly.

A deterministic preference model

The model converts the combined scores into a relative share across 10,000 repeatable iterations. It does not contact people, observe YouTube traffic, or add evidence beyond the scores already on screen.

See scoring details
What it measures
A 10,000-iteration allocation generated by a deterministic PRNG (mulberry32), seeded from the combined scores and weighted with a power of 2.3. The same inputs produce the same output. The count is a display resolution for the model, not a sample size.
Low score means
A low model share means this variant is weaker than the other options in this run under the tool's fixed rules. It says nothing about how people on YouTube will respond.
How to raise it
You cannot tune the model share directly. Improve the underlying cover or title signals, then re-run the same set to compare the revision.

Relative preflight ranking

The set is ordered by its combined preflight scores. The top candidate is the strongest under this tool's rules, not a guaranteed winner against other thumbnails on YouTube.

See scoring details
What it measures
A best-to-worst sort of the variants in this run, computed entirely from local pixel and title signals. It has no reference to another channel, published performance, or audience behavior.
Low score means
Ranking last means only that the option trails this particular set. Add or remove a concept and the relative context changes.
How to raise it
Revise the lowest-ranked option, then compare it against the current leader again. Use the ranking to build a shortlist, not as proof of performance.

Prioritized fix suggestions

The tool flags the weakest dimension for each variant and suggests a focused change. Start with the biggest scoring gap instead of polishing a signal that already scores well.

See scoring details
What it measures
The weakest dimension on each variant, listed in priority order. The list is built from the same five cover and seven title scores, so the named dimension is the one currently costing the variant the most points.
Low score means
Nothing. The fix list is not a score; it is a work order. The dimension it names is the one with the most headroom, not a verdict on the cover.
How to raise it
Apply the named change, re-upload, and re-score. The new result shows whether that revision improved the fixed signals.

Local and private by default

The scoring flow runs in your browser. Your original images stay local during the test. If you log in and save, account sync sends only a downsized report preview and score data.

See scoring details
What it measures
Nothing. Privacy is a property of how the tool runs, not a score it produces. The cover read, scoring, and deterministic preference model all run inside your browser tab.
Low score means
Nothing to lower. There is no privacy score. Your original images are read and scored locally whatever the result.
How to raise it
Nothing to raise. No YouTube account, no extension, no OAuth, and no original image upload for scoring.

Pre-publish workflow

How a test runs

  1. Upload 2 to 6 pairs

    Drop in between two and six thumbnails, each with its title. Every pair is treated as one unit, because that is how viewers see them in the feed.

  2. Get an explainable shortlist

    Each pair gets a preflight scorecard across five cover dimensions and seven title angles. The set is ranked, and the top three appear in two reduced layouts before you choose what deserves another revision or a real YouTube test.

  3. Fix the weakest dimension

    Open the lowest-ranked variant, check which dimension dragged it down, and apply the suggested change before you publish, instead of guessing after the fact.

What the tool does not do

The scorer sees only the files and title text you provide. It cannot see channel history or published performance.

  • Does not test on people. The 10,000 iterations allocate model share from the scores already calculated. No person, panel, or YouTube audience is sampled.
  • Does not predict your YouTube CTR. The preflight score and model share are comparison aids. Neither is a forecast of the click-through rate your published video will earn.
  • Cannot see your real audience, channel, or topic. Your subscriber base, watch history, niche, and the specific video topic never reach the tool. They are the reasons no absolute CTR can be promised.
  • Does not touch your live video. It is a pre-publish scoring tool with no connection to your YouTube channel, so running it cannot affect a live video's statistics.

Feature FAQ

The questions that come up most often about what each capability actually measures.

What exactly does the cover score measure?

Five dimensions from your pixels: salience (30%), emotional pull (25%), color quality (20%), composition (15%), and readability (10%). Each cover also gets a small plus or minus 5 fingerprint so near-identical covers do not collide on the same total. The full method is on the methodology page.

Does the tool analyze the rule of thirds or spatial layout?

No. Composition here measures only the share of the frame taken up by skin tones, as a proxy for a prominent subject. It does not detect where the subject sits, apply the rule of thirds, or read text. If your layout is strong but no skin tone is present, composition still scores low.

Is the model share a predicted CTR?

No. It is a deterministic allocation across the variants in one run. A 43% model share means the fixed model assigned 43% of its iterations to that option, not that 43% of people would click.

How many variants can I compare at once?

Between 2 and 6. The tool needs at least 2 variants to rank anything, and supports up to 6 in a single test.

Do my images leave my device?

Original images stay local for scoring. Account sync sends a downsized preview, title text, scores, and result metadata only after sign-in and an explicit save.

Can the score tell me if my video will perform?

No. A high score means the concept fits the fixed salience, color, focus, readability, and title rules. It cannot predict topic quality, audience fit, clicks, or retention.

Can I check how the finalists read at smaller sizes?

Yes. After a comparison, the top three appear in a reduced feed card and compact list. This checks size and readability; it is not a pixel-perfect YouTube interface or a viewer test.

See the scorecard on your own covers

Drop in 2 to 6 thumbnails and titles to compare 5 visual signals and 7 title signals, then see the relative ranking and first fix for each. Free, no login required for the core test.

Test your thumbnails