What the score is
The Cleats Hub Score is a single number between 0 and 100 that summarises the weight of independent evidence about a product. It is not an editor's opinion, not an average of star ratings, and not influenced by advertising or affiliate revenue.
Think of it as a meta-score, the same idea behind Metacritic for films or Rotten Tomatoes for critics' verdicts. We read every published review we can find, extract what each one actually says, and blend those verdicts into one transparent, reproducible number.
We don't test products ourselves, and we won't pretend we do. That honesty is a feature, not a limitation. It means the score is never coloured by a single tester's preferences, and it scales across every product we cover.
Where the data comes from
Every score is built from three independent pools of evidence:
- Published critic reviews, specialist sports-equipment publications, general sports media, and YouTube channels whose reviewers skate, lace up, or play in the gear they cover. We record the outlet, the author, the date, any numeric rating given, and the scale it was on.
- Owner feedback, retailer review aggregates (with source counts), forum discussions, Reddit threads, and Cleats Hub user submissions. Owner data captures long-term, real-world use that critics who review pre-release samples often can't.
- Manufacturer specs, objective, verifiable figures: weight, listed price, materials, stiffness rating, boot height. These are facts, not sentiment, and they anchor the score against things that can be independently confirmed.
Every source we use is listed on the product's page, with an outbound link so you can read it yourself.
How sources are weighted
Not all reviews carry equal weight. A decade-old YouTube video from a casual player counts less than a 2025 deep-dive from a specialist hockey skate reviewer. We apply three multipliers to each source:
- Authority tier. Tier 1: dedicated specialist reviewers with a track record for this category. Tier 2: generalist sports media, mainstream YouTube. Tier 3: retailer-hosted and anonymous owner reviews. Higher tiers receive a larger weight multiplier.
- Recency, scores decay with age, resetting when a new model generation launches. A review from this season outweighs one from three years ago, which may cover a different construction.
- Depth, a 3,000-word long-term review after six months of use carries more signal than a first-impression unboxing. For owner aggregates, the sample size is factored in (logarithmically, a pool of 400 reviews is meaningful, but not 20x more meaningful than 20).
Layer weights
The three evidence layers are blended at the following default weights:
These weights adjust automatically when a layer has thin coverage. A product with only two critic reviews leans more heavily on owners and specs until the critic pool grows.
How quotes are extracted and verified
We use a structured extraction step to turn long-form reviews into comparable data points. For each source and each characteristic (fit, traction, durability, etc.), we record:
- A numeric stance on the 0–10 scale
- A verbatim quote that supports that stance
- An extraction confidence score (0–1)
The verbatim quote is the critical safeguard. Every claim on a product page traces back to a real sentence from a real published source. You can expand the evidence drawer on any characteristic to see the exact quote, the outlet it came from, and its date. If the quote doesn't support the stance, the data point is wrong, and we want to know about it.
The scoring arithmetic itself is pure code: deterministic, auditable, never hallucinated. The extraction step is the only place where natural language processing is involved, and it only reads and quotes, it never invents a number.
Confidence levels explained
Every score carries a confidence badge, Low, Medium, or High, which tells you how much to trust the precision of that number.
- Low confidence, fewer than the target number of sources, or sources that are old or in sharp disagreement. The score is a best estimate from sparse data. Treat it as a directional indicator, not a fine-grained ranking.
- Medium confidence, reasonable coverage across at least two source types, with broadly agreeing signals. The score is reliable for most buying decisions.
- High confidence, strong consensus across multiple authority sources, recent data, and low disagreement. The score is as reliable as this methodology can produce.
We never suppress a low-confidence score, but we do flag it clearly so you can weigh it accordingly. Hiding uncertainty would make us feel authoritative and make you worse-informed.
What we don't do
- No hands-on testing. We don't skate in the skates or run routes in the cleats. Our credibility comes from the breadth and transparency of what we aggregate, not from pretending otherwise.
- No paid placements. A brand cannot pay to appear in a list, to have its score boosted, or to have a competitor's score reduced. There is no such product and there never will be.
- No score-for-cash. Affiliate commissions (see our Affiliate Disclosure) are earned when a reader clicks through and buys. They are the same regardless of where a product ranks. The score is never a factor in which links we include or which retailers we surface.
How affiliate links relate to scores
Product pages include links to retailers. If you click through and buy, we may earn a commission at no extra cost to you. That commission is identical whether the product scores 91 or 54, we have no financial incentive to inflate or deflate any score.
The retailer links are selected to give you the best available prices, not to favour retailers that pay higher rates. Full details are on our Affiliate Disclosure page.
When scores are updated
Scores are re-run whenever a new review is added to the evidence pool or a new model-year variant launches. Every product page shows a "last updated" date so you know how current the score is. If you spot a review we've missed, use the feedback link on any product page to tell us.