Method

Everything the engine does, in the open

A verdict you cannot check is just an opinion. This page documents every rule the engine applies. The same inputs always produce the same output, and no language model is involved in reaching the conclusion — text extraction, where used, only pulls a published number across and always shows it next to its source.

1. Global evidence collection

Every factual finding is a published figure attributed to the publisher that produced it: expert reviews, laboratory measurements, benchmark database entries, professional evaluations and manufacturer commitments. The publisher registry currently holds 35 independent publishers across 15 countries and is designed to scale to hundreds, with a per-publisher collector that discovers documents and extracts figures — never estimates them. The corpus behind the live smartphone category is 576 documents and 2520 findings across 24 models.

Review scores are never averaged into a rating. For each metric the engine takes the median of the independent published figures and reports the spread between them: agreement is strong under 5%, moderate under 15%, and anything wider is flagged as contested so you can see that the testers do not agree.

Tier 1

Lab measurement or public benchmark database

Tier 2

Independent editorial testing with a documented procedure

Tier 3

Manufacturer specification, used only for stated commitments

Live collection is being rolled out publisher by publisher. Until a publisher is live, its findings are marked "seeded corpus" in the analysis so a demonstration figure is never presented as a verified one.

2. Metrics and base weights

MetricGroupBetterWeightThreshold
Active use runtimeBattery & charginghigher1.008%
0–100% charge timeBattery & charginglower0.5512%
CPU multi-corePerformancehigher0.8512%
Sustained GPUPerformancehigher1.0012%
Stills qualityCamerahigher1.004%
Zoom qualityCamerahigher0.706%
Video qualityCamerahigher0.804%
Peak brightnessDisplayhigher1.0015%
Refresh rateDisplayhigher0.6010%
Remaining supportLongevityhigher1.0012%
WeightHandlinglower1.006%

A difference smaller than the threshold is reported as "no real difference" and contributes exactly zero points, however good it looks on a spec sheet.

3. Your priorities become weights

×0.0Ignore
×0.5Low
×1.0Normal
×1.6High
×2.4Critical

Final weight = base metric weight × priority multiplier. All weights are then normalised so the score always lands on the same −100…+100 scale.

4. Your phone is scored as it is today

  • Battery: measured runtime × your battery health percentage.
  • Performance: benchmark figures reduced by a condition allowance — flawless 0%, good 2%, worn 5%, damaged 12% — covering thermal and throttling behaviour of a used device.
  • Longevity: only the support years still remaining count, for both devices.
  • Resale: launch price depreciated 24% per year, then adjusted for condition, battery health and storage.

5. The Upgrade Value Score and the verdict

Each metric's relative change is capped at ±60% — past that point more is not meaningfully better — then multiplied by its normalised weight and scaled to points. The Upgrade Value Score is the sum.

The break-even score is net cost ÷ 14, so a more expensive upgrade has to prove more. Verdict bands:

  • score ≤ 8, or below 70% of break-even → Keep your current phone
  • below break-even → Marginal upgrade
  • at or above break-even → Worth it
  • ≥ 1.6× break-even and ≥ 35 points → Clear upgrade

6. Limits we state openly

The engine only judges what is measurable and published. Preferences like design, ecosystem lock-in or brand loyalty are not scored. Where a model lacks a published figure for a metric, that metric is marked unknown and contributes nothing rather than being guessed.