Research
Recommender benchmark
27 recommender systems on 10 public datasets, given the same data and judged on the slate they produce. Quality, cost and latency measured separately. Protocol, confidence intervals and negative results are published with every release.
Live leaderboard · Recommender Benchmark v3.2
Best recommender systems
27 complete recommender systems - classical, neural, graph, sequential, multi-stage and generative - given the same data and judged on the ranked slate they produce, across 10 public datasets.
- 1ItemKNN + Multi-task ranker68.8
- 2ItemKNN + MLP ranker66.9
- 3VS-KNN65.2
- 4EASEᴿ64.5
- 5SASRec62.9
- 6VS-KNN + GBDT60.8
- 7Recency-Weighted Popularity60.2
- 8LightGCN59.8
- 9ItemKNN58.1
- 10BERT4Rec57.2
Scale. A score is the percentile of that system’s measurement within a frozen distribution of every entrant’s measurement on the same metric and dataset. No system is fixed at an endpoint, and adding one does not move an existing score. ItemKNN places 9th (58.1), popularity 25th (38.7) of 27. Index v3.1 replaced a two-point scale that fixed these at 100 and 0. Method.
Recommender Performance Index. Six weighted dimensions: relevance, cold start, long tail, novelty and diversity, multi-objective value, robustness. Quality only; cost and latency are reported separately. Colour encodes architecture family only. Adjacent ranks are often statistical ties; the leaderboard marks them. Generated 19 September 2026.
The reports behind it
Each one is the full method and the complete results, including what did not replicate.
Which Complete Recommender System Produces the Best Recommendations?
Twenty-three complete recommender systems - classical, neural, graph, sequential, multi-stage and generative - given identical data and judged on the ranked slate they produce, across ten public datasets.
23
Complete systems compared
frozen field
How it is scaled
kept apart
Quality and cost
Which Recommender Architecture Wins, at Which Stage, Under Which Data Regime
Seventeen complete recommendation pipelines measured stage by stage across nine datasets, including feeds whose logs record impressions rather than only engagements.
12 of 12
Top systems that are hybrids
−0.30
Offline vs. true ranking
−11.7%
Cost of diversity reranking
Comparing Recommender Systems Under a Leakage-Free Evaluation Protocol
Nine recommendation approaches measured on a chronological split, ranked against the full catalogue, with confidence intervals and negative results.
39.9%
Split protocol
+44%
Reported score inflation
3.4×
Retrieval objective
Related
The benchmark’s methodology, metric definitions, changelog and downloadable results are published in full and licensed CC BY 4.0. For the category these results sit in, see recommendation engines and the recommendation engine API; for the products in it, the platform comparisons.