Research

Recommender benchmark

27 recommender systems on 10 public datasets, given the same data and judged on the slate they produce. Quality, cost and latency measured separately. Protocol, confidence intervals and negative results are published with every release.

Live leaderboard · Recommender Benchmark v3.2

Best recommender systems

27 complete recommender systems - classical, neural, graph, sequential, multi-stage and generative - given the same data and judged on the ranked slate they produce, across 10 public datasets.

  1. 1ItemKNN + Multi-task ranker68.8
  2. 2ItemKNN + MLP ranker66.9
  3. 3VS-KNN65.2
  4. 4EASEᴿ64.5
  5. 5SASRec62.9
  6. 6VS-KNN + GBDT60.8
  7. 7Recency-Weighted Popularity60.2
  8. 8LightGCN59.8
  9. 9ItemKNN58.1
  10. 10BERT4Rec57.2
genGenerativewhere ItemKNN landedFull leaderboard, charts and caveats →
Multi-stageNeighbourhoodLinear autoencoderSequentialNon-personalisedGraphGenerative · semantic IDMulti-stage (neural)Matrix factorisationOption scoringGenerative · sequentialFoundation modelUnified transformer

Scale. A score is the percentile of that system’s measurement within a frozen distribution of every entrant’s measurement on the same metric and dataset. No system is fixed at an endpoint, and adding one does not move an existing score. ItemKNN places 9th (58.1), popularity 25th (38.7) of 27. Index v3.1 replaced a two-point scale that fixed these at 100 and 0. Method.

Recommender Performance Index. Six weighted dimensions: relevance, cold start, long tail, novelty and diversity, multi-objective value, robustness. Quality only; cost and latency are reported separately. Colour encodes architecture family only. Adjacent ranks are often statistical ties; the leaderboard marks them. Generated 19 September 2026.

The reports behind it

Each one is the full method and the complete results, including what did not replicate.

Related

The benchmark’s methodology, metric definitions, changelog and downloadable results are published in full and licensed CC BY 4.0. For the category these results sit in, see recommendation engines and the recommendation engine API; for the products in it, the platform comparisons.