NSL Research · Recommender Benchmark v3.2 · 19 September 2026
Best Recommender Systems
27 complete recommender systems - classical, neural, graph, sequential, multi-stage and generative - given the same data and the same catalogue, and judged on the only thing a user ever sees: the ranked slate they produce. Quality, cost and latency are measured separately, because the trade-off between them is yours to make.
NSL Recommender Performance Index
Relevance, cold start, long tail, novelty and diversity, multi-objective value and robustness, weighted into one score. Quality only.
Filter
| System | |||||
|---|---|---|---|---|---|
| 1 | ItemKNN + Multi-task ranker Multi-stage | 68.8 | $$$$ | ▮▮▮▮▮ | |
| 2 | ItemKNN + MLP ranker Multi-stage | 66.9 | $$$$ | ▮▮▮▮▮ | |
| 3 | VS-KNN Neighbourhood | 65.2 | $$$ | ▮▮▮▮ | |
| 4 | EASEᴿ Linear autoencoder | 64.5 | $$ | ▮▮ | |
| 5 | SASRec Sequential | 62.9 | $ | ▮ | |
| 6 | VS-KNN + GBDT Multi-stage | 60.8 | $$$$$ | ▮▮▮▮▮ | |
| 7 | Recency-Weighted Popularity Non-personalised | 60.2 | $ | ▮ | |
| 8 | LightGCN Graph | 59.8 | $ | ▮ | |
| 9 | ItemKNN Neighbourhood | 58.1 | $$ | ▮▮ | |
| 10 | BERT4Rec Sequential | 57.2 | $ | ▮ | |
| 11 | PLUM-style SID generative retrieval Generative · semantic ID | 55.0 | $$$$$ | ▮▮▮▮▮ | |
| 12 | Two-Tower + DCN-v2 Multi-stage (neural) | 53.6 | $$$$ | ▮▮▮▮ | |
| 13 | GRU4Rec Sequential | 52.2 | $ | ▮ | |
| 14 | Cheap full pass + Option-Attention rerankerNew Multi-stage (neural) | 52.0 | $$ | ▮ | |
| 15 | TIGER Generative · semantic ID | 51.6 | $$$ | ▮▮▮▮ | |
| 16 | Two-Tower + RankMixer-style ranker Multi-stage (neural) | 51.3 | $$$$ | ▮▮▮▮ | |
| 17 | iALS Matrix factorisation | 51.0 | $$ | ▮ | |
| 18 | EASEᴿ + LambdaMART Multi-stage | 50.9 | $$$$ | ▮▮▮▮▮ | |
| 19 | Option-Attention Scorer (full catalogue)New Option scoring | 50.3 | $$ | ▮▮ | |
| 20 | HSTU (Generative Recommenders) Generative · sequential | 49.6 | $$ | ▮ | |
| 21 | Calibrated Multi-Outcome Option ScorerNew Option scoring | 47.5 | $$ | ▮▮ | |
| 22 | Netflix-style foundation model Foundation model | 46.0 | $$ | ▮ | |
| 23 | OneRec-style session generator Generative · sequential | 42.4 | $$$$$ | ▮▮▮▮▮ | |
| 24 | ItemKNN + Option-Attention rerankerNew Multi-stage (neural) | 40.2 | $$ | ▮▮ | |
| 25 | Popularity Non-personalised | 38.7 | $ | ▮ | |
| 26 | Content TF-IDF Neighbourhood | 37.7 | $$ | ▮▮ | |
| 27 | Phoenix-style unified transformer Unified transformer | 27.1 | $$ | ▮▮ |
Showing 27 of 27 systems, scored on the best-achievable compute budget; open a row to see the same system on the compute-constrained budget. Click any row for what the system is, where it comes from, and what its score is made of. The Recommender Performance Index is quality only - cost and latency sit beside it, never inside it.
Where each architecture wins
Scenario winners
The same results sliced by the question a team usually has: what happens to a user with four interactions, to an item nobody has touched, to the long tail, to a feed with eight competing engagement signals.
Cold users
cohort:cold_users_1_4:ndcg@10 · 7 datasetsUsers with 1–4 prior interactions
- 1Recency-Weighted Popularity0.0983
- 2VS-KNN0.0864
- 3ItemKNN + Multi-task ranker0.0847
- 4EASEᴿ0.0838
- 5LightGCN0.0814
- 6ItemKNN + MLP ranker0.0800
- 7ItemKNN0.0796
- 8VS-KNN + GBDT0.0793
Cold items
cohort:cold_items:ndcg@10 · 4 datasetsTarget items with no fit-window interactions
- 1Content TF-IDF0.0166
- 2Two-Tower + RankMixer-style ranker0.0055
- 3Two-Tower + DCN-v20.0039
- =Popularity0.0000
- =ItemKNN0.0000
- =Recency-Weighted Popularity0.0000
- =EASEᴿ0.0000
- =iALS0.0000
24 of 27 score exactly 0.0000 here. They are marked = rather than numbered: their order in this list is the sort’s, not the measurement’s.
Sparse histories
cohort:sparse_history:ndcg@10 · 10 datasetsBottom-quartile history length
- 1Recency-Weighted Popularity0.0947
- 2EASEᴿ0.0911
- 3VS-KNN0.0873
- 4ItemKNN + Multi-task ranker0.0855
- 5ItemKNN0.0844
- 6LightGCN0.0839
- 7ItemKNN + MLP ranker0.0834
- 8VS-KNN + GBDT0.0719
Dense histories
cohort:dense_history:ndcg@10 · 9 datasetsTop-quartile history length
- 1Recency-Weighted Popularity0.0756
- 2ItemKNN + Multi-task ranker0.0591
- 3BERT4Rec0.0587
- 4VS-KNN0.0585
- 5ItemKNN + MLP ranker0.0576
- 6EASEᴿ0.0566
- 7EASEᴿ + LambdaMART0.0566
- 8VS-KNN + GBDT0.0550
Long tail
cohort:tail_target_items:ndcg@10 · 8 datasetsTargets below median popularity
- 1Content TF-IDF0.0103
- 2VS-KNN + GBDT0.0039
- 3ItemKNN + MLP ranker0.0036
- 4ItemKNN + Multi-task ranker0.0035
- 5EASEᴿ + LambdaMART0.0033
- 6EASEᴿ0.0027
- 7ItemKNN0.0027
- 8Two-Tower + DCN-v20.0025
Catalogue coverage
coverage@20 · 10 datasetsShare of the catalogue ever served
- 1GRU4Rec0.5115
- 2iALS0.4409
- 3LightGCN0.3839
- 4HSTU (Generative Recommenders)0.3451
- 5Content TF-IDF0.3400
- 6EASEᴿ0.3280
- 7Netflix-style foundation model0.3244
- 8PLUM-style SID generative retrieval0.3157
Novelty
novelty@20 · 10 datasetsMean self-information of served items
- 1Content TF-IDF14.1858
- 2GRU4Rec11.4019
- 3SASRec11.1379
- 4iALS11.0139
- 5Netflix-style foundation model10.9912
- 6HSTU (Generative Recommenders)10.8948
- 7Phoenix-style unified transformer10.8884
- 8PLUM-style SID generative retrieval10.8528
Diversity
ild@20 · 10 datasetsIntra-list tag diversity
- 1ItemKNN + Option-Attention reranker0.7941
- 2Recency-Weighted Popularity0.7910
- 3iALS0.7760
- 4Popularity0.7728
- 5VS-KNN0.7695
- 6ItemKNN0.7671
- 7HSTU (Generative Recommenders)0.7651
- 8LightGCN0.7605
Multi-objective value
mo_weighted_utility · 3 datasetsWeighted value across all logged signals
- 1SASRec0.0582
- 2BERT4Rec0.0560
- 3Two-Tower + RankMixer-style ranker0.0497
- 4Content TF-IDF0.0472
- 5ItemKNN + Multi-task ranker0.0448
- 6HSTU (Generative Recommenders)0.0443
- 7ItemKNN + MLP ranker0.0438
- 8GRU4Rec0.0435
Winners, in words
Best at…
Best overall
ItemKNN + Multi-task ranker
Recommender Performance Index
Runner-up: ItemKNN + MLP ranker
Best generative recommender
PLUM-style SID generative retrieval
Recommender Performance Index, generative entrants only
Runner-up: TIGER
Best cold start
ItemKNN + Multi-task ranker
cold-start dimension
Runner-up: Recency-Weighted Popularity
Best long tail
Content TF-IDF
long-tail dimension
Runner-up: GRU4Rec
Best multi-objective
SASRec
multi-objective dimension
Runner-up: VS-KNN + GBDT
Best under a latency budget
ItemKNN + Multi-task ranker
Recommender Performance Index among systems serving under 25 ms/user offline
Runner-up: ItemKNN + MLP ranker
Best quality per dollar
Recency-Weighted Popularity
Recommender Value Index
Runner-up: GRU4Rec
Best retriever
EASEᴿ
candidate recall@100, systems with a separable retrieval stage
Runner-up: EASEᴿ + LambdaMART
Best ranker
Cheap full pass + Option-Attention reranker
NDCG@10 on identical candidate sets
Runner-up: ItemKNN + Multi-task ranker
Best novelty & diversity
GRU4Rec
novelty, diversity and coverage
Runner-up: iALS
Build your own
Pick the stages, or take a model that does both
Not run end to end
The benchmark measures whole systems, and it never ran this pair as one. There is no quality index, cost or latency for the combination, and this page will not invent one — what it knows about each stage separately is below.
Retrieval stage · does the answer survive?
of held-out items reach this stage’s top 100. Whatever it drops, no ranker can recover.
Ranking stage · how well does it order?
NDCG@10 over one shared candidate set, identical for every ranking stage, so this compares rankers to each other rather than pipelines to each other.
What the results say
Key findings
Every sentence below is derived from the stored results rather than written by hand, so a rerun of the benchmark rewrites them.
ItemKNN + Multi-task ranker leads
It scores 68.8 on the Recommender Performance Index - 1.9 points ahead of ItemKNN + MLP ranker - on the best-achievable compute track across 10 datasets. A score is a percentile in a frozen distribution of every entrant's measurements, so it is read against the field rather than against a chosen system: the twenty-five-year-old ItemKNN baseline scores 58.1 and places 9th of 27, which is the more useful thing to calibrate against.
No system wins everywhere
Across 9 scenario leaderboards - cold users, cold items, sparse and dense histories, long tail, coverage, novelty, diversity and multi-objective value - 5 different systems take first place. The right question is not which recommender is best; it is which is best for the part of the distribution your product actually lives in.
Generative architectures do not yet lead at this scale
5 entrants decode item identifiers rather than scoring the catalogue. Their best index is 55.0 against 68.8 for the rest, and they take the top spot in 0 of 9 scenarios. The industrial reports behind these architectures all rest on a scaling claim that a public-data benchmark of this size cannot test.
The last few points of quality are the expensive ones
SASRec comes within 5.9 index points of the winner for 35.2× less - a difference most products would struggle to detect.
Offline evaluation and randomised exposure pick different winners
One dataset in the panel inserts randomly-chosen items into real feeds. Exposure on those rows is independent of the incumbent recommender, so a plain average over them is an unbiased estimate of a slate's value - no propensity model, no assumption. Across 23 systems the rank correlation between the ordinary offline ordering and that unbiased one is 0.17. Offline picks SASRec; randomised exposure picks LightGCN. Offline evaluation here does not merely inflate scores - it can choose a different system. Read every number on this page with that in mind.
Quality against cost
What the top of the leaderboard actually costs
Quality is on the vertical axis and never mixed into the horizontal one. The dashed line is the Pareto frontier: nothing above and to the left of a point on it is both better and cheaper. Most deployment decisions are made against this shape rather than against the ranking.
One confound worth knowing. The multi-stage systems build their ranker’s features online, in Python, for every request here - work a deployed pipeline precomputes or caches. Their position on the horizontal axis is an upper bound, and the cost gap between them and the single-stage systems is overstated by an amount this benchmark does not try to estimate. Quality is measured on the served slate and carries no such caveat, which is why the index excludes cost entirely.
The comparison everyone asks for
Generative against traditional
Semantic-ID decoders, autoregressive user models and session generators against neighbourhood models, matrix factorisation, graphs and multi-stage pipelines - on identical data, with identical budgets and identical evaluation.
5
Generative entrants
22
Traditional entrants
55.0
Best generative index
68.8
Best traditional index
Scenario by scenario, each normalised to its winner
Three of the generative entrants are architecture proxies for proprietary systems and one is trained from scratch where the original is adapted from a pre-trained language model; the comparison is between architectures at benchmark scale, not between the companies' production systems.
Plain English
What the leading systems actually are
A short description of the top of the table, including what our implementation is and - where the system is proprietary - what it is not.
ItemKNN + Multi-task ranker
ReproductionA neighbourhood model produces the candidates, and the ranking stage carries one output head per logged engagement signal — click, long view, like, share, follow, and the negative signals — combined into a single order by weights fixed before any result was measured (Ma et al., KDD 2018). On a dataset that logs only one signal there is only one head, and it is reported as a pointwise ranker there.
Multi-task ranking over several engagement signals · Ma et al., 'Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts', KDD 2018; Zhao et al., 'Recommending What Video to Watch Next', RecSys 2019.
Index 68.8 · high cost · slowest · GPU
ItemKNN + MLP ranker
ReferenceA neighbourhood model produces a few hundred candidates and a multilayer perceptron reorders them, over features combining the retrieval score, item and user statistics, and content similarity. The retriever is fitted on the training window, the ranker is trained on candidates labelled from the validation window, and the retriever is refitted for serving; every part of that cost is charged to the system.
Classical retrieval with a neural pointwise ranker · Sarwar et al., WWW 2001 (retrieval); Covington, Adams & Sargin, RecSys 2016 (the deep-ranking stage this follows).
Index 66.9 · high cost · slowest · GPU
VS-KNN
ReferenceSession-based nearest neighbours: it finds past sessions that resemble the one in progress, weighting the current session by position, and scores what those sessions went on to contain. It holds no long-term profile, so it produces recommendations for traffic with no identity attached to it.
Vector-multiplication session-based kNN · Ludewig & Jannach, 'Evaluation of Session-based Recommendation Algorithms', UMUAI 2018.
Index 65.2 · moderate cost · slow · CPU only
EASEᴿ
ReferenceA linear item-item autoencoder with a closed-form solution: one regularised Gram-matrix inverse, no iterations, and no hyperparameter beyond the ridge term (Steck, WWW 2019). Its memory and time are quadratic in catalogue size, so it is not attempted above roughly 25,000 items; the datasets where that applies are recorded as skipped rather than left out.
Embarrassingly Shallow AutoEncoder · Steck, 'Embarrassingly Shallow Autoencoders for Sparse Data', WWW 2019.
Index 64.5 · low cost · fast · CPU only
SASRec
ReproductionA causal transformer over the ordered item sequence (Kang & McAuley, ICDM 2018), trained with a softmax over the full catalogue rather than the original binary objective — the formulation Klenitskiy & Vasilev (RecSys 2023) found accounts for most of the reported gap to bidirectional models.
Self-attentive sequential recommendation · Kang & McAuley, 'Self-Attentive Sequential Recommendation', ICDM 2018; trained with the full softmax of Klenitskiy & Vasilev, RecSys 2023.
Index 62.9 · lowest cost · fastest · GPU
VS-KNN + GBDT
ReferenceSession-based nearest neighbours produce the candidates and a pointwise gradient-boosted ranker orders them. The sequence signal enters through the retrieval stage rather than through a neural sequence model, which separates the contribution of sequence information from the contribution of the architecture that usually carries it.
Session-based neighbourhood retrieval with a boosted ranker · Ludewig & Jannach, UMUAI 2018; Ke et al., NeurIPS 2017 (LightGBM).
Index 60.8 · highest cost · slowest · CPU only
Recency-Weighted Popularity
ReferenceCatalogue-wide interaction count with an exponential recency kernel applied across the training window, so recently popular items outrank items that were popular earlier. Like plain popularity it returns the same order for every user, and it has nothing to train beyond the decay constant.
Standard non-personalised baseline with a recency kernel · Ji et al., 'A Re-visit of the Popularity Baseline', SIGIR 2020.
Index 60.2 · lowest cost · fastest · CPU only
LightGCN
ReproductionRounds of neighbour averaging over the user-item bipartite graph, with the nonlinearities and feature transforms of general graph convolution removed (He et al., SIGIR 2020). Propagation reaches items two or more hops away, which no direct co-occurrence connects to the user; on a sparse graph those hops tend to arrive at broadly popular items.
Simplified graph convolution for collaborative filtering · He et al., 'LightGCN: Simplifying and Powering Graph Convolution Network for Recommendation', SIGIR 2020.
Index 59.8 · lowest cost · fastest · GPU
Read this before quoting anything above
What this benchmark cannot tell you
These are not disclaimers. They are the boundary of what the measurements support, and a result quoted outside that boundary is a wrong result.
- 01
On the one dataset that inserts randomly-chosen items into real feeds — the only place an unbiased estimate is available — the offline ordering and the unbiased ordering correlate at 0.17 and name different winners. That is a limit on what any offline recommender benchmark, including this one, can tell you.
- 02
Public offline datasets are not production environments. Offline relevance is a proxy for user value, and the mapping between them is platform-specific.
- 03
Several entrants are architecture proxies for proprietary systems. They implement a published mechanism at benchmark scale and are not those companies' production systems.
- 04
No single recommender wins every scenario. The winner changes 5 ways across the scenario boards.
- 05
Entrants added after the calibrating release are measured in their own run rather than beside the systems they are ranked against. Every board is unaffected — a score is a percentile in a frozen distribution, not a rank against whoever ran that day — but a head-to-head significance cell exists only where both systems' result files carry per-user vectors. Which cells those are is published in `pairwise_provenance`. Cost and latency use the same instance types across runs, not the same job.
- 06
Serving latency is offline batch throughput on the benchmark host, not a production p99. Multi-stage systems additionally build their ranker's features online here, which a deployed pipeline precomputes — so their measured serving cost is an upper bound and the cost gap to single-stage systems is overstated.
- 07
Scale is the untested variable: the central claim of the industrial generative systems is a scaling claim this benchmark is too small to test.
- 08
The index scores each measurement by its percentile in a frozen, published distribution. That records where a system placed, not by how much: doubling the best cold-start result and beating it by a hair score the same. Raw metric values are published beside every score, and the scenario boards are stated in raw metrics.
- 09
This is index version v3.1 and its scores are not comparable to v3.0's. The v3.0 scale rescaled every metric so Popularity read 0 and ItemKNN 100; it was withdrawn before publication because on two of ten datasets Popularity beat ItemKNN, which inverted the scale, and where the two were close a worse result could reach -381. The effect was to flatter ItemKNN by about six places. Under v3.1 nothing is pinned and ItemKNN places 9th.
Methodology
How the index is built
What the labels mean
- Official
- Official implementation — Authors' released implementation, run as published.
- Reference
- Reference implementation — Faithful implementation of a fully specified public method.
- Reproduction
- Reproduction — Our implementation from the paper, using the paper's stated design.
- Proxy
- Architecture proxy — The target system is proprietary and only partially described; we implement the published mechanism at benchmark scale. Not the production system.
What the index measures
Recommendation quality and nothing else. Cost, latency, model size and retraining burden are reported beside it and combined only in the separate Recommender Value Index.
- Relevance35%
- Cold start15%
- Long tail15%
- Novelty & diversity10%
- Multi-objective10%
- Robustness15%
Weights frozen 2026-08-22, fingerprint 52706396658de50c, so you can check they were not adjusted after the results existed.
How it is scaled
A system’s score on a metric is its percentile in a frozen field: the distribution of what every entrant measured on that same metric and dataset, pooled across both compute tracks and then frozen. Nothing is scaled against a chosen system, so no system sits at a fixed position - the two reference systems included.
Freezing the distribution is what makes adding a new system to this leaderboard leave every existing score exactly where it was. A percentile taken against the current field would move all of them, which makes historical comparison impossible.
The field is 167 published constants, fingerprinted 05eb9f3f645e581b, so a score can be recomputed - or your own system scored against this benchmark - without rerunning anything. Open any row for the unnormalised measurements behind its score.
Re-scoring the identical results against the live field, which depends on no published constant at all, gives a rank correlation of 0.98 with this one, so the ordering is a property of the measurements rather than of the frozen constants.
Index version v3.1 replaced a scale that was withdrawn before publication. v3.0 used a two-point anchored scale dividing by (itemknn - popularity). It was withdrawn: on two of ten datasets that difference was negative, inverting the scale, and where it was small a worse result could reach -381 before being floored. Scores are not comparable across the change. Neither failure could touch the reference system itself, which sat on 100 by arithmetic - the effect was to flatter it by about six places against three independent checks that use no normalisation at all.
What the cost figures assume
A stated reference deployment: 1,000,000 active users, 10 slates per user per month (10,000,000 slates), retrained 4.35 times a month, at on-demand cloud rates. The per-slate and per-retrain figures it multiplies are measured; the multipliers are assumptions, published so you can redo the arithmetic for your own traffic.
Datasets
All public, used under their stated licences, none redistributed. Chronological splits, full-catalogue metrics with no sampled negatives, user-level bootstrap intervals, paired permutation tests with multiplicity correction.
2 systems could not be attempted on 4 datasets and are recorded as skipped rather than dropped - for example: catalogue has 39979 items; this system is not attempted above 25000 by policy
NSL Recommender Benchmark v3.2 · generated 19 September 2026 · best-achievable compute track · scoring weights frozen 2026-08-22, fingerprint 52706396658de50c. The leaderboard is meant to grow: a new architecture is a class in the harness and a rerun, and because every score is a percentile in a frozen, published distribution rather than a rank against the current field, adding one leaves every score already published exactly where it was. If a system you care about is missing, tell us. Every definition behind these numbers - metrics, cohorts, cold-start cuts, reproduction, model sources and the release changelog - is on the full methodology page.
Use these results
Download, reproduce and cite
The results are free, ungated and licensed CC BY 4.0. The versioned path never changes once published, so a number quoted from it stays checkable.
Machine-readable results
- Leaderboard (CSV)/research/benchmark/v3.2/leaderboard.csv
- Leaderboard (JSON)/research/benchmark/v3.2/leaderboard.json
- Per-dataset measurements (CSV)/research/benchmark/v3.2/per-dataset-metrics.csv
- Scenario boards (CSV)/research/benchmark/v3.2/scenarios.csv
- Datasets (CSV)/research/benchmark/v3.2/datasets.csv
- Full payload (JSON)/research/benchmark/v3.2/benchmark-full.json
Full metric and cohort definitions, the reproduction path, the model sources and the changelog are on the methodology page.
How to cite this benchmark
NeuronSearchLab. "NSL Recommender Leaderboard v3.2: which complete recommender system produces the best recommendations?" 19 September 2026. https://www.neuronsearchlab.com/research/recommender-leaderboard
A BibTeX entry, the licence terms and guidance on quoting a single number are on the methodology page. Cite the version rather than the page: the page shows the current release, and scores are not comparable across every release.
Go deeper
The full academic paper
Inside the paper
- Twenty-three complete recommender systems ranked by an index whose weights were fixed and hashed before any result existed.
- Every entrant labelled official, reference, reproduction or architecture proxy - with each proxy stating exactly what it is not.
- Scenario leaderboards for cold users, cold items, sparse and dense histories, long tail, coverage, novelty, diversity and multi-objective value.
- Stage-level diagnostics explaining why each system wins, including a batch-invariance test for ranking architectures that claim it.
- A full cost study: training cost, inference cost per thousand slates, latency, model size and retraining requirements against a stated reference deployment.
- Exposure-bias analysis on the datasets whose logs record impressions rather than only engagements, including a randomised-exposure check of whether offline evaluation picks the same winner as an unbiased one.
- Two compute tracks reported separately, so a weak architecture can be told apart from an under-trained one.
- The whole ranking re-scored under a second normalisation as a robustness check, plus the one defect we found and fixed before publishing.
The leaderboard, the headline findings, the scenario winners and every caveat on this page stay free and ungated, permanently. Other NSL research is at neuronsearchlab.com/research.