NSL Research · Recommender Benchmark v3.0 · 23 August 2026
Best Recommender Systems
23 complete recommender systems — classical, neural, graph, sequential, multi-stage and generative — given the same data and the same catalogue, and judged on the only thing a user ever sees: the ranked slate they produce. Quality, cost and latency are measured separately, because the trade-off between them is yours to make.
ItemKNN → Multi-task ranker
Overall winner
Recommender Performance Index 68.8 — a percentile in a frozen, published distribution of every entrant's measurements. Nothing is pinned to a value: ItemKNN places 9 of 23 and the popularity baseline 21.
4
Different scenario winners
Across 9 scenario leaderboards. No architecture dominates every regime.
5
Generative entrants
Best is PLUM-style SID generative retrieval, ranked 11 overall.
10
Public datasets
Media, feeds, news, e-commerce, music and books; dense and sparse; small and very large catalogues.
The leaderboard
Ranked by recommendation quality
Sort by any column, filter by architecture family, and open a row for what the system actually is and where its score comes from. The Recommender Performance Index covers relevance, cold start, long tail, novelty and diversity, multi-objective value and robustness across data regimes — and nothing else. Cost and latency sit beside it.
| System | Architecture | |||||||
|---|---|---|---|---|---|---|---|---|
| 1 | ItemKNN → Multi-task rankerReproductionTied with nextFrontier Multi-stageTraditional | 68.8 | 68 | $64 | 3.45 ms | 7 | Multi-stageTraditional | |
| 2 | ItemKNN → MLP rankerReferenceTied with nextFrontier Multi-stageTraditional | 66.9 | 65 | $64 | 3.42 ms | 7 | Multi-stageTraditional | |
| 3 | VS-KNNReferenceTied with nextFrontier NeighbourhoodTraditional | 65.2 | 63 | $35 | 1.88 ms | 13 | NeighbourhoodTraditional | |
| 4 | EASEᴿReferenceTied with nextFrontier Linear autoencoderTraditional | 64.5 | 64 | $4.76 | 0.25 ms | 91 | Linear autoencoderTraditional | |
| 5 | SASRecReproductionTied with nextFrontier SequentialTraditional | 62.9 | 59 | $1.83 | 0.12 ms | 230 | SequentialTraditional | |
| 6 | VS-KNN → GBDTReferenceTied with next Multi-stageTraditional | 60.8 | 57 | $94 | 5.07 ms | 4 | Multi-stageTraditional | |
| 7 | Recency-Weighted PopularityReferenceTied with nextFrontier Non-personalisedTraditional | 60.2 | 58 | $1.45 | 0.09 ms | 277 | Non-personalisedTraditional | |
| 8 | LightGCNReproductionTied with next GraphTraditional | 59.8 | 59 | $1.88 | 0.09 ms | 213 | GraphTraditional | |
| 9 | ItemKNNReferenceTied with next NeighbourhoodTraditional | 58.1 | 56 | $3.89 | 0.25 ms | 100 | NeighbourhoodTraditional | |
| 10 | BERT4RecReproductionTied with next SequentialTraditional | 57.2 | 54 | $1.71 | 0.11 ms | 223 | SequentialTraditional | |
| 11 | PLUM-style SID generative retrievalProxyTied with next Generative · semantic IDGenerative | 55.0 | 53 | $92 | 4.81 ms | 4 | Generative · semantic IDGenerative | |
| 12 | Two-Tower → DCN-v2ReferenceTied with next Multi-stage (neural)Traditional | 53.6 | 48 | $58 | 3.07 ms | 6 | Multi-stage (neural)Traditional | |
| 13 | GRU4RecReproductionTied with nextFrontier SequentialTraditional | 52.2 | 48 | $1.42 | 0.11 ms | 246 | SequentialTraditional | |
| 14 | TIGERReproductionTied with next Generative · semantic IDGenerative | 51.6 | 47 | $33 | 1.67 ms | 11 | Generative · semantic IDGenerative | |
| 15 | Two-Tower → RankMixer-style rankerProxyTied with next Multi-stage (neural)Traditional | 51.3 | 45 | $59 | 3.14 ms | 6 | Multi-stage (neural)Traditional | |
| 16 | iALSReferenceTied with next Matrix factorisationTraditional | 51.0 | 49 | $2.86 | 0.13 ms | 119 | Matrix factorisationTraditional | |
| 17 | EASEᴿ → LambdaMARTReferenceTied with next Multi-stageTraditional | 50.9 | 47 | $76 | 3.87 ms | 4 | Multi-stageTraditional | |
| 18 | HSTU (Generative Recommenders)ReproductionTied with next Generative · sequentialGenerative | 49.6 | 45 | $7.63 | 0.18 ms | 44 | Generative · sequentialGenerative | |
| 19 | Netflix-style foundation modelProxyTied with next Foundation modelGenerative | 46.0 | 42 | $3.70 | 0.20 ms | 83 | Foundation modelGenerative | |
| 20 | OneRec-style session generatorProxyTied with next Generative · sequentialGenerative | 42.4 | 37 | $115 | 6.15 ms | 2 | Generative · sequentialGenerative | |
| 21 | PopularityReference Non-personalisedTraditional | 38.7 | 35 | $1.77 | 0.12 ms | 146 | Non-personalisedTraditional | |
| 22 | Content TF-IDFReferenceTied with next NeighbourhoodTraditional | 37.7 | 32 | $3.59 | 0.22 ms | 70 | NeighbourhoodTraditional | |
| 23 | Phoenix-style unified transformerProxy Unified transformerTraditional | 27.1 | 21 | $5.62 | 0.29 ms | 32 | Unified transformerTraditional |
Showing 23 of 23 systems, scored on the best-achievable compute budget; open a row to see the same system on the compute-constrained budget. Click any row for what the system is, where it comes from, and what its score is made of. The Recommender Performance Index is quality only — cost and latency sit beside it, never inside it.
What the results say
Key findings
Every sentence below is derived from the stored results rather than written by hand, so a rerun of the benchmark rewrites them.
ItemKNN → Multi-task ranker leads
It scores 68.8 on the Recommender Performance Index — 1.9 points ahead of ItemKNN → MLP ranker — on the best-achievable compute track across 10 datasets. A score is a percentile in a frozen distribution of every entrant's measurements, so it is read against the field rather than against a chosen system: the twenty-five-year-old ItemKNN baseline scores 58.1 and places 9th of 23, which is the more useful thing to calibrate against.
No system wins everywhere
Across 9 scenario leaderboards — cold users, cold items, sparse and dense histories, long tail, coverage, novelty, diversity and multi-objective value — 4 different systems take first place. The right question is not which recommender is best; it is which is best for the part of the distribution your product actually lives in.
Generative architectures do not yet lead at this scale
5 entrants decode item identifiers rather than scoring the catalogue. Their best index is 55.0 against 68.8 for the rest, and they take the top spot in 0 of 9 scenarios. The industrial reports behind these architectures all rest on a scaling claim that a public-data benchmark of this size cannot test.
The last few points of quality are the expensive ones
SASRec comes within 5.9 index points of the winner at $1.83 a month against $64 — 35.2× cheaper for a difference most products would struggle to detect.
Offline evaluation and randomised exposure pick different winners
One dataset in the panel inserts randomly-chosen items into real feeds. Exposure on those rows is independent of the incumbent recommender, so a plain average over them is an unbiased estimate of a slate's value — no propensity model, no assumption. Across 23 systems the rank correlation between the ordinary offline ordering and that unbiased one is 0.17. Offline picks SASRec; randomised exposure picks LightGCN. Offline evaluation here does not merely inflate scores — it can choose a different system. Read every number on this page with that in mind.
Quality against cost
What the top of the leaderboard actually costs
Quality is on the vertical axis and never mixed into the horizontal one. The dashed line is the Pareto frontier: nothing above and to the left of a point on it is both better and cheaper. Most deployment decisions are made against this shape rather than against the ranking.
One confound worth knowing. The multi-stage systems build their ranker’s features online, in Python, for every request here — work a deployed pipeline precomputes or caches. Their position on the horizontal axis is an upper bound, and the cost gap between them and the single-stage systems is overstated by an amount this benchmark does not try to estimate. Quality is measured on the served slate and carries no such caveat, which is why the index excludes cost entirely.
Where each architecture wins
Scenario winners
The same results sliced by the question a team usually has: what happens to a user with four interactions, to an item nobody has touched, to the long tail, to a feed with eight competing engagement signals.
Best at…
Best overall
ItemKNN → Multi-task ranker
Recommender Performance Index
Runner-up: ItemKNN → MLP ranker
Best generative recommender
PLUM-style SID generative retrieval
Recommender Performance Index, generative entrants only
Runner-up: TIGER
Best cold start
ItemKNN → Multi-task ranker
cold-start dimension
Runner-up: Recency-Weighted Popularity
Best long tail
Content TF-IDF
long-tail dimension
Runner-up: GRU4Rec
Best multi-objective
SASRec
multi-objective dimension
Runner-up: VS-KNN → GBDT
Best under a latency budget
ItemKNN → Multi-task ranker
Recommender Performance Index among systems serving under 25 ms/user offline
Runner-up: ItemKNN → MLP ranker
Best quality per dollar
Recency-Weighted Popularity
Recommender Value Index
Runner-up: GRU4Rec
Best retriever
EASEᴿ
candidate recall@100, systems with a separable retrieval stage
Runner-up: EASEᴿ → LambdaMART
Best ranker
ItemKNN → Multi-task ranker
NDCG@10 on identical candidate sets
Runner-up: ItemKNN → MLP ranker
Best novelty & diversity
GRU4Rec
novelty, diversity and coverage
Runner-up: iALS
The comparison everyone asks for
Generative against traditional
Semantic-ID decoders, autoregressive user models and session generators against neighbourhood models, matrix factorisation, graphs and multi-stage pipelines — on identical data, with identical budgets and identical evaluation.
5
Generative entrants
18
Traditional entrants
55.0
Best generative index
68.8
Best traditional index
Scenario by scenario, each normalised to its winner
What this comparison can support. Three of the generative entrants are architecture proxies for proprietary systems and one is trained from scratch where the original is adapted from a pre-trained language model; the comparison is between architectures at benchmark scale, not between the companies' production systems. The central claim of every industrial generative recommender is a scaling claim, and a benchmark of this size is structurally unable to test it. What it can test is whether the mechanism buys anything when the scale is absent — which is the situation almost every team is actually in.
Plain English
What the leading systems actually are
A short description of the top of the table, including what our implementation is and — where the system is proprietary — what it is not.
ItemKNN → Multi-task ranker
ReproductionThe same retrieval stage, but the ranker has one head per engagement signal — click, long view, like, share, follow, and the negative signals — combined with weights fixed before any result was seen. On feed datasets that log more than a click this is the system that can trade a little relevance for a lot of downstream value; on datasets with a single signal it collapses to a pointwise ranker and is reported as such.
Multi-task ranking over several engagement signals · Ma et al., 'Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts', KDD 2018; Zhao et al., 'Recommending What Video to Watch Next', RecSys 2019.
Index 68.8 · $64/mo · 3.45 ms p95 · GPU
ItemKNN → MLP ranker
ReferenceThe most common shape a real recommender takes when a team first adds machine learning on top of an existing collaborative-filtering service: keep the neighbourhood model as the candidate source, learn an MLP over retrieval scores, item statistics, user statistics and content-similarity features to reorder the top few hundred.
Classical retrieval with a neural pointwise ranker · Sarwar et al., WWW 2001 (retrieval); Covington, Adams & Sargin, RecSys 2016 (the deep-ranking stage this follows).
Index 66.9 · $64/mo · 3.42 ms p95 · GPU
VS-KNN
ReferenceSession-based nearest neighbours with position-decayed weighting of the current session. Ludewig & Jannach found it beat every neural session recommender they tested; it is the sequential family's classical control.
Vector-multiplication session-based kNN · Ludewig & Jannach, 'Evaluation of Session-based Recommendation Algorithms', UMUAI 2018.
Index 65.2 · $35/mo · 1.88 ms p95 · CPU only
EASEᴿ
ReferenceA linear item-item autoencoder with a closed-form solution: one regularised Gram-matrix inverse, no iterations, no hyperparameters beyond the ridge term. Steck's result — that this beats most of the deep models published against it — has held up in every independent replication since. Its cost is quadratic in catalogue size, so it is not attempted above roughly 25,000 items, and that skip is reported rather than hidden.
Embarrassingly Shallow AutoEncoder · Steck, 'Embarrassingly Shallow Autoencoders for Sparse Data', WWW 2019.
Index 64.5 · $4.76/mo · 0.25 ms p95 · CPU only
SASRec
ReproductionA causal transformer over the item sequence, trained with a softmax over the whole catalogue rather than the original binary objective. This is the strongest formulation of the architecture that the replication literature has converged on, and it is the reference point every generative entrant is really competing with.
Self-attentive sequential recommendation · Kang & McAuley, 'Self-Attentive Sequential Recommendation', ICDM 2018; trained with the full softmax of Klenitskiy & Vasilev, RecSys 2023.
Index 62.9 · $1.83/mo · 0.12 ms p95 · GPU
VS-KNN → GBDT
ReferenceSession-aware retrieval with a pointwise gradient-boosted ranker. Included because the sequential signal enters through the retriever rather than through a transformer, which separates 'sequence modelling helps' from 'transformers help'.
Session-based neighbourhood retrieval with a boosted ranker · Ludewig & Jannach, UMUAI 2018; Ke et al., NeurIPS 2017 (LightGBM).
Index 60.8 · $94/mo · 5.07 ms p95 · CPU only
Recency-Weighted Popularity
ReferencePopularity with an exponential recency kernel over the training window. On fast-moving catalogues — news, short video, fashion — this single change is worth more than most architectural ones, and it costs nothing to serve.
Standard non-personalised baseline with a recency kernel · Ji et al., 'A Re-visit of the Popularity Baseline', SIGIR 2020.
Index 60.2 · $1.45/mo · 0.09 ms p95 · CPU only
LightGCN
ReproductionRounds of neighbour averaging over the user–item graph, with every nonlinearity and feature transform removed. The high-order connectivity is the point: an item two hops away reaches a user who never touched anything it co-occurs with directly, which is where it beats a neighbourhood model — and where it is most exposed on sparse graphs, since two hops on a sparse graph mostly finds popularity.
Simplified graph convolution for collaborative filtering · He et al., 'LightGCN: Simplifying and Powering Graph Convolution Network for Recommendation', SIGIR 2020.
Index 59.8 · $1.88/mo · 0.09 ms p95 · GPU
Read this before quoting anything above
What this benchmark cannot tell you
These are not disclaimers. They are the boundary of what the measurements support, and a result quoted outside that boundary is a wrong result.
- 01
On the one dataset that inserts randomly-chosen items into real feeds — the only place an unbiased estimate is available — the offline ordering and the unbiased ordering correlate at 0.17 and name different winners. That is a limit on what any offline recommender benchmark, including this one, can tell you.
- 02
Public offline datasets are not production environments. Offline relevance is a proxy for user value, and the mapping between them is platform-specific.
- 03
Several entrants are architecture proxies for proprietary systems. They implement a published mechanism at benchmark scale and are not those companies' production systems.
- 04
No single recommender wins every scenario. The winner changes 4 ways across the scenario boards.
- 05
Serving latency is offline batch throughput on the benchmark host, not a production p99. Multi-stage systems additionally build their ranker's features online here, which a deployed pipeline precomputes — so their measured serving cost is an upper bound and the cost gap to single-stage systems is overstated.
- 06
Scale is the untested variable: the central claim of the industrial generative systems is a scaling claim this benchmark is too small to test.
- 07
The index scores each measurement by its percentile in a frozen, published distribution. That records where a system placed, not by how much: doubling the best cold-start result and beating it by a hair score the same. Raw metric values are published beside every score, and the scenario boards are stated in raw metrics.
- 08
This is index version v3.1 and its scores are not comparable to v3.0's. The v3.0 scale rescaled every metric so Popularity read 0 and ItemKNN 100; it was withdrawn before publication because on two of ten datasets Popularity beat ItemKNN, which inverted the scale, and where the two were close a worse result could reach -381. The effect was to flatter ItemKNN by about six places. Under v3.1 nothing is pinned and ItemKNN places 9th.
The benchmark disagrees with itself, and says so. On the one dataset that inserts randomly-chosen items into real feeds — the only place an unbiased estimate is available — the offline ordering and the unbiased ordering correlate at just 0.17, and they name different winners. That is a limit on what any offline recommender benchmark, including this one, can tell you. It is also the reason the full paper reports it rather than the leaderboard hiding it.
5 entrants are architecture proxies. A row named after a company’s system is an NSL implementation of the architecture that company published, run at benchmark scale on public data. It does not carry their scale, their data, their multimodal encoders or their unpublished components. No number on this page is evidence about the quality of any company’s production recommender.
Methodology
How the index is built
What the index measures
Recommendation quality and nothing else. Cost, latency, model size and retraining burden are reported beside it and combined only in the separate Recommender Value Index.
- Relevance35%
- Cold start15%
- Long tail15%
- Novelty & diversity10%
- Multi-objective10%
- Robustness15%
Weights frozen 2026-08-22, fingerprint 52706396658de50c, so you can check they were not adjusted after the results existed.
How it is scaled
A system’s score on a metric is its percentile in a frozen field: the distribution of what every entrant measured on that same metric and dataset, pooled across both compute tracks and then frozen. Nothing is scaled against a chosen system, so no system sits at a fixed position — the two reference systems included.
Freezing the distribution is what makes adding a new system to this leaderboard leave every existing score exactly where it was. A percentile taken against the current field would move all of them, which makes historical comparison impossible.
The field is 167 published constants, fingerprinted 05eb9f3f645e581b, so a score can be recomputed — or your own system scored against this benchmark — without rerunning anything. Open any row for the unnormalised measurements behind its score.
Re-scoring the identical results against the live field, which depends on no published constant at all, gives a rank correlation of 0.98 with this one, so the ordering is a property of the measurements rather than of the frozen constants.
Index version v3.1 replaced a scale that was withdrawn before publication. v3.0 used a two-point anchored scale dividing by (itemknn - popularity). It was withdrawn: on two of ten datasets that difference was negative, inverting the scale, and where it was small a worse result could reach -381 before being floored. Scores are not comparable across the change. Neither failure could touch the reference system itself, which sat on 100 by arithmetic — the effect was to flatter it by about six places against three independent checks that use no normalisation at all.
What the cost figures assume
A stated reference deployment: 1,000,000 active users, 10 slates per user per month (10,000,000 slates), retrained 4.35 times a month, at on-demand cloud rates. The per-slate and per-retrain figures it multiplies are measured; the multipliers are assumptions, published so you can redo the arithmetic for your own traffic.
Datasets
All public, used under their stated licences, none redistributed. Chronological splits, full-catalogue metrics with no sampled negatives, user-level bootstrap intervals, paired permutation tests with multiplicity correction.
2 systems could not be attempted on 4 datasets and are recorded as skipped rather than dropped — for example: catalogue has 39979 items; this system is not attempted above 25000 by policy
NSL Recommender Benchmark v3.0 · generated 23 August 2026 · best-achievable compute track · scoring weights frozen 2026-08-22, fingerprint 52706396658de50c. The leaderboard is meant to grow: a new architecture is a class in the harness and a rerun, and because every score is a percentile in a frozen, published distribution rather than a rank against the current field, adding one leaves every score already published exactly where it was. If a system you care about is missing, tell us.
Go deeper
The full academic paper
Inside the paper
- Twenty-three complete recommender systems ranked by an index whose weights were fixed and hashed before any result existed.
- Every entrant labelled official, reference, reproduction or architecture proxy — with each proxy stating exactly what it is not.
- Scenario leaderboards for cold users, cold items, sparse and dense histories, long tail, coverage, novelty, diversity and multi-objective value.
- Stage-level diagnostics explaining why each system wins, including a batch-invariance test for ranking architectures that claim it.
- A full cost study: training cost, inference cost per thousand slates, latency, model size and retraining requirements against a stated reference deployment.
- Exposure-bias analysis on the datasets whose logs record impressions rather than only engagements, including a randomised-exposure check of whether offline evaluation picks the same winner as an unbiased one.
- Two compute tracks reported separately, so a weak architecture can be told apart from an under-trained one.
- The whole ranking re-scored under a second normalisation as a robustness check, plus the one defect we found and fixed before publishing.
The leaderboard, the headline findings, the scenario winners and every caveat on this page stay free and ungated, permanently. Other NSL research is at neuronsearchlab.com/research.