Research · v1
Comparing Recommender Systems Under a Leakage-Free Evaluation Protocol
Nine recommendation approaches measured on a chronological split, ranked against the full catalogue, with confidence intervals and negative results.
Most disagreements between recommender benchmarks come from the evaluation, not the models. We rebuilt the measurement first: one global chronological cut-off, full-catalogue ranking with no sampled negatives, tuned classical baselines, and paired significance tests over identical users.
Under that protocol the ordering is not the one recent literature predicts. Neighbourhood methods lead, non-personalised popularity is far more competitive than its reputation, and the training objective — not the architecture — turns out to carry almost all of the difference for two-tower retrieval.
What the report measures
39.9%
Split protocol
of training examples under a random-hash split contained data the same split had reserved for testing.
+44%
Reported score inflation
higher NDCG@10 from that split than from a chronological one, on identical data and an identical model.
3.4×
Retrieval objective
improvement from changing the retrieval loss alone, replicated across three random seeds.
Inside
- Nine recommendation approaches compared on identical data and splits.
- A leakage audit quantifying what each split protocol costs in credibility.
- The accuracy price of diversity, calibration and supplier caps.
- An ablation across three seeds isolating which change to a two-tower retrieval model actually matters.
- Results that did not replicate, reported alongside the ones that did.
- Full method: metrics, cohorts, significance testing and stated limitations.
Datasets
All datasets are public and used under their respective research licences. None is redistributed, and no customer data was used in any experiment.