Oliver Wakeford
All projects
Information Retrieval
Information RetrievalShipped2026

Scientific Citation Retrieval

0.749 NDCG@10 on the final leaderboard, team of four

Information RetrievalNLPEnsemblesEmbeddings

Publication

Course challenge, hosted on Codabench

Final held-out score: 0.749 NDCG@10

About 0.02 behind the top score. Team of four.

Built with

PythonSentence Transformersrank-bm25FAISSscikit-learn

A course challenge on Codabench, done by our team of four. Given a query paper, return the 100 papers from a 20,000-paper corpus it most likely cites, scored on NDCG@10 over held-out queries. Each query has about seven relevant papers, spread across 19 fields.

The pipeline we presented mid-competition fuses eight rankings with weighted reciprocal rank fusion. Three come from dense encoders, including SPECTER2 and MiniLM. The other five are lexical, mostly BM25 run over different parts of each paper, down to the sentences around each citation marker in the query. BM25 on the full text was the best single signal, ahead of every dense model in that pipeline. Fusion took the MiniLM baseline from 0.507 to 0.615 on the public queries.

Only 18 of the 736 citation pairs in the public set cross a field boundary. So we multiplied the scores of same-field candidates by 10, and same-venue ones by 2. That one change added about 0.10, roughly half of everything we gained. The presented pipeline scored 0.697 on the held-out queries, against 0.71 to 0.72 locally, and part of that gap is overfitting, since the public queries were also our tuning set.

We kept submitting until the deadline. The final submission scored 0.749 NDCG@10 on the held-out queries, about 0.02 behind the top score on the leaderboard.

A lot of what we tried made things worse. A cross-encoder reranker trained on web search dropped the score from 0.592 to 0.476, and fine-tuning it on our own judgements didn't move it at all. An LLM listwise reranker got 0.557. The one worth describing is a publication-year filter, on the idea that a paper can't cite the future. It fell to 0.378, because 68.5% of the papers labelled as cited are newer than the query, by 4.3 years on average. The labels aren't a plain citation graph, and we should have checked that before assuming it.

What this does not show

  • A four-person team project, so the write-up says 'we'.
  • A course leaderboard with 37 accounts on it, which is a small field next to a public benchmark.
  • Fusion weights and boosts were tuned on the 100 public queries, so local scores (0.71 to 0.72) run above held-out.
  • The public repository reproduces the pipeline we presented (0.697 held out). The last-day code behind the final submission isn't in it.
  • About 10% of relevant papers (73 of 736) never reach any retriever's top 300, which caps what reranking can do.