Skip to content

Retrieval & RAG

Semantic Search

A two-stage retrieval pipeline: a bi-encoder pulls candidates out of a vector index, then a cross-encoder reranker reorders them before they reach the user.

Illustration. A printed page of prose on a desk with one passage picked out in yellow highlighter, and a small search control resting beside it.

What it does

  • Two-stage retrieval (bi-encoder MiniLM-L6 for recall, cross-encoder reranker for precision) over 100K MS MARCO passages.
  • Reported NDCG@10 of 0.692 on that benchmark: 143% above BM25 and 12% above the bi-encoder alone.
  • Evaluation harness measuring NDCG@10, MRR@10, and Recall@100 across three retrieval systems end to end.
  • Production API served with FastAPI and Qdrant, with Apple Silicon MPS acceleration, containerised with Docker.

Built with

  • sentence-transformers
  • Qdrant
  • cross-encoder
  • FastAPI
  • MS MARCO
  • Docker

← All projects