Retrieval & RAG
Semantic Search
A two-stage retrieval pipeline: a bi-encoder pulls candidates out of a vector index, then a cross-encoder reranker reorders them before they reach the user.

What it does
- Two-stage retrieval (bi-encoder MiniLM-L6 for recall, cross-encoder reranker for precision) over 100K MS MARCO passages.
- Reported NDCG@10 of 0.692 on that benchmark: 143% above BM25 and 12% above the bi-encoder alone.
- Evaluation harness measuring NDCG@10, MRR@10, and Recall@100 across three retrieval systems end to end.
- Production API served with FastAPI and Qdrant, with Apple Silicon MPS acceleration, containerised with Docker.
Built with
- sentence-transformers
- Qdrant
- cross-encoder
- FastAPI
- MS MARCO
- Docker