Modern AI Engineering

Lesson 10.5 · 25 min

Rerankers: Re-Scoring Retrieved Results by Relevance

Vector search found 50 "relevant" passages in 20 milliseconds, but the one that actually answers the question is sitting at position 17. How do we get it to the top?

In short: A reranker is a second, more precise model that re-scores a short list of candidates from a fast first-stage search. First-stage retrievers (BM25, bi-encoders) score documents independently of the query, which is fast but loose; a cross-encoder reranker reads the query and each document together and judges relevance much more accurately, at a much higher cost per document. Two-stage retrieval gets most of the accuracy of the slow model at close to the speed of the fast one.

What is a reranker?

Our running example: a RAG assistant for a company's HR policies. An employee asks "Can I carry unused vacation days into next year?" Vector search returns 50 passages about vacation: how to request leave, public holidays, sick days, and somewhere in there the carry-over rule. The LLM will only see the top 5. If the carry-over passage is 17th, the answer will be wrong or vague.

A reranker is a model that takes the query and a short list of candidate documents and reorders them by how well each one actually answers the query. It does not search the whole collection; it only re-scores what the first search found. Its output is a new ranking (and usually a relevance score per document), and we keep the top few.

Think of it like hiring A recruiter skims 1,000 CVs for keywords in an afternoon and shortlists 30. Then a hiring manager reads those 30 carefully, maybe interviews them, and picks the best 3. The skim is fast but rough; the careful read is slow but accurate. Nobody can interview 1,000 people, and nobody should hire from keywords alone.

Where a reranker sits: the two-stage retrieval idea

Two-stage retrieval splits search into a cheap, wide first stage and an expensive, narrow second stage.

The first stage is optimised for recall: make sure the right document is somewhere in the candidate list. The second stage is optimised for precision: make sure the very top positions are the right ones. Recall means "did we find it at all?"; precision@k means "of the top k results, how many are relevant?".

Why first-stage retrieval is fast but not precise

The speed of vector search comes from one design choice: the document vectors are computed before any query arrives. The query is embedded alone, the documents were embedded alone, and the search just compares two vectors. The model never sees the query and the document together.

That means one vector of 768 numbers must summarise everything a passage could be relevant for, without knowing what will be asked. Fine details get squeezed out: negation ("days cannot be carried over"), conditions ("only for staff hired before 2024"), which entity does what to whom. Two passages about vacation days look almost identical as vectors even if only one answers this question. BM25 has a similar limit: it counts word matches but does not understand them.

Pause and think: Why can a bi-encoder store document vectors in advance, while a cross-encoder cannot?

A bi-encoder encodes each document independently of any query, so the vector is the same for every future query and can be precomputed. A cross-encoder's score depends on the query and document read together, so it can only be computed once the query is known.

Bi-encoder vs cross-encoder

An encoder here is a Transformer model (BERT-like) that reads text and produces vectors. The two designs differ in what it reads at once.

A bi-encoder encodes the query and document separately into one vector each and scores them with cosine similarity or a dot product. It is used for the first stage. A cross-encoder concatenates them into one input, like [CLS] query [SEP] document [SEP], runs the full Transformer over both, and outputs a single relevance score from a small classification head. Inside, every query word can attend to every document word in every layer (this is cross-attention between the two texts via self-attention over the joined input), so it can notice that "carry into next year" matches "roll over to the following calendar year" and that "cannot" flips the meaning.

The landmark demonstration that a BERT cross-encoder makes a strong reranker was Nogueira and Cho's 2019 work on passage re-ranking (often called monoBERT), which reranked BM25 candidates on the MS MARCO benchmark and improved results by a large margin over earlier methods.

How a reranker scores documents, step by step

Inside a cross-encoder reranker

  1. Take the candidates: Receive the query and the top N candidates (say 50) from the first stage.
  2. Build pairs: Create 50 inputs: [CLS] query [SEP] candidate_i [SEP]. Long candidates are truncated to the model's maximum length (often 512 tokens).
  3. Run the model on each pair: The Transformer reads each pair jointly; pairs are processed in batches on a GPU for speed.
  4. Read off a score: A small head on top outputs one number per pair, a relevance logit, sometimes squashed to 0..1 with a sigmoid.
  5. Sort and cut: Sort the 50 candidates by this score and keep the top k (say 5). Optionally drop any below a minimum score.

Small example with illustrative numbers. Stage-1 cosine scores: "How to request leave" 0.82, "Public holidays 2026" 0.80, "Carry-over of unused days" 0.78. A cross-encoder might score the same three as 0.10, 0.03 and 0.94: once the model reads the question and the passage together, the carry-over passage is obviously the answer and moves from 3rd to 1st.

Scores are for sorting, not for comparing across queries Raw reranker scores are mainly meaningful within one query's candidate list. Some vendors calibrate them to 0..1 so a threshold is usable, but a "0.6" for one query is not necessarily as good as a "0.6" for another. Test before using a fixed cut-off.

The accuracy vs latency and cost trade-off

Every extra candidate we rerank costs one more model pass. More candidates mean a better chance that the right document is in the list (higher recall going in), but also more latency and more money. The simulation below makes this concrete.

two_stage_simulation.py

import numpy as np
rng = np.random.default_rng(42)
N_DOCS, N_QUERIES, TOP = 10_000, 200, 5
def simulate(first_k):
hits, calls = 0, 0
for _ in range(N_QUERIES):
relevant = np.zeros(N_DOCS)
relevant[rng.choice(N_DOCS, 5, replace=False)] = 1.0   # 5 truly relevant docs
# Stage 1 (bi-encoder): cheap but noisy score for every doc
fast = relevant + rng.normal(0, 0.3, N_DOCS)
if first_k == 0:                       # no reranker at all
final = np.argsort(-fast)[:TOP]
else:
cand = np.argsort(-fast)[:first_k]  # shortlist
# Stage 2 (cross-encoder): precise but one model call per candidate
precise = relevant[cand] + rng.normal(0, 0.15, len(cand))
calls += len(cand)
final = cand[np.argsort(-precise)[:TOP]]
hits += relevant[final].sum()
return hits / (N_QUERIES * TOP), calls / N_QUERIES
for k in [0, 20, 100, 1000]:
p, c = simulate(k)
label = "stage 1 only" if k == 0 else f"rerank top {k}"
print(f"{label:15} precision@5 = {p:.2f}   reranker calls/query = {c:.0f}")

Output:

stage 1 only    precision@5 = 0.44   reranker calls/query = 0
rerank top 20   precision@5 = 0.68   reranker calls/query = 20
rerank top 100  precision@5 = 0.84   reranker calls/query = 100
rerank top 1000 precision@5 = 0.99   reranker calls/query = 1000

In real systems, typical shortlist sizes are 20 to 200. Small cross-encoders (MiniLM-sized, tens of millions of parameters) can rerank 100 short passages in tens of milliseconds on a GPU; larger rerankers or LLM-based rerankers are slower and pricier. The right number depends on your latency budget, so measure end-to-end quality and latency together.

Pause and think: In the simulation, why does reranking the top 20 give lower precision than reranking the top 100, even though the reranker is equally accurate in both cases?

Because the reranker can only reorder what stage 1 handed it. With a shortlist of 20, some relevant documents never made it into the list, so no amount of reranking can recover them. A larger shortlist raises the recall going into stage 2.

Late-interaction models like ColBERT

There is a middle ground between bi- and cross-encoders. Late-interaction models such as ColBERT (Khattab and Zaharia, 2020) encode the query and document separately, like a bi-encoder, but keep one vector per token instead of one per text. At scoring time, each query token finds its most similar document token (the MaxSim operation), and these best matches are summed. Document token vectors can be precomputed, so it is much faster than a cross-encoder, and the token-level matching makes it more precise than a single-vector bi-encoder. The price is a much larger index. ColBERT gets its own lesson next.

Where each model type fits (general behaviour, not benchmark numbers).
Model typeInteractionPrecompute docs?Typical role
Bi-encoderNone (one vector each)YesFirst-stage retrieval
Late interaction (ColBERT)Token-level, after encodingYes (many vectors)Retrieval or reranking
Cross-encoderFull, inside every layerNoReranking a shortlist
LLM rerankerFull, plus reasoningNoHigh-value, low-volume reranking

Real rerankers, and why they matter for RAG

Rerankers you will meet Open models: the cross-encoder/ms-marco-MiniLM family in the Sentence-Transformers library, BAAI's bge-reranker models, Jina's rerankers and mixedbread's rerankers. Hosted APIs: Cohere Rerank, Voyage AI rerankers and reranking features in cloud search services. LLMs can also rerank by being prompted to order passages (listwise reranking, as in RankGPT). Names and versions change often, so check current model cards and benchmarks such as BEIR or MTEB-style reranking leaderboards.

Why do rerankers matter so much for RAG specifically? An LLM's context is limited and costly, and LLMs tend to use information at the start and end of the context better than the middle. So what we put in the top 3–5 slots largely decides the answer quality. A reranker lets us retrieve generously (high recall) but pass only a few, highly relevant passages (high precision). It often also reduces hallucinations, because irrelevant but similar-looking passages are pushed out of the context.

Common mistakes Reranking too few candidates (the right document never reaches the reranker). Reranking too many (latency explodes). Feeding chunks longer than the reranker's max length, so the important part gets truncated. Using an English-only reranker on multilingual content. Assuming reranker scores are calibrated across queries. Never measuring: compare answer quality with and without the reranker on real questions.

When might we skip a reranker? If the first stage is already precise enough for the task (measure it), if latency budgets are extremely tight, or if the collection is tiny enough to send everything to the LLM. Otherwise, adding a reranker is one of the cheapest, most reliable upgrades to a RAG system.

Worked example, step by step

How many candidates should we rerank? We do not have to guess. Two measurements, a latency budget and the recall of stage 1 at different depths, give the answer. All numbers here are illustrative.

Choosing the shortlist size

  1. Set the budget: Our HR assistant must have its passages ready within 200 ms, before the LLM starts writing. Stage 1 takes 30 ms, so the reranker may use 170 ms.
  2. Measure reranker speed: On our hardware the reranker scores 100 passages in 120 ms, so about 1.2 ms per passage. 170 / 1.2 ≈ 140 passages fit in the budget.
  3. Measure stage-1 recall by depth: On a test set, check how often the right passage is somewhere in the top N of stage 1, for several N. See the table.
  4. Read the ceiling: Stage-1 recall at N is a ceiling for the whole system. If the right passage is in the shortlist only 70% of the time, no reranker can do better than 70%.
  5. Pick N: Going from 20 to 50 buys 18 points of recall for 36 ms. Going from 100 to 200 buys 2 points for 120 ms and breaks the budget. We pick N = 100, or 50 if we want headroom.
Illustrative measurements for one system.
Shortlist NStage-1 recall at NRerank time
200.7024 ms
500.8860 ms
1000.93120 ms
2000.95240 ms

One more thing to check while we are measuring: truncation. Suppose the reranker reads at most 512 tokens per pair. With a 20-token query and a few special tokens, only about the first 490 tokens of the passage are read. If our chunks are 800 tokens long and the carry-over rule sits in the last paragraph, the reranker never sees it and scores the chunk low. To spot this, log the token length of each pair. To fix it, use shorter chunks, or split a long chunk into pieces and keep the best piece score.

Practice: try it yourself

We will run a complete two-stage search on five real sentences. Stage 1 uses pretend bi-encoder scores. Stage 2 is a tiny stand-in for a cross-encoder: it reads the query and the passage together and counts the short phrases (neighbouring word pairs) they share. A real cross-encoder is a neural network, but the shape of the pipeline is the same.

practice_two_stage.py

query = "can i carry unused vacation days into next year"
# (passage, pretend stage-1 cosine score from a bi-encoder). Illustrative numbers.
candidates = [
("how to request vacation days for next year", 0.84),
("vacation days and sick days are tracked separately", 0.81),
("unused vacation days carry into next year up to five days", 0.79),
("public holidays for next year are in the calendar", 0.70),
("parking permits are renewed every year", 0.35),
]
def pairs(text):                       # neighbouring word pairs = short phrases
w = text.split()
return set(zip(w, w[1:]))
def careful_score(q, passage):         # stand-in cross-encoder: reads both together
return len(pairs(q) & pairs(passage))
def two_stage(shortlist_size):
shortlist = sorted(candidates, key=lambda c: -c[1])[:shortlist_size]     # stage 1
reranked = sorted(shortlist, key=lambda c: -careful_score(query, c[0]))  # stage 2
return reranked[0][0], len(shortlist)
for passage, s1 in candidates:
print(f"stage1={s1:.2f}  careful={careful_score(query, passage)}  {passage}")
print("stage 1 only ->", candidates[0][0])
for n in [2, 3, 5]:
best, calls = two_stage(n)
print(f"rerank top {n} -> {best}  ({calls} reranker calls)")

Output:

stage1=0.84  careful=2  how to request vacation days for next year
stage1=0.81  careful=1  vacation days and sick days are tracked separately
stage1=0.79  careful=4  unused vacation days carry into next year up to five days
stage1=0.70  careful=1  public holidays for next year are in the calendar
stage1=0.35  careful=0  parking permits are renewed every year
stage 1 only -> how to request vacation days for next year
rerank top 2 -> how to request vacation days for next year  (2 reranker calls)
rerank top 3 -> unused vacation days carry into next year up to five days  (3 reranker calls)
rerank top 5 -> unused vacation days carry into next year up to five days  (5 reranker calls)

Now change it:

  • Change the stage-1 score of the carry-over passage from 0.79 to 0.60. Predict the smallest shortlist size that still finds it.
  • Change the query to "are sick days tracked separately". Work out the careful score of each passage by hand, then predict the winner when we rerank the top 5.
  • Replace the body of careful_score with a plain shared-word count: len(set(q.split()) & set(passage.split())). Predict whether the right passage still wins, and which passages now tie.

Pause and think: Reranking the top 5 cost 5 calls and gave the same top result as reranking the top 3. Does that mean a shortlist of 3 is the right setting?

Not from one query. Here the right passage happened to be third in stage 1. For the next query it may be 4th or 40th, and we cannot know in advance. The shortlist size should come from stage-1 recall measured over many queries, weighed against the latency budget.

Pause and think: With a shortlist of 2 the pipeline returned the same wrong passage as stage 1 alone, even though the careful scorer gives the right passage the highest score of all (4). What does that tell us about where to look when a reranked system fails?

Look at stage 1 first. The reranker never saw the right passage, so its quality did not matter. Before blaming or swapping the reranker, check whether the right passage is in the shortlist at all.

Key takeaways

  • A reranker re-scores a shortlist from a fast first stage with a slower, more precise model.
  • Two-stage retrieval: stage 1 maximises recall over the whole collection; stage 2 maximises precision at the top.
  • Bi-encoders encode query and document separately (fast, precomputable); cross-encoders read them together (accurate, expensive).
  • Shortlist size trades quality against latency and cost; a reranker cannot recover documents stage 1 missed.
  • Late-interaction models like ColBERT sit between the two, matching at the token level with precomputed vectors.

Key terms

  • Reranker: A model that reorders a short list of retrieved candidates by relevance to the query.
  • Two-stage retrieval: A cheap wide search for candidates followed by an expensive precise re-scoring.
  • Bi-encoder: A model that encodes query and document separately into vectors compared by similarity.
  • Cross-encoder: A model that reads query and document together and outputs a single relevance score.
  • Precision@k: The fraction of the top k results that are relevant.
  • Late interaction: Encoding texts separately into token vectors and matching them at scoring time, as in ColBERT.

← 10.4 Hybrid Search: Combining Sparse and Dense Retrieval · 10.6 ColBERT: Token-Level Late Interaction for Retrieval →