Choosing the Reranker

Six rerankers evaluated, one vendor benchmark thrown out, and what the evidence does not support

Embedder Benchmarks ← Back to search

Correction, 2 October 2026. The October re-test below first reported that ettin-reranker-400m ranks better than the deployed model, nDCG@5 0.6564 against 0.5412. Those labelled figures came from six template queries. Re-scored against 20 real questions labelled by hand, the gap halved to 0.6562 against 0.5942 and is inside noise. The claim is withdrawn. The decision to keep the deployed model is unchanged.

Correction, September 2026. This page recorded Qwen3-Reranker-4B as the deployed reranker. It is not what runs. The live endpoint serves Qwen3-Reranker-0.6B-seq-cls, confirmed from the endpoint's own /v1/models response and the application configuration.

The swap was deliberate, not accidental. On this workload the 4B scored roughly 4% better and ran about 125% slower. That is not a trade worth making for an interactive search box. The page simply never caught up with the decision, and the endpoint is still named for the 4B, which is how it went unnoticed.

Experiments have since been run against this corpus. All are reported below, including the ones that failed. See Measured on this corpus, and the October 2026 re-test in which the deployed model was checked against three alternatives and kept.

This page shows how the embedding and reranking models running this demo were selected, what was measured, and — the part most write-ups skip — what the evidence does not support. Every number below is either measured against the live endpoints or a published benchmark with its source named. The two are never mixed.

The finding worth three orders of magnitude

Qwen3-Reranker scores through its chat template. Testing that wrapping applied client-side against posting raw query/document text — same model, same endpoint, same documents:

4,600×relevant vs irrelevant separation, chat-template wrapping
3.3×the same test, raw prompt
918 mswarm query, end to end

Both calls hit the same live endpoint, which at the time of this July 2026 test served Qwen3-Reranker-4B-seq-cls, 329 prompt tokens, three documents and one query:

Instruct: Given a code search query, retrieve relevant code chunks that answer the query.

Query: how is user authentication handled

DocumentChat-template wrappedRaw
PasswordSignInAsync(...) — the answer0.16188657280.375985
LogTimestamp(...) — logging helper0.00003464250.105656
CalculateBasketTotal(...) — unrelated0.00000532750.113646
Separation, top vs next4,673×3.31×

The takeaway isn't the model — and it isn't that raw "fails." Raw still ranked the authentication method first. It got the answer right. What it lost was discrimination: it scored a basket-total calculation and a logging helper within 8% of each other, and both within a factor of 3.5 of the correct answer.

Wrapped, the correct answer sits 4,673× above the next candidate and 30,388× above the worst. That margin is what makes a score threshold possible at all. Three orders of magnitude of usable signal lived in the prompt formatting, not the weights.

What the demo showed, and why it was not a test

Correction. This section previously said there was “no keyword overlap engineered in.” That was wrong. Every one of these six questions contains a word that appears in its own answer’s filename — send email against EmailSender.cs, the basket total against Basket.cs. Plain keyword matching can win them with no embedding model involved, so they show the system works end to end but prove nothing about retrieval quality. Noticing that is what prompted the experiments below and a question set built to discriminate instead.

Six plain-English questions fired at the live index — 254 files, 605 chunks, health green. Kept here because they are what the demo actually showed, including the one it got wrong.

QuestionTop results
how does the app send emailEmailSender.cs · Program.cs · OrderController.cs
how is the basket total calculatedBasket.cs · BasketViewModel.cs · Order.cs · OrderTotal.cs
how are users authenticated at sign inAuthenticateEndpoint.cs · ManageController.cs · CustomAuthStateProvider.cs
how are catalog items filtered by brand and typeCatalogItemService.cs · CatalogViewModelService.cs · CatalogBrand.cs
where is the database connection configuredAppIdentityDbContextSeed.cs · IdentityHostingStartup.cs
what happens when a user checks outConfigureCookieSettings.cs · ApiHealthCheck.cs — miss

Five hit, one missed — and the miss is the useful one. "What happens when a user checks out" returned cookie settings and health checks: nothing to do with checkout. Rewording found the boundary:

RewordedTop results
create an order from the basketOrder.cs · Basket.cs · OrderService.cs
how is an order placedOrder.cs · GetMyOrdersHandler.cs · GetOrderDetailsHandler.cs
checkout controller actionBaseApiController.cs — still weak

The corpus has no "checkout" concept — it models the same idea as orders. The retrieval wasn't broken; the vocabulary was. Naming the domain concept fixes it instantly, and no amount of reranking would have.

The keyword leg, tested separately

Exact identifiers should not need semantics, and they don't: EmailSender → EmailSender.cs, IEmailSender.cs, EmailSenderExtensions.cs · OrderService → OrderService.cs, IOrderService.cs · IRepository → IRepository.cs. PaymentMethods returned exactly one result — Buyer.cs — rather than padding the list with near-misses.

Six rerankers evaluated

ModelParamsCode benchmarkOutcome
Qwen3-Reranker-4B4BMTEB-Code 81.20chosen on paper; ~4% better, ~125% slower, so not deployed
Qwen3-Reranker-8B8BMTEB-Code 81.22+0.02 for ~2× cost — rejected
Qwen3-Reranker-0.6B0.6BCoIR 65.18deployed: the 4B's quality edge did not pay for its latency
mxbai-rerank-base-v20.5BCoIR 65.71≈ tie, kept as A/B arm
bge-reranker-v2-m3 (incumbent)0.57BCoIR 36.28replaced
jina-reranker-v3 / jina-code—strongCC-BY-NC — license-blocked
zerank-1-small / zerank-21.7B / 4BnDCG@10 0.6528CC-BY-NC or ~3× cost

The incumbent was not bad — it remains strong on multilingual text (MIRACL 69.32, best in its table). It was bad specifically on code, at roughly half the achievable nDCG.

Why the headline benchmark was thrown out

The widely-quoted figure for this comparison is MTEB-Code 73.42 vs 41.38. It did not survive checking: that number is the vendor scoring its own model on its own retrieved candidates. Not independent.

So the claim was verified from a primary source instead — arXiv 2509.25085 v3, Table 2, published by a rival lab:

BenchmarkQwen3-Reranker-0.6Bbge-reranker-v2-m3Read
CoIR (code)65.1836.28~29 pt code gap — real
BEIR (general)56.2856.51parity
MIRACL (multilingual)57.7069.32incumbent wins

The gap is genuine, independent, and same-pipeline. Worth noting: two of the nine research agents reported that this table did not exist — they had read an earlier revision of the paper. Pinning the revision caught it.

And then the benchmarks were demoted anyway

MTEB-Code and CoIR contain zero named C#. No model in this field has any published C#-specific evaluation. Qwen was fine-tuned on CodeSearchNet; the incumbent had no code training at all. CoREB (2026) separately found no off-the-shelf reranker is net-positive across code tasks.

So published benchmarks can pick candidates. They cannot decide a winner. Every option is equally unproven on C#. That is the entire reason a hand-labelled golden set is mandatory rather than nice-to-have.

Two calls that went against the obvious answer

Staging was inverted — embedder first, reranker second

The instinct is to swap the reranker first, since its benchmark gap is larger. Rejected: the reranker is the risky integration — chat-template handling plus yes/no logit extraction, and both vLLM and llama.cpp had shipped silent-wrong-score bugs. The embedder is deterministic, parity-testable, and a drop-in dimension. Do the verifiable one first.

A cost claim was corrected by hand

One research pass asserted the new reranker cost "+35% FLOPs." Working the embedding-table math directly — the incumbent's XLM-R-large table is ~250,002 × 1024 ≈ 256M of its 568M parameters and carries no FLOPs, leaving ~312M non-embedding, against Qwen3-0.6B's ~440M — gives ≈1.41×. The full-vocab lm_head at only the last position is ~0.3 GFLOP, negligible. "+35%" was true only for a naive all-position export.

Speed decided as much as accuracy

MeasurementValue
Warm query, end to end918 ms
Cold start — scaled-to-zero endpoint + auto-paused SQL116,466 ms (1 m 56 s)
Endpoint warm via heartbeat1 s
Full reindex — 254 files / 605 chunks, staging + atomic swap1 m 13 s

Cold start, not accuracy, was the real product risk. A visitor arriving after an idle period gets a first query that can take two minutes. Someone who assumes it's broken and leaves at second 20 never sees the retrieval quality at all. That produced the warm-up retry, the heartbeat, and honest "waking up the server…" messaging instead of a spinner that lies.

Speed also drove the hardware. Sized then for the 4B, which fits a cheaper L4, A10G was chosen for ~2× the memory bandwidth and better latency. Always-on serving was costed at ~$1,150/month; scale-to-zero brings real demo usage to a few dollars a day.

Measured on this corpus

Everything above this point is published benchmarks plus one live prompt-format test. This section is different: it is this system, this corpus, measured in September 2026. Both experiments are reported, including the one that was wrong.

A question set built to discriminate, not to demo

The original six demo questions each contained a word from their own answer's filename, so plain keyword matching could win them outright and no embedder was ever really under test. The replacement set has fourteen questions with zero token overlap between question and answer filename, verified per question at run time, plus four exact class names as a lexical control and two unanswerable questions to check abstention.

Experiment 1: the chunk header. Wrong.

Every indexed chunk carried a context header naming the file, namespace, class and fields. Measured across 605 chunks it was 33.3% of all embedded text, and 36% of chunks carried it twice. The theory was that this collapsed a file's chunks toward one vector describing the header rather than the code. It was stripped, the corpus fully reindexed, and the identical battery rerun.

Semantic setrank 1top 5MRR
Header left in place3 / 146 / 140.304
Header stripped3 / 144 / 140.238

It got worse, and not one of the six questions that both stacks missed moved. The header was not inert noise: // Class: RegisterModel was the only occurrence of "register" in that chunk, so Register.cshtml.cs fell from rank 1 to absent. Class and file names are real signal for natural-language questions. The change was reverted.

Experiment 2: is a code-specialized embedder worth it?

The comparison that had never been run. Local e5-large-v2 (general-purpose, 1024-dim, ONNX on CPU) against the hosted code embedder (1536-dim, GPU). Both corpora verified identical first: 605 chunks, 605 vectors, 351 carrying the header, on each side. The measurement is the exact cosine rank of the correct file across all 605 chunks, with no index approximation, no fusion, no reranker and no score floor, so nothing but the embedder is being compared.

Question’s correct answere5-large-v2code embedderBetter
TransferBasket.cs9315code, by 78
CatalogFilterPaginatedSpecification.cs6017code, by 43
CacheHelpers.cs5633code, by 23
HomePageHealthCheck.cs26261code, by 201
GetMyOrdersHandler.cs2499e5, by 75
Logout.cshtml.cs5995e5, by 36
EmptyBasketOnCheckoutException.cs (control)1124e5, by 123
Register.cshtml.cs (control)251e5, by 49

The code-specialized embedder wins four of the six hard questions, and on one of them the margin is large. It still does not make them retrievable. Its best rank anywhere in that set is 15 of 605, so every one of these questions still misses a top-5 result. And it is substantially worse on the two questions the general-purpose model answers at rank 1 and rank 2.

The honest conclusion is that neither embedder can answer these six, that the two are good at different questions rather than one being better, and that the remaining gap is not something a better embedding model fixes. Chunking strategy is the open candidate.

Experiment 3: which reranker, on identical input

The experiment that should have come first. A cross-encoder is a pure function of (query, documents), so handing two of them the same candidate list removes the embedder, the index, the fusion weights and the score floors from the comparison entirely. Nothing is left varying but the model. Shortlists came from one retrieval call per question at the same pool depth the server reranks at.

Question’s correct answerfirst stagebge-v2-m3Qwen3-0.6BBetter
EmptyBasketOnCheckoutException.cs811tie
ToastComponent.cs441Qwen3
ImageValidators.cs621Qwen3
CatalogContextSeed.cs111tie
ExceptionMiddleware.cs551Qwen3
IdentityTokenClaimService.cs1825bge
GetMyOrdersHandler.cs2774Qwen3
CatalogFilterPaginatedSpecification.cs2894Qwen3
Register.cshtml.cs213bge
OrderBuilder.cs (identifier)231Qwen3
ExceptionMiddleware.cs (identifier)111tie
CacheHelpers.cs (identifier)111tie
BasketQueryService.cs (identifier)111tie
0.452MRR, retrieval alone
+37%bge-reranker-v2-m3
+71%Qwen3-Reranker-0.6B

Qwen3-Reranker-0.6B takes six questions to bge’s two, with five ties, and lifts MRR from 0.452 to 0.772 against bge’s 0.618. Same shortlists, same corpus, same questions.

This is where the retrieval quality on this corpus actually lives. Changing the embedding model moved nothing that mattered: no question became retrievable and both controls got substantially worse. Changing the reranker, on byte-identical input, moves MRR by 71%.

The clearest cases are the ones the first stage nearly lost. GetMyOrdersHandler.cs came out of retrieval at rank 27 and CatalogFilterPaginatedSpecification.cs at 28, both effectively invisible to a user, and Qwen3 pulled them to 4. ExceptionMiddleware.cs sat at 5 and went to 1, where bge left it at 5 and never moved it at all.

bge is not useless. It still adds 37% over raw retrieval and wins two questions outright, including rescuing IdentityTokenClaimService.cs from 18 to 2 where Qwen3 only managed 5. Across the set it is simply the weaker model. And this settles the question the correction at the top of this page raises: the 0.6B was deployed over the 4B for latency, but on measured retrieval quality it also beats the reranker this project previously shipped.

Latency is deliberately not scored here. The two models ran on different hardware, roughly 250 ms per query against roughly 140 s, which measures an A10G against a laptop CPU and says nothing about either model.

Experiment 4: does the embedder earn its place at all?

A cross-encoder never reads a vector. It reads the query and the document text together, so what it needs is a candidate list, not embeddings. Full-text search can produce one. That makes “full-text, then rerank” a complete architecture with no embedding model, no vector column, no index, and no reindex when models change. Given Experiment 3, that deserved a measurement rather than an assumption.

Same reranker on both sides. The only difference is how the shortlist was built.

First stageAnswer in shortlistrank 1top 5MRR
hybrid (vector + full-text)13 of 189130.557
full-text only, no embedder10 of 189100.519

The embedder contributes recall, and nothing else. Where full-text finds the answer but buries it, the reranker erases the difference completely:

Expected answerhybrid first stagefull-text first stagefinal rank, both
EmptyBasketOnCheckoutException.cs8191
ImageValidators.cs6251
ExceptionMiddleware.cs5221
Register.cshtml.cs2273

Rank 25 out of retrieval, rank 1 after reranking, identical to the hybrid run that started at 6. So the embedder is not improving the ordering of anything once a strong cross-encoder is in play. What it does is put three answers in front of the reranker that full-text never surfaces at all: IdentityTokenClaimService.cs, GetMyOrdersHandler.cs and CatalogFilterPaginatedSpecification.cs. Those are pure recall, and no reranker recovers a document it was never shown.

The honest trade, then: the embedding endpoint buys about 17% more answerable questions and zero ranking improvement. That is a real gain and a real cost, and which way it goes depends on the corpus and the budget rather than on anything a benchmark table can tell you.

Re-tested October 2026: four rerankers, and a near miss

The choice above was made on published benchmarks plus a latency judgement. In October 2026 it was re-tested on this corpus against three alternatives, with the retrieval, fusion weights, score floors and embedding model all held constant so that only the reranker varied. Twenty questions came from this demo's own usage log and four were natural negative controls the corpus cannot answer. The same twenty real questions were then labelled by hand at file level, 59 graded labels, to give nDCG, recall and MRR.

MetricQwen3-Reranker-0.6B
deployed
mxbai-rerank-base-v2ettin-reranker-400mettin-reranker-150m
Real questions, top-10.67510.96900.99680.9971
Negative controls (lower is better)0.05930.96070.78970.7906
Separation0.61580.00830.20710.2065
Labelled nDCG@50.59420.36120.65620.6051
Labelled recall@50.52250.35580.57420.5283
Labelled MRR0.81670.50000.85830.7750
Results surviving the score floors7.723.825.025.0
Parameters0.6B0.49B0.39B0.15B

On ranking, the best challenger is inside noise. None of them can decline

ettin-reranker-400m posts the highest labelled numbers: nDCG@5 0.6562 against 0.5942, MRR 0.8583 against 0.8167. That gap is not established. Paired question by question, ettin wins 7, the deployed model wins 5, and 8 tie. The 95% bootstrap interval on the nDCG difference runs from −0.03 to +0.17, so it includes zero. mxbai-rerank-base-v2 is clearly last.

An earlier version of this section scored the same runs against six template queries, where ettin read 0.6564 against 0.5412 and was reported as better. Real labels removed most of that gap.

The negative controls are where it falls apart. ettin-reranker-150m scores "how many bits in a byte" at 0.79 against a corpus of C# e-commerce code that contains no answer to it. mxbai-rerank-base-v2 is worse still, with a separation of 0.0083: it scores the unanswerable queries almost exactly as highly as the real ones, which makes it uninformative as a relevance signal here.

The practical consequence is the last row. All three challengers keep 24 or 25 of 25 candidates after the score floors, because the relevance floor is a fraction of the top score and nothing is filtered when everything scores above 0.95. Every query, including the ones with no answer, would return a full page. The deployed model keeps 7.7.

Why the deployed model stays

Ranking is half the job. The other half is representing absence, and a search box that answers everything confidently is worse than one that returns nothing. Qwen3-Reranker is trained for an instruction-conditioned binary judgement, which is what produces usable calibration. The ettin models are MS MARCO passage rerankers that take no instruction at all, and mxbai is a general multilingual reranker. None of them were built to represent absence.

This page previously said a better-ranking model exists and only calibration stopped its adoption. The labelled set does not show that. Re-deriving the score floors around ettin is only worth trying if a larger labelled set shows a real ranking gain.

Two of the three challengers nearly produced fabricated results.

The ettin models load through the standard sequence-classification path with a randomly initialised classification head. The loader reports classifier.weight, classifier.bias and head.dense.weight as missing, then returns scores anyway. On an easy pair the random head still ranked the relevant document first, by chance. Reading the load report rather than the output is what caught it.

The mxbai package fails outright on current transformers, and its plain sequence-classification path has no head either, so its scoring was reimplemented from the package source: a sigmoid over logit("1") − logit("0") at the final position.

What this re-test cannot tell you

The reranker is also the judge for the unlabelled questions, which is circular when the thing being compared is a reranker. A model that simply emits high numbers wins the relevance row for free, which is exactly what happened. The same inflation applies to the negative controls, so separation stays honest where the raw score does not, and the labelled queries are independent of all of it. The conclusion rests on those two and not on the relevance row.

Twenty labelled questions and four negative controls is a small set. Differences under about 0.1 nDCG are inside noise at this size. The separation result is not subtle. The ranking result between the top three is.

An earlier version of this section carried figures measured against an index built with a tokenization bug. They have been replaced with figures from the corrected index.

What this evidence does not support

The model selection is defensible on independent benchmarks plus a decisive live prompt-format measurement. A hand-labelled set of 20 real questions now exists and is reported above. It is enough to identify the clear loser and too small to split the close calls.