Correction, September 2026. This page recorded
Qwen3-Reranker-4B as the deployed reranker. It is not what runs. The live endpoint
serves Qwen3-Reranker-0.6B-seq-cls, confirmed from the endpoint's own
/v1/models response and the application configuration.
The swap was deliberate, not accidental. On this workload the 4B scored roughly 4% better and ran about 125% slower. That is not a trade worth making for an interactive search box. The page simply never caught up with the decision, and the endpoint is still named for the 4B, which is how it went unnoticed.
Two experiments have since been run against this corpus. Both are reported below, including the one that failed. See Measured on this corpus.
This page shows how the embedding and reranking models running this demo were selected, what was measured, and — the part most write-ups skip — what the evidence does not support. Every number below is either measured against the live endpoints or a published benchmark with its source named. The two are never mixed.
The finding worth three orders of magnitude
Qwen3-Reranker scores through its chat template. Testing that wrapping applied client-side against posting raw query/document text — same model, same endpoint, same documents:
Both calls hit the same live endpoint, which at the time of this July 2026 test served Qwen3-Reranker-4B-seq-cls, 329 prompt tokens,
three documents and one query:
Instruct: Given a code search query, retrieve relevant code chunks that answer the query.
Query: how is user authentication handled
| Document | Chat-template wrapped | Raw |
|---|---|---|
PasswordSignInAsync(...) — the answer | 0.1618865728 | 0.375985 |
LogTimestamp(...) — logging helper | 0.0000346425 | 0.105656 |
CalculateBasketTotal(...) — unrelated | 0.0000053275 | 0.113646 |
| Separation, top vs next | 4,673× | 3.31× |
The takeaway isn't the model — and it isn't that raw "fails." Raw still ranked the authentication method first. It got the answer right. What it lost was discrimination: it scored a basket-total calculation and a logging helper within 8% of each other, and both within a factor of 3.5 of the correct answer.
Wrapped, the correct answer sits 4,673× above the next candidate and 30,388× above the worst. That margin is what makes a score threshold possible at all. Three orders of magnitude of usable signal lived in the prompt formatting, not the weights.
What the demo showed, and why it was not a test
Correction. This section previously said there was “no keyword overlap
engineered in.” That was wrong. Every one of these six questions contains a word that appears
in its own answer’s filename — send email against
EmailSender.cs, the basket total against
Basket.cs. Plain keyword matching can win them with no embedding model involved, so
they show the system works end to end but prove nothing about retrieval quality. Noticing that is
what prompted the experiments below and a question set built to
discriminate instead.
Six plain-English questions fired at the live index — 254 files, 605 chunks, health green. Kept here because they are what the demo actually showed, including the one it got wrong.
| Question | Top results |
|---|---|
| how does the app send email | EmailSender.cs · Program.cs · OrderController.cs |
| how is the basket total calculated | Basket.cs · BasketViewModel.cs · Order.cs · OrderTotal.cs |
| how are users authenticated at sign in | AuthenticateEndpoint.cs · ManageController.cs · CustomAuthStateProvider.cs |
| how are catalog items filtered by brand and type | CatalogItemService.cs · CatalogViewModelService.cs · CatalogBrand.cs |
| where is the database connection configured | AppIdentityDbContextSeed.cs · IdentityHostingStartup.cs |
| what happens when a user checks out | ConfigureCookieSettings.cs · ApiHealthCheck.cs — miss |
Five hit, one missed — and the miss is the useful one. "What happens when a user checks out" returned cookie settings and health checks: nothing to do with checkout. Rewording found the boundary:
| Reworded | Top results |
|---|---|
| create an order from the basket | Order.cs · Basket.cs · OrderService.cs |
| how is an order placed | Order.cs · GetMyOrdersHandler.cs · GetOrderDetailsHandler.cs |
| checkout controller action | BaseApiController.cs — still weak |
The corpus has no "checkout" concept — it models the same idea as orders. The retrieval wasn't broken; the vocabulary was. Naming the domain concept fixes it instantly, and no amount of reranking would have.
The keyword leg, tested separately
Exact identifiers should not need semantics, and they don't:
EmailSender → EmailSender.cs, IEmailSender.cs, EmailSenderExtensions.cs ·
OrderService → OrderService.cs, IOrderService.cs ·
IRepository → IRepository.cs.
PaymentMethods returned exactly one result — Buyer.cs — rather than
padding the list with near-misses.
Six rerankers evaluated
| Model | Params | Code benchmark | Outcome |
|---|---|---|---|
| Qwen3-Reranker-4B | 4B | MTEB-Code 81.20 | chosen on paper; ~4% better, ~125% slower, so not deployed |
| Qwen3-Reranker-8B | 8B | MTEB-Code 81.22 | +0.02 for ~2× cost — rejected |
| Qwen3-Reranker-0.6B | 0.6B | CoIR 65.18 | deployed: the 4B's quality edge did not pay for its latency |
| mxbai-rerank-base-v2 | 0.5B | CoIR 65.71 | ≈ tie, kept as A/B arm |
| bge-reranker-v2-m3 (incumbent) | 0.57B | CoIR 36.28 | replaced |
| jina-reranker-v3 / jina-code | — | strong | CC-BY-NC — license-blocked |
| zerank-1-small / zerank-2 | 1.7B / 4B | nDCG@10 0.6528 | CC-BY-NC or ~3× cost |
The incumbent was not bad — it remains strong on multilingual text (MIRACL 69.32, best in its table). It was bad specifically on code, at roughly half the achievable nDCG.
Why the headline benchmark was thrown out
The widely-quoted figure for this comparison is MTEB-Code 73.42 vs 41.38. It did not survive checking: that number is the vendor scoring its own model on its own retrieved candidates. Not independent.
So the claim was verified from a primary source instead — arXiv 2509.25085 v3, Table 2, published by a rival lab:
| Benchmark | Qwen3-Reranker-0.6B | bge-reranker-v2-m3 | Read |
|---|---|---|---|
| CoIR (code) | 65.18 | 36.28 | ~29 pt code gap — real |
| BEIR (general) | 56.28 | 56.51 | parity |
| MIRACL (multilingual) | 57.70 | 69.32 | incumbent wins |
The gap is genuine, independent, and same-pipeline. Worth noting: two of the nine research agents reported that this table did not exist — they had read an earlier revision of the paper. Pinning the revision caught it.
And then the benchmarks were demoted anyway
MTEB-Code and CoIR contain zero named C#. No model in this field has any published C#-specific evaluation. Qwen was fine-tuned on CodeSearchNet; the incumbent had no code training at all. CoREB (2026) separately found no off-the-shelf reranker is net-positive across code tasks.
So published benchmarks can pick candidates. They cannot decide a winner. Every option is equally unproven on C#. That is the entire reason a hand-labelled golden set is mandatory rather than nice-to-have.
Two calls that went against the obvious answer
Staging was inverted — embedder first, reranker second
The instinct is to swap the reranker first, since its benchmark gap is larger. Rejected: the reranker is the risky integration — chat-template handling plus yes/no logit extraction, and both vLLM and llama.cpp had shipped silent-wrong-score bugs. The embedder is deterministic, parity-testable, and a drop-in dimension. Do the verifiable one first.
A cost claim was corrected by hand
One research pass asserted the new reranker cost "+35% FLOPs." Working the embedding-table math
directly — the incumbent's XLM-R-large table is ~250,002 × 1024 ≈ 256M of its
568M parameters and carries no FLOPs, leaving ~312M non-embedding, against Qwen3-0.6B's ~440M — gives
≈1.41×. The full-vocab lm_head at only the last position is
~0.3 GFLOP, negligible. "+35%" was true only for a naive all-position export.
Speed decided as much as accuracy
| Measurement | Value |
|---|---|
| Warm query, end to end | 918 ms |
| Cold start — scaled-to-zero endpoint + auto-paused SQL | 116,466 ms (1 m 56 s) |
| Endpoint warm via heartbeat | 1 s |
| Full reindex — 254 files / 605 chunks, staging + atomic swap | 1 m 13 s |
Cold start, not accuracy, was the real product risk. A visitor arriving after an idle period gets a first query that can take two minutes. Someone who assumes it's broken and leaves at second 20 never sees the retrieval quality at all. That produced the warm-up retry, the heartbeat, and honest "waking up the server…" messaging instead of a spinner that lies.
Speed also drove the hardware. Sized then for the 4B, which fits a cheaper L4, A10G was chosen for ~2× the memory bandwidth and better latency. Always-on serving was costed at ~$1,150/month; scale-to-zero brings real demo usage to a few dollars a day.
Measured on this corpus
Everything above this point is published benchmarks plus one live prompt-format test. This section is different: it is this system, this corpus, measured in September 2026. Both experiments are reported, including the one that was wrong.
A question set built to discriminate, not to demo
The original six demo questions each contained a word from their own answer's filename, so plain keyword matching could win them outright and no embedder was ever really under test. The replacement set has fourteen questions with zero token overlap between question and answer filename, verified per question at run time, plus four exact class names as a lexical control and two unanswerable questions to check abstention.
Experiment 1: the chunk header. Wrong.
Every indexed chunk carried a context header naming the file, namespace, class and fields. Measured across 605 chunks it was 33.3% of all embedded text, and 36% of chunks carried it twice. The theory was that this collapsed a file's chunks toward one vector describing the header rather than the code. It was stripped, the corpus fully reindexed, and the identical battery rerun.
| Semantic set | rank 1 | top 5 | MRR |
|---|---|---|---|
| Header left in place | 3 / 14 | 6 / 14 | 0.304 |
| Header stripped | 3 / 14 | 4 / 14 | 0.238 |
It got worse, and not one of the six questions that both stacks missed moved. The header was not inert
noise: // Class: RegisterModel was the only occurrence of "register" in that chunk, so
Register.cshtml.cs fell from rank 1 to absent. Class and file names are real signal for
natural-language questions. The change was reverted.
Experiment 2: is a code-specialized embedder worth it?
The comparison that had never been run. Local e5-large-v2 (general-purpose, 1024-dim,
ONNX on CPU) against the hosted code embedder (1536-dim, GPU). Both corpora verified identical first:
605 chunks, 605 vectors, 351 carrying the header, on each side. The measurement is the exact cosine
rank of the correct file across all 605 chunks, with no index approximation, no fusion, no reranker
and no score floor, so nothing but the embedder is being compared.
| Question’s correct answer | e5-large-v2 | code embedder | Better |
|---|---|---|---|
| TransferBasket.cs | 93 | 15 | code, by 78 |
| CatalogFilterPaginatedSpecification.cs | 60 | 17 | code, by 43 |
| CacheHelpers.cs | 56 | 33 | code, by 23 |
| HomePageHealthCheck.cs | 262 | 61 | code, by 201 |
| GetMyOrdersHandler.cs | 24 | 99 | e5, by 75 |
| Logout.cshtml.cs | 59 | 95 | e5, by 36 |
| EmptyBasketOnCheckoutException.cs (control) | 1 | 124 | e5, by 123 |
| Register.cshtml.cs (control) | 2 | 51 | e5, by 49 |
The code-specialized embedder wins four of the six hard questions, and on one of them the margin is large. It still does not make them retrievable. Its best rank anywhere in that set is 15 of 605, so every one of these questions still misses a top-5 result. And it is substantially worse on the two questions the general-purpose model answers at rank 1 and rank 2.
The honest conclusion is that neither embedder can answer these six, that the two are good at different questions rather than one being better, and that the remaining gap is not something a better embedding model fixes. Chunking strategy is the open candidate.
Experiment 3: which reranker, on identical input
The experiment that should have come first. A cross-encoder is a pure function of (query, documents), so handing two of them the same candidate list removes the embedder, the index, the fusion weights and the score floors from the comparison entirely. Nothing is left varying but the model. Shortlists came from one retrieval call per question at the same pool depth the server reranks at.
| Question’s correct answer | first stage | bge-v2-m3 | Qwen3-0.6B | Better |
|---|---|---|---|---|
| EmptyBasketOnCheckoutException.cs | 8 | 1 | 1 | tie |
| ToastComponent.cs | 4 | 4 | 1 | Qwen3 |
| ImageValidators.cs | 6 | 2 | 1 | Qwen3 |
| CatalogContextSeed.cs | 1 | 1 | 1 | tie |
| ExceptionMiddleware.cs | 5 | 5 | 1 | Qwen3 |
| IdentityTokenClaimService.cs | 18 | 2 | 5 | bge |
| GetMyOrdersHandler.cs | 27 | 7 | 4 | Qwen3 |
| CatalogFilterPaginatedSpecification.cs | 28 | 9 | 4 | Qwen3 |
| Register.cshtml.cs | 2 | 1 | 3 | bge |
| OrderBuilder.cs (identifier) | 2 | 3 | 1 | Qwen3 |
| ExceptionMiddleware.cs (identifier) | 1 | 1 | 1 | tie |
| CacheHelpers.cs (identifier) | 1 | 1 | 1 | tie |
| BasketQueryService.cs (identifier) | 1 | 1 | 1 | tie |
Qwen3-Reranker-0.6B takes six questions to bge’s two, with five ties, and lifts MRR from 0.452 to 0.772 against bge’s 0.618. Same shortlists, same corpus, same questions.
This is where the retrieval quality on this corpus actually lives. Changing the embedding model moved nothing that mattered: no question became retrievable and both controls got substantially worse. Changing the reranker, on byte-identical input, moves MRR by 71%.
The clearest cases are the ones the first stage nearly lost. GetMyOrdersHandler.cs came out of retrieval at rank 27 and CatalogFilterPaginatedSpecification.cs at 28, both effectively invisible to a user, and Qwen3 pulled them to 4. ExceptionMiddleware.cs sat at 5 and went to 1, where bge left it at 5 and never moved it at all.
bge is not useless. It still adds 37% over raw retrieval and wins two questions outright, including rescuing IdentityTokenClaimService.cs from 18 to 2 where Qwen3 only managed 5. Across the set it is simply the weaker model. And this settles the question the correction at the top of this page raises: the 0.6B was deployed over the 4B for latency, but on measured retrieval quality it also beats the reranker this project previously shipped.
Latency is deliberately not scored here. The two models ran on different hardware, roughly 250 ms per query against roughly 140 s, which measures an A10G against a laptop CPU and says nothing about either model.
Experiment 4: does the embedder earn its place at all?
A cross-encoder never reads a vector. It reads the query and the document text together, so what it needs is a candidate list, not embeddings. Full-text search can produce one. That makes “full-text, then rerank” a complete architecture with no embedding model, no vector column, no index, and no reindex when models change. Given Experiment 3, that deserved a measurement rather than an assumption.
Same reranker on both sides. The only difference is how the shortlist was built.
| First stage | Answer in shortlist | rank 1 | top 5 | MRR |
|---|---|---|---|---|
| hybrid (vector + full-text) | 13 of 18 | 9 | 13 | 0.557 |
| full-text only, no embedder | 10 of 18 | 9 | 10 | 0.519 |
The embedder contributes recall, and nothing else. Where full-text finds the answer but buries it, the reranker erases the difference completely:
| Expected answer | hybrid first stage | full-text first stage | final rank, both |
|---|---|---|---|
| EmptyBasketOnCheckoutException.cs | 8 | 19 | 1 |
| ImageValidators.cs | 6 | 25 | 1 |
| ExceptionMiddleware.cs | 5 | 22 | 1 |
| Register.cshtml.cs | 2 | 27 | 3 |
Rank 25 out of retrieval, rank 1 after reranking, identical to the hybrid run that started at 6. So the embedder is not improving the ordering of anything once a strong cross-encoder is in play. What it does is put three answers in front of the reranker that full-text never surfaces at all: IdentityTokenClaimService.cs, GetMyOrdersHandler.cs and CatalogFilterPaginatedSpecification.cs. Those are pure recall, and no reranker recovers a document it was never shown.
The honest trade, then: the embedding endpoint buys about 17% more answerable questions and zero ranking improvement. That is a real gain and a real cost, and which way it goes depends on the corpus and the budget rather than on anything a benchmark table can tell you.
What this evidence does not support
- There is no nDCG or recall@k measured on this corpus. The benchmark figures are published results on public datasets; the 4,600× result is a score-separation test on a hand-built three-document probe.
- The six rerankers were never run head-to-head on the same query. That table is published benchmarks, not a bake-off on this corpus. The live measurement above tested one model against itself in two prompt formats — it says nothing about how the other five would have scored.
- There are no per-model latency numbers. The 918 ms is whole-system, end to end. Nothing here isolates reranker time from embedding, SQL, or fusion.
- The same is true of the embedders. The retrieval battery above ran after the embedding swap, so it reflects one model only. The previous embedder was never run against those six questions. This one is runnable, though — both embedder paths ship in the codebase (local ONNX e5-large-v2 at 1024-dim, hosted bge-code-v1 at 1536), so it is two indexes and two harness runs, not a dead end. It is an unrun experiment, not an impossible one.
- The model-vs-model A/B now exists, on 18 questions rather than a full golden set. Both an embedder and a reranker comparison have been run (Experiments 2 and 3). A larger hand-labelled golden set would tighten the confidence interval; 13 scored questions is enough to separate a 71% MRR gain from a 37% one, and not enough to split hairs.
- Comparing embedders is not a sweep. Each embedding model requires a full reindex, because the query vector must match the stored vector geometry.
The model selection is defensible on independent benchmarks plus a decisive live prompt-format measurement. The end-to-end retrieval-quality number is the one thing still missing, and it is gated on labelling that golden set. Stating that plainly is more useful than a chart that implies otherwise.