How the Models Were Chosen

Retrieval-quality engineering behind this demo · evidence, and its limits

Benchmarks ← Back to search

Correction, September 2026. This page recorded Qwen3-Reranker-4B as the deployed reranker. It is not what runs. The live endpoint serves Qwen3-Reranker-0.6B-seq-cls, confirmed from the endpoint's own /v1/models response and the application configuration.

The swap was deliberate, not accidental. On this workload the 4B scored roughly 4% better and ran about 125% slower. That is not a trade worth making for an interactive search box. The page simply never caught up with the decision, and the endpoint is still named for the 4B, which is how it went unnoticed.

Two experiments have since been run against this corpus. Both are reported below, including the one that failed. See Measured on this corpus.

This page shows how the embedding and reranking models running this demo were selected, what was measured, and — the part most write-ups skip — what the evidence does not support. Every number below is either measured against the live endpoints or a published benchmark with its source named. The two are never mixed.

The finding worth three orders of magnitude

Qwen3-Reranker scores through its chat template. Testing that wrapping applied client-side against posting raw query/document text — same model, same endpoint, same documents:

4,600×relevant vs irrelevant separation, chat-template wrapping
3.3×the same test, raw prompt
918 mswarm query, end to end

Both calls hit the same live endpoint, which at the time of this July 2026 test served Qwen3-Reranker-4B-seq-cls, 329 prompt tokens, three documents and one query:

Instruct: Given a code search query, retrieve relevant code chunks that answer the query.

Query: how is user authentication handled

DocumentChat-template wrappedRaw
PasswordSignInAsync(...) — the answer0.16188657280.375985
LogTimestamp(...) — logging helper0.00003464250.105656
CalculateBasketTotal(...) — unrelated0.00000532750.113646
Separation, top vs next4,673×3.31×

The takeaway isn't the model — and it isn't that raw "fails." Raw still ranked the authentication method first. It got the answer right. What it lost was discrimination: it scored a basket-total calculation and a logging helper within 8% of each other, and both within a factor of 3.5 of the correct answer.

Wrapped, the correct answer sits 4,673× above the next candidate and 30,388× above the worst. That margin is what makes a score threshold possible at all. Three orders of magnitude of usable signal lived in the prompt formatting, not the weights.

What the demo showed, and why it was not a test

Correction. This section previously said there was “no keyword overlap engineered in.” That was wrong. Every one of these six questions contains a word that appears in its own answer’s filename — send email against EmailSender.cs, the basket total against Basket.cs. Plain keyword matching can win them with no embedding model involved, so they show the system works end to end but prove nothing about retrieval quality. Noticing that is what prompted the experiments below and a question set built to discriminate instead.

Six plain-English questions fired at the live index — 254 files, 605 chunks, health green. Kept here because they are what the demo actually showed, including the one it got wrong.

QuestionTop results
how does the app send emailEmailSender.cs · Program.cs · OrderController.cs
how is the basket total calculatedBasket.cs · BasketViewModel.cs · Order.cs · OrderTotal.cs
how are users authenticated at sign inAuthenticateEndpoint.cs · ManageController.cs · CustomAuthStateProvider.cs
how are catalog items filtered by brand and typeCatalogItemService.cs · CatalogViewModelService.cs · CatalogBrand.cs
where is the database connection configuredAppIdentityDbContextSeed.cs · IdentityHostingStartup.cs
what happens when a user checks outConfigureCookieSettings.cs · ApiHealthCheck.cs — miss

Five hit, one missed — and the miss is the useful one. "What happens when a user checks out" returned cookie settings and health checks: nothing to do with checkout. Rewording found the boundary:

RewordedTop results
create an order from the basketOrder.cs · Basket.cs · OrderService.cs
how is an order placedOrder.cs · GetMyOrdersHandler.cs · GetOrderDetailsHandler.cs
checkout controller actionBaseApiController.cs — still weak

The corpus has no "checkout" concept — it models the same idea as orders. The retrieval wasn't broken; the vocabulary was. Naming the domain concept fixes it instantly, and no amount of reranking would have.

The keyword leg, tested separately

Exact identifiers should not need semantics, and they don't: EmailSender → EmailSender.cs, IEmailSender.cs, EmailSenderExtensions.cs · OrderService → OrderService.cs, IOrderService.cs · IRepository → IRepository.cs. PaymentMethods returned exactly one result — Buyer.cs — rather than padding the list with near-misses.

Six rerankers evaluated

ModelParamsCode benchmarkOutcome
Qwen3-Reranker-4B4BMTEB-Code 81.20chosen on paper; ~4% better, ~125% slower, so not deployed
Qwen3-Reranker-8B8BMTEB-Code 81.22+0.02 for ~2× cost — rejected
Qwen3-Reranker-0.6B0.6BCoIR 65.18deployed: the 4B's quality edge did not pay for its latency
mxbai-rerank-base-v20.5BCoIR 65.71≈ tie, kept as A/B arm
bge-reranker-v2-m3 (incumbent)0.57BCoIR 36.28replaced
jina-reranker-v3 / jina-code—strongCC-BY-NC — license-blocked
zerank-1-small / zerank-21.7B / 4BnDCG@10 0.6528CC-BY-NC or ~3× cost

The incumbent was not bad — it remains strong on multilingual text (MIRACL 69.32, best in its table). It was bad specifically on code, at roughly half the achievable nDCG.

Why the headline benchmark was thrown out

The widely-quoted figure for this comparison is MTEB-Code 73.42 vs 41.38. It did not survive checking: that number is the vendor scoring its own model on its own retrieved candidates. Not independent.

So the claim was verified from a primary source instead — arXiv 2509.25085 v3, Table 2, published by a rival lab:

BenchmarkQwen3-Reranker-0.6Bbge-reranker-v2-m3Read
CoIR (code)65.1836.28~29 pt code gap — real
BEIR (general)56.2856.51parity
MIRACL (multilingual)57.7069.32incumbent wins

The gap is genuine, independent, and same-pipeline. Worth noting: two of the nine research agents reported that this table did not exist — they had read an earlier revision of the paper. Pinning the revision caught it.

And then the benchmarks were demoted anyway

MTEB-Code and CoIR contain zero named C#. No model in this field has any published C#-specific evaluation. Qwen was fine-tuned on CodeSearchNet; the incumbent had no code training at all. CoREB (2026) separately found no off-the-shelf reranker is net-positive across code tasks.

So published benchmarks can pick candidates. They cannot decide a winner. Every option is equally unproven on C#. That is the entire reason a hand-labelled golden set is mandatory rather than nice-to-have.

Two calls that went against the obvious answer

Staging was inverted — embedder first, reranker second

The instinct is to swap the reranker first, since its benchmark gap is larger. Rejected: the reranker is the risky integration — chat-template handling plus yes/no logit extraction, and both vLLM and llama.cpp had shipped silent-wrong-score bugs. The embedder is deterministic, parity-testable, and a drop-in dimension. Do the verifiable one first.

A cost claim was corrected by hand

One research pass asserted the new reranker cost "+35% FLOPs." Working the embedding-table math directly — the incumbent's XLM-R-large table is ~250,002 × 1024 ≈ 256M of its 568M parameters and carries no FLOPs, leaving ~312M non-embedding, against Qwen3-0.6B's ~440M — gives ≈1.41×. The full-vocab lm_head at only the last position is ~0.3 GFLOP, negligible. "+35%" was true only for a naive all-position export.

Speed decided as much as accuracy

MeasurementValue
Warm query, end to end918 ms
Cold start — scaled-to-zero endpoint + auto-paused SQL116,466 ms (1 m 56 s)
Endpoint warm via heartbeat1 s
Full reindex — 254 files / 605 chunks, staging + atomic swap1 m 13 s

Cold start, not accuracy, was the real product risk. A visitor arriving after an idle period gets a first query that can take two minutes. Someone who assumes it's broken and leaves at second 20 never sees the retrieval quality at all. That produced the warm-up retry, the heartbeat, and honest "waking up the server…" messaging instead of a spinner that lies.

Speed also drove the hardware. Sized then for the 4B, which fits a cheaper L4, A10G was chosen for ~2× the memory bandwidth and better latency. Always-on serving was costed at ~$1,150/month; scale-to-zero brings real demo usage to a few dollars a day.

Measured on this corpus

Everything above this point is published benchmarks plus one live prompt-format test. This section is different: it is this system, this corpus, measured in September 2026. Both experiments are reported, including the one that was wrong.

A question set built to discriminate, not to demo

The original six demo questions each contained a word from their own answer's filename, so plain keyword matching could win them outright and no embedder was ever really under test. The replacement set has fourteen questions with zero token overlap between question and answer filename, verified per question at run time, plus four exact class names as a lexical control and two unanswerable questions to check abstention.

Experiment 1: the chunk header. Wrong.

Every indexed chunk carried a context header naming the file, namespace, class and fields. Measured across 605 chunks it was 33.3% of all embedded text, and 36% of chunks carried it twice. The theory was that this collapsed a file's chunks toward one vector describing the header rather than the code. It was stripped, the corpus fully reindexed, and the identical battery rerun.

Semantic setrank 1top 5MRR
Header left in place3 / 146 / 140.304
Header stripped3 / 144 / 140.238

It got worse, and not one of the six questions that both stacks missed moved. The header was not inert noise: // Class: RegisterModel was the only occurrence of "register" in that chunk, so Register.cshtml.cs fell from rank 1 to absent. Class and file names are real signal for natural-language questions. The change was reverted.

Experiment 2: is a code-specialized embedder worth it?

The comparison that had never been run. Local e5-large-v2 (general-purpose, 1024-dim, ONNX on CPU) against the hosted code embedder (1536-dim, GPU). Both corpora verified identical first: 605 chunks, 605 vectors, 351 carrying the header, on each side. The measurement is the exact cosine rank of the correct file across all 605 chunks, with no index approximation, no fusion, no reranker and no score floor, so nothing but the embedder is being compared.

Question’s correct answere5-large-v2code embedderBetter
TransferBasket.cs9315code, by 78
CatalogFilterPaginatedSpecification.cs6017code, by 43
CacheHelpers.cs5633code, by 23
HomePageHealthCheck.cs26261code, by 201
GetMyOrdersHandler.cs2499e5, by 75
Logout.cshtml.cs5995e5, by 36
EmptyBasketOnCheckoutException.cs (control)1124e5, by 123
Register.cshtml.cs (control)251e5, by 49

The code-specialized embedder wins four of the six hard questions, and on one of them the margin is large. It still does not make them retrievable. Its best rank anywhere in that set is 15 of 605, so every one of these questions still misses a top-5 result. And it is substantially worse on the two questions the general-purpose model answers at rank 1 and rank 2.

The honest conclusion is that neither embedder can answer these six, that the two are good at different questions rather than one being better, and that the remaining gap is not something a better embedding model fixes. Chunking strategy is the open candidate.

Experiment 3: which reranker, on identical input

The experiment that should have come first. A cross-encoder is a pure function of (query, documents), so handing two of them the same candidate list removes the embedder, the index, the fusion weights and the score floors from the comparison entirely. Nothing is left varying but the model. Shortlists came from one retrieval call per question at the same pool depth the server reranks at.

Question’s correct answerfirst stagebge-v2-m3Qwen3-0.6BBetter
EmptyBasketOnCheckoutException.cs811tie
ToastComponent.cs441Qwen3
ImageValidators.cs621Qwen3
CatalogContextSeed.cs111tie
ExceptionMiddleware.cs551Qwen3
IdentityTokenClaimService.cs1825bge
GetMyOrdersHandler.cs2774Qwen3
CatalogFilterPaginatedSpecification.cs2894Qwen3
Register.cshtml.cs213bge
OrderBuilder.cs (identifier)231Qwen3
ExceptionMiddleware.cs (identifier)111tie
CacheHelpers.cs (identifier)111tie
BasketQueryService.cs (identifier)111tie
0.452MRR, retrieval alone
+37%bge-reranker-v2-m3
+71%Qwen3-Reranker-0.6B

Qwen3-Reranker-0.6B takes six questions to bge’s two, with five ties, and lifts MRR from 0.452 to 0.772 against bge’s 0.618. Same shortlists, same corpus, same questions.

This is where the retrieval quality on this corpus actually lives. Changing the embedding model moved nothing that mattered: no question became retrievable and both controls got substantially worse. Changing the reranker, on byte-identical input, moves MRR by 71%.

The clearest cases are the ones the first stage nearly lost. GetMyOrdersHandler.cs came out of retrieval at rank 27 and CatalogFilterPaginatedSpecification.cs at 28, both effectively invisible to a user, and Qwen3 pulled them to 4. ExceptionMiddleware.cs sat at 5 and went to 1, where bge left it at 5 and never moved it at all.

bge is not useless. It still adds 37% over raw retrieval and wins two questions outright, including rescuing IdentityTokenClaimService.cs from 18 to 2 where Qwen3 only managed 5. Across the set it is simply the weaker model. And this settles the question the correction at the top of this page raises: the 0.6B was deployed over the 4B for latency, but on measured retrieval quality it also beats the reranker this project previously shipped.

Latency is deliberately not scored here. The two models ran on different hardware, roughly 250 ms per query against roughly 140 s, which measures an A10G against a laptop CPU and says nothing about either model.

Experiment 4: does the embedder earn its place at all?

A cross-encoder never reads a vector. It reads the query and the document text together, so what it needs is a candidate list, not embeddings. Full-text search can produce one. That makes “full-text, then rerank” a complete architecture with no embedding model, no vector column, no index, and no reindex when models change. Given Experiment 3, that deserved a measurement rather than an assumption.

Same reranker on both sides. The only difference is how the shortlist was built.

First stageAnswer in shortlistrank 1top 5MRR
hybrid (vector + full-text)13 of 189130.557
full-text only, no embedder10 of 189100.519

The embedder contributes recall, and nothing else. Where full-text finds the answer but buries it, the reranker erases the difference completely:

Expected answerhybrid first stagefull-text first stagefinal rank, both
EmptyBasketOnCheckoutException.cs8191
ImageValidators.cs6251
ExceptionMiddleware.cs5221
Register.cshtml.cs2273

Rank 25 out of retrieval, rank 1 after reranking, identical to the hybrid run that started at 6. So the embedder is not improving the ordering of anything once a strong cross-encoder is in play. What it does is put three answers in front of the reranker that full-text never surfaces at all: IdentityTokenClaimService.cs, GetMyOrdersHandler.cs and CatalogFilterPaginatedSpecification.cs. Those are pure recall, and no reranker recovers a document it was never shown.

The honest trade, then: the embedding endpoint buys about 17% more answerable questions and zero ranking improvement. That is a real gain and a real cost, and which way it goes depends on the corpus and the budget rather than on anything a benchmark table can tell you.

What this evidence does not support

The model selection is defensible on independent benchmarks plus a decisive live prompt-format measurement. The end-to-end retrieval-quality number is the one thing still missing, and it is gated on labelling that golden set. Stating that plainly is more useful than a chart that implies otherwise.