Benchmarks

Published figures for every model considered · and how they held up when measured

← Back to search

These are published third-party benchmark figures, not measurements taken on this system. They are what the model choices were originally made from. What happened when those choices were actually tested against this corpus is a different page, and the two disagree in places: how the models were chosen.

Rerankers

Cross-encoders scored on code retrieval. MTEB-Code is rerank-top-100; CoIR is the code information retrieval benchmark. License matters as much as score here: three of the strongest models are CC-BY-NC and therefore unusable in a commercial product.

ModelParamsMTEB-CodeCoIRLicenseOutcome
Qwen3-Reranker-8B8B81.22—Apache-2.0+0.02 over the 4B for roughly double the GPU cost — rejected
Qwen3-Reranker-4B4B81.20—Apache-2.0chosen on paper; ~4% better and ~125% slower in practice, never deployed
Qwen3-Reranker-0.6B0.6B73.4265.18Apache-2.0deployed
mxbai-rerank-base-v20.5B—65.71Apache-2.0roughly a tie with the 0.6B; kept as a second A/B arm
bge-reranker-v2-m3 (incumbent)0.57B41.3835.97Apache-2.0replaced
jina-reranker-v3——63.28CC-BY-NClicense-blocked
jina-reranker-v2——56.14CC-BY-NClicense-blocked
zerank-24BnDCG@10 0.6528—CC-BY-NClicense-blocked
zerank-1-small1.7B——Apache-2.0roughly 3× the serving cost

The benchmark that was thrown out

A vendor page cited a single headline figure implying a large general-purpose win. Splitting it by benchmark family shows the gap is real but only on code, and that the incumbent is actually the stronger model on multilingual retrieval:

Benchmark familyQwen3-Reranker-0.6Bbge-reranker-v2-m3Read
CoIR — code65.1836.28~29 point code gap, real
BEIR — general text56.2856.51parity
MIRACL — multilingual57.7069.32incumbent wins outright

The CoIR figures were taken from arXiv 2509.25085v3, Table 2 — the jina-reranker-v3 paper, a rival lab — rather than from either model's own marketing, so neither vendor is grading its own homework.

Embedders

ModelDimsCode benchmarkLicenseOutcome
bge-code-v11536CoIR ~81.8Apache-2.0hosted path; Qwen2.5-Coder-1.5B backbone, 32k context
e5-large-v2 (general text)1024—MITlocal ONNX path; no GPU, no API cost
Qwen3-Embedding-4B2560MTEB-Code ~80.06Apache-2.0rejected on a schema limit, not quality — see below

Qwen3-Embedding-4B was ruled out because 2560 dimensions exceeds SQL Server's VECTOR(n) ceiling of roughly 1998. Using it would have meant truncating to 1536 or 1024 and losing quality anyway, at around 2.5× the serving cost for comparable code retrieval. A storage constraint decided that one, not a leaderboard.

What the serving hardware costs

Benchmarks say nothing about the bill. Scale-to-zero is what makes a demo affordable: these are per-hour rates while an endpoint is actually awake.

GPUVRAM$/hrNote
CPU—0.03 – 0.54viable for the local ONNX path only
T4 ×116 GB0.50ruled out: experimental Turing image, Flash-Attention precision problems
L4 ×124 GB0.80fits, but roughly half the memory bandwidth
A10G ×124 GB1.00chosen — ~2× the bandwidth of the L4 for $0.20/hr
L40S ×148 GB1.80unnecessary at this model size
A100 ×180 GB2.50unnecessary at this model size
H200 ×1141 GB5.00unnecessary at this model size

How well these predicted reality

Every figure above is someone else's measurement on someone else's corpus. Four experiments were later run on this one, and the benchmarks turned out to be a good guide in one case and a poor one in another.

What the benchmarks predictedWhat measurement foundVerdict
Qwen3-Reranker-0.6B should beat bge-reranker-v2-m3 substantially on code — a 29 point CoIR gap. On identical shortlists it won 6 questions to 2 and lifted MRR from 0.618 to 0.772. held up
bge-code-v1 is state of the art for code retrieval at CoIR ~81.8, so it should clearly beat a general-purpose text embedder. It won 4 of 6 hard questions but made none of them retrievable, and was far worse on both controls — rank 1 became rank 124. did not transfer
A bigger reranker is a better reranker: the 4B outscores the 0.6B by ~8 MTEB-Code points. The 4B measured ~4% better and ~125% slower on this workload, so the 0.6B was deployed instead. true but not worth it

The pattern is that benchmarks predicted the reranker ranking well and the embedder ranking badly. A leaderboard is a reasonable starting filter and a poor substitute for measuring your own corpus. The measurements, including the two experiments that disproved what was assumed going in, are on how the models were chosen.

See what was actually measured →