These are published third-party benchmark figures, not measurements taken on this system. They are what the model choices were originally made from. What happened when those choices were actually tested against this corpus is a different page, and the two disagree in places: how the models were chosen.
Rerankers
Cross-encoders scored on code retrieval. MTEB-Code is rerank-top-100; CoIR is the code information retrieval benchmark. License matters as much as score here: three of the strongest models are CC-BY-NC and therefore unusable in a commercial product.
| Model | Params | MTEB-Code | CoIR | License | Outcome |
|---|---|---|---|---|---|
| Qwen3-Reranker-8B | 8B | 81.22 | — | Apache-2.0 | +0.02 over the 4B for roughly double the GPU cost — rejected |
| Qwen3-Reranker-4B | 4B | 81.20 | — | Apache-2.0 | chosen on paper; ~4% better and ~125% slower in practice, never deployed |
| Qwen3-Reranker-0.6B | 0.6B | 73.42 | 65.18 | Apache-2.0 | deployed |
| mxbai-rerank-base-v2 | 0.5B | — | 65.71 | Apache-2.0 | roughly a tie with the 0.6B; kept as a second A/B arm |
| bge-reranker-v2-m3 (incumbent) | 0.57B | 41.38 | 35.97 | Apache-2.0 | replaced |
| jina-reranker-v3 | — | — | 63.28 | CC-BY-NC | license-blocked |
| jina-reranker-v2 | — | — | 56.14 | CC-BY-NC | license-blocked |
| zerank-2 | 4B | nDCG@10 0.6528 | — | CC-BY-NC | license-blocked |
| zerank-1-small | 1.7B | — | — | Apache-2.0 | roughly 3× the serving cost |
The benchmark that was thrown out
A vendor page cited a single headline figure implying a large general-purpose win. Splitting it by benchmark family shows the gap is real but only on code, and that the incumbent is actually the stronger model on multilingual retrieval:
| Benchmark family | Qwen3-Reranker-0.6B | bge-reranker-v2-m3 | Read |
|---|---|---|---|
| CoIR — code | 65.18 | 36.28 | ~29 point code gap, real |
| BEIR — general text | 56.28 | 56.51 | parity |
| MIRACL — multilingual | 57.70 | 69.32 | incumbent wins outright |
The CoIR figures were taken from arXiv 2509.25085v3, Table 2 — the jina-reranker-v3 paper, a rival lab — rather than from either model's own marketing, so neither vendor is grading its own homework.
Embedders
| Model | Dims | Code benchmark | License | Outcome |
|---|---|---|---|---|
| bge-code-v1 | 1536 | CoIR ~81.8 | Apache-2.0 | hosted path; Qwen2.5-Coder-1.5B backbone, 32k context |
| e5-large-v2 (general text) | 1024 | — | MIT | local ONNX path; no GPU, no API cost |
| Qwen3-Embedding-4B | 2560 | MTEB-Code ~80.06 | Apache-2.0 | rejected on a schema limit, not quality — see below |
Qwen3-Embedding-4B was ruled out because 2560 dimensions exceeds SQL Server's
VECTOR(n) ceiling of roughly 1998. Using it would have meant truncating to 1536
or 1024 and losing quality anyway, at around 2.5× the serving cost for comparable code
retrieval. A storage constraint decided that one, not a leaderboard.
What the serving hardware costs
Benchmarks say nothing about the bill. Scale-to-zero is what makes a demo affordable: these are per-hour rates while an endpoint is actually awake.
| GPU | VRAM | $/hr | Note |
|---|---|---|---|
| CPU | — | 0.03 – 0.54 | viable for the local ONNX path only |
| T4 ×1 | 16 GB | 0.50 | ruled out: experimental Turing image, Flash-Attention precision problems |
| L4 ×1 | 24 GB | 0.80 | fits, but roughly half the memory bandwidth |
| A10G ×1 | 24 GB | 1.00 | chosen — ~2× the bandwidth of the L4 for $0.20/hr |
| L40S ×1 | 48 GB | 1.80 | unnecessary at this model size |
| A100 ×1 | 80 GB | 2.50 | unnecessary at this model size |
| H200 ×1 | 141 GB | 5.00 | unnecessary at this model size |
How well these predicted reality
Every figure above is someone else's measurement on someone else's corpus. Four experiments were later run on this one, and the benchmarks turned out to be a good guide in one case and a poor one in another.
| What the benchmarks predicted | What measurement found | Verdict |
|---|---|---|
| Qwen3-Reranker-0.6B should beat bge-reranker-v2-m3 substantially on code — a 29 point CoIR gap. | On identical shortlists it won 6 questions to 2 and lifted MRR from 0.618 to 0.772. | held up |
| bge-code-v1 is state of the art for code retrieval at CoIR ~81.8, so it should clearly beat a general-purpose text embedder. | It won 4 of 6 hard questions but made none of them retrievable, and was far worse on both controls — rank 1 became rank 124. | did not transfer |
| A bigger reranker is a better reranker: the 4B outscores the 0.6B by ~8 MTEB-Code points. | The 4B measured ~4% better and ~125% slower on this workload, so the 0.6B was deployed instead. | true but not worth it |
The pattern is that benchmarks predicted the reranker ranking well and the embedder ranking badly. A leaderboard is a reasonable starting filter and a poor substitute for measuring your own corpus. The measurements, including the two experiments that disproved what was assumed going in, are on how the models were chosen.
See what was actually measured →