Re-scored, 2 October 2026. The labelled rows below were first measured on six template queries, where bge-code-v1 led on nDCG and MRR. Re-scored against 20 real questions labelled by hand, Qwen3-Embedding-0.6B leads on both, and the gap is outside noise. Ranking is no longer evidence against the switch. The negative-control figure turned out to overstate the cost as well; see What it costs below.
Correction, 2 October 2026. The conclusion first published here, that Qwen3-Embedding-0.6B beat the shipping model on every metric, was produced by a build error and is withdrawn. The numbers below are the corrected ones.
Qwen3-Embedding's tokenizer already appends
<|endoftext|> (id 151643) through a post-processor. Its
eos_token_id, however, is 151645 (<|im_end|>), a different
token. The build appended that second token on top, so pooling read the wrong position. The
index built that way was invalid, and so was every figure taken from it.
Caught by comparing against the live inference endpoint: the corrected build matches it at cosine 0.999997, the buggy one at 0.8238. The corrected result is a trade rather than an upgrade, and the recommendation to switch has been withdrawn.
This demo has shipped bge-code-v1 since launch. In October 2026 it was measured
against two alternatives on this corpus, with the chunking, retrieval and reranker held
identical so that only the embedding model varied. It lost.
The two constraints that come before quality
Most published comparisons of code embedding models are not usable here, because two constraints disqualify candidates before accuracy is considered.
SQL Server caps a vector at 1998 dimensions
Vectors in this system live in a SQL Server VECTOR column, which is limited to
1998 dimensions. That removes two of the most frequently recommended code embedders outright:
nomic-embed-code at 3584 dimensions and
Salesforce/SFR-Embedding-Code-2B_R at 2304. Neither fits without truncation, and
truncation is a change that would itself need measuring.
A dedicated vector store does not have this limit. Qdrant accepts 2304, 3584 and 4096, verified directly rather than assumed. Moving the vectors out of SQL Server would take the vector leg out of the fusion procedure and require rebuilding reciprocal rank fusion in application code, so it is an option held in reserve rather than a default.
License
Several strong candidates are CC-BY-NC-4.0, which makes them unusable in a commercial product regardless of how they score. They remain useful as measurement references.
| Model | Params | Dims | Base | License | Eligible |
|---|---|---|---|---|---|
BAAI/bge-code-v1 | 1.54B | 1536 | Qwen2 | Apache-2.0 | yes |
Qwen/Qwen3-Embedding-0.6B | 0.6B | 1024 | Qwen3 | Apache-2.0 | yes |
jinaai/jina-code-embeddings-1.5b | 1.54B | 1536 | Qwen2.5-Coder | CC-BY-NC | measurement only |
Salesforce/SFR-Embedding-Code-2B_R | 2.6B | 2304 | Gemma2 | CC-BY-NC | no, dims and license |
nomic-ai/nomic-embed-code | 7.07B | 3584 | Qwen2.5-Coder-7B | Apache-2.0 | no, dims and 14 GB |
Method
Three databases were built from the same 254 files and 605 chunks of eShopOnWeb. Chunk text was copied verbatim between them, so the chunker, the corpus, the hybrid retrieval procedure, the fusion weights, the score floors and the reranker are identical across all three. The embedding model is the only variable.
Each model was given its own correct query convention, because they are not interchangeable:
bge-code-v1 takes an <instruct>/<query> wrapper, Qwen3
takes an Instruct:/Query: prefix, and the Jina model takes its documented
nl2code prefixes. They also differ on whether an end-of-sequence token is
appended before pooling. Getting any of this wrong produces vectors that look reasonable and
retrieve badly, so each configuration was verified on a relevant-versus-irrelevant ordering
check before a database was built.
The questions
Twenty questions taken from this demo's own usage log, which is to say questions real people actually typed, not questions written to be answered. Four natural negative controls came with them, queries the corpus cannot answer at all. The same twenty real questions, labelled by hand at file level with graded relevance (59 labels), provide nDCG, recall and MRR.
For the unlabelled questions the cross-encoder reranker scores the results. It is independent of the embedding model, which is what makes it usable as a judge, but it is still a model's opinion rather than ground truth.
Results
Both models as the inference stack actually serves them, on identical corpus, chunking, retrieval, fusion weights, score floors and reranker.
| Metric | bge-code-v1 shipping | Qwen3-Embedding-0.6B |
|---|---|---|
| Real questions, top-1 | 0.5083 | 0.6751 |
| Real questions, mean top-5 | 0.3314 | 0.4284 |
| Separation (real minus negative) | 0.5012 | 0.6158 |
| Labelled recall@5 | 0.3600 | 0.5225 |
| Labelled nDCG@5 | 0.4309 | 0.5942 |
| Labelled MRR | 0.6000 | 0.8167 |
| Negative controls (4 queries, lower is better) | 0.0072 | 0.0593 |
| Embedding throughput | 16.8/s | 25.2/s |
| Resident VRAM | 3.1 GB | 1.2 GB |
| Dimensions | 1536 | 1024 |
The third candidate, jina-code-embeddings-1.5b, is competitive on relevance but
scores 0.0596 on the four negative controls, effectively tied with Qwen3 and well above the
shipping model's 0.0072 on that small set. It is CC-BY-NC in any case, so it was only ever a measurement
reference.
What was decided
The embedder was switched to Qwen3-Embedding-0.6B. It went against one metric, the negative controls, and that is worth stating plainly rather than presenting the measurement as unanimous, even though that metric turned out to overstate the cost.
What it buys: better top-1, better mean top-5, better separation, and better recall, nDCG and MRR on the labelled set, at 0.6B parameters against 1.54B, 1024 dimensions against 1536, and 25.2 chunks per second against 16.8. Cheaper to run, cheaper to store, and it finds more of the answers.
What it costs: on the four negative controls the mean top score is roughly eight times worse,
0.0593 against 0.0072. That overstates it. Almost all of the rise is one question: "parity check"
now matches an Equals implementation at 0.23, which is arguably a fair hit.
Re-measured on 38 verified-unanswerable questions through the production relevance gate, both
models let the same 28 through under the score floors then in place. Under the floors re-derived
since, bge-code-v1 would let 7 through and Qwen3 lets 11. That is a small real cost, not eight
times.
On the labelled set the ranking gap is real at this sample size. nDCG@5 0.5942 against 0.4309: paired question by question, Qwen3 wins 7, bge wins 2, and 11 tie, with a 95% bootstrap interval of +0.05 to +0.29. MRR 0.8167 against 0.6000, interval +0.05 to +0.40. An earlier version of this page, scored on six template queries, had bge ahead on both, 0.5587 against 0.5412.
The reasoning for taking that trade: recall and ranking are what a user notices. The score floors have since been re-derived on this model, on 40 answerable and 38 unanswerable questions. Unanswerable questions that return results fell from 28 of 38 to 11, and no answerable question lost its answer.
This page previously recommended keeping the old model. That recommendation was based on the same numbers, read more conservatively.
What this evidence does not support
Latency was not measurable. Embedding took about 2,500 ms per query over the network and 33 ms on the machine itself. All three models landed within noise of each other, which is not a finding about the models. It means the network dominated so completely that inference time was invisible. The throughput and VRAM figures above were measured on the machine and are real; any per-query latency comparison from that run is not.
The question set is small. Twenty real questions, labelled by hand, plus four negative controls. Large gaps show up at that size and small ones do not.
The judge is a model. For the unlabelled questions the reranker decides what is good. It is independent of the embedder, so it cannot simply flatter whichever model produced the candidates, but it is not a human label.
The hand-labelled set this page used to name as outstanding now exists, and the labelled rows above come from it. A larger one would tighten every interval here.
How the reranker was chosen →← Back to search