Semantic Search Principles

What embedding similarity captures and misses, why queries and documents must share one space, and how to diagnose bad semantic search results.

Last generated

Lesson 2 of 38 available15 practice questions

SPACED REPETITION Β· 15 practice questions

Make this lesson stick.

Try 3 questions now. No account needed. Sample answers aren't saved.

The rule that expired and still won

You have built a search box for an appliance shop's help center. It holds sixteen short chunks of text: returns, refunds, delivery, two water filters, a password page, a forum post, kettle care. Each chunk becomes a vector through all-MiniLM-L6-v2, a small open embedding model, and each customer question returns the chunks whose vectors lie closest to the question's vector. Before shipping, you write down fourteen real customer questions, mark the chunk that answers each one, and run them.

Eleven of the fourteen come back right at rank 1. Not bad for a few dozen lines of Python. The three misses are more interesting than the hits:

Question Came back first (score) Should have come first (score)
"Can I return it after 45 days?" "Returns policy (2023, replaced): items can be returned within 60 days, opened or not." (0.600) "Unopened items can be returned within 30 days of delivery for a full refund." (0.551)
"How do I reset my password?" "Forum post: How do I reset my password? I have tried everything and nothing works." (0.813) "To reset your password, choose Forgot password on the sign-in page and follow the link we email you." (0.798)
"I want to send back the kettle I tried" "Descale your kettle once a month in hard-water areas: …" (0.500) "Opened or used small appliances cannot be refunded. If one is faulty, we exchange it within 14 days." (0.365)

The first miss is the one that costs money. A chatbot reads the top chunk and tells the customer "Yes, you have 60 days." That rule was replaced; the chunk even says so. Yet the model did its job. Of the sixteen texts, the old policy is the one most like the question: it is about returns, it counts days, it even covers opened items. Whether a rule is still in force is not something "alike" can see.

The other two misses fail differently. In the second, the model found a text that asks the same question instead of one that answers it. In the third, the shared word "kettle" outweighed the meaning "return something I used", even though the right chunk, which shares no word at all with the question, still came second.

That is semantic search in practice. It ranks texts by how alike they are, in the sense its training taught it. Often that is exactly relevance. Sometimes it is not, and you can learn to predict when. Each move below starts with a symptom you will see in real results.

# Symptom in the results Move
1 You can't predict what a query will match Read the geometry
2 The top hit is on topic but has the wrong number, part, polarity or version Know the blind spots: similar β‰  relevant β‰  true
3 Questions retrieve other questions, not answers Search asymmetrically: prefixes and instructions
4 Relevance hinges on one detail of the query meeting one detail of the passage Know what the bi-encoder trades away
5 Worse in production than in the notebook, and nothing raised an error Keep the indexing/querying contract
6 Fine for English and everyday words, poor for other languages or jargon Check domain and language fit
7 "Search is bad" Diagnose before you fix

Why keyword search misses paraphrases, and the full path a query takes through a search system, are covered in Foundations of Modern AI Search. How embedding models are built and chosen is in Vector Embeddings; how the scores are computed is in Cosine Similarity & Distance Metrics. This lesson is about what the scores mean.

Before any move: give your search an exam

Every claim in this lesson comes with a measurement, and your own system deserves the same treatment. The cheapest measuring device in retrieval is a labeled query set: real questions, each marked with the chunk that answers it. Here is the help center and its fourteen-question exam. The code in this lesson runs top to bottom in one Python session (Python 3.12, sentence-transformers 6.1).

import numpy as np
from sentence_transformers import SentenceTransformer

CHUNKS = [
    "Unopened items can be returned within 30 days of delivery for a full refund.",                                   # 0
    "Opened or used small appliances cannot be refunded. If one is faulty, we exchange it within 14 days.",           # 1
    "Refunds go back to the original payment method within 5 business days after we receive the item.",              # 2
    "Standard delivery takes 3 to 5 business days. Express delivery arrives the next working day if you order before 2 pm.",  # 3
    "The XK-4471 replacement filter fits the AquaPure 300 jug. Change it every 2 months.",                            # 4
    "The XK-4417 replacement filter fits the AquaPure 500 tap unit. Change it every 6 months.",                       # 5
    "Motors in blenders and mixers carry a 2-year manufacturer warranty. Claims go directly to the manufacturer.",    # 6
    "To change your delivery address, open the order and choose Edit address. This works until the parcel is dispatched.",  # 7
    "Returns policy (2023, replaced): items can be returned within 60 days, opened or not.",                          # 8
    "To reset your password, choose Forgot password on the sign-in page and follow the link we email you.",          # 9
    "Forum post: How do I reset my password? I have tried everything and nothing works.",                             # 10
    "Descale your kettle once a month in hard-water areas: boil equal parts water and white vinegar, then rinse twice.",  # 11
    "Gift cards are valid for 3 years and cannot be exchanged for cash.",                                             # 12
    "We match a lower price from another UK retailer if the item is identical and in stock.",                        # 13
    "You can cancel an order free of charge until it is dispatched. After that, use the returns process.",            # 14
    "We deliver to addresses in the UK and Ireland only.",                                                            # 15
]
# Real-looking customer questions, each labeled with the chunk that answers it.
LABELED = [
    ("I want to send back the kettle I tried", 1),
    ("Can I return it after 45 days?", 0),
    ("when will I get my money back", 2),
    ("how fast is next-day shipping", 3),
    ("Which AquaPure model takes the XK-4417 filter?", 5),
    ("my mixer motor stopped working after a year", 6),
    ("I moved house, can I update where my order goes", 7),
    ("How do I reset my password?", 9),
    ("white crust inside my kettle", 11),
    ("do gift vouchers expire", 12),
    ("I found it cheaper elsewhere", 13),
    ("can I stop my order before it ships", 14),
    ("do you deliver to Norway", 15),
    ("how often should I replace the AquaPure 300 filter", 4),
]

model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")   # 384-d, returns unit vectors

def evaluate(doc_model, query_model=None, doc_prefix="", query_prefix=""):
    """Rank the labeled chunk for every question. Returns hit@1, MRR and the misses."""
    query_model = query_model or doc_model
    D = doc_model.encode([doc_prefix + c for c in CHUNKS], normalize_embeddings=True)
    Q = query_model.encode([query_prefix + q for q, _ in LABELED], normalize_embeddings=True)
    ranks, misses = [], []
    for (question, gold), scores in zip(LABELED, Q @ D.T):
        rank = 1 + int((scores > scores[gold]).sum())     # 1 means the right chunk came first
        ranks.append(rank)
        if rank > 1:
            misses.append((question, rank, int(scores.argmax())))
    ranks = np.array(ranks)
    return int((ranks == 1).sum()), round(float((1 / ranks).mean()), 3), misses

hits, mrr, misses = evaluate(model)
print(f"hit@1 {hits}/{len(LABELED)}   MRR {mrr}")
for question, rank, top in misses:
    print(f"  rank {rank}: {question!r} -> chunk {top} came first")
# hit@1 11/14   MRR 0.893
#   rank 2: 'I want to send back the kettle I tried' -> chunk 11 came first
#   rank 2: 'Can I return it after 45 days?' -> chunk 8 came first
#   rank 2: 'How do I reset my password?' -> chunk 10 came first

hit@1 is the share of questions whose labeled chunk came first. MRR (mean reciprocal rank) averages 1/rank, so a chunk at rank 2 earns 0.5: here (11 Γ— 1 + 3 Γ— 0.5) / 14 = 0.893. Retrieval Metrics covers these and their relatives properly.

Fourteen questions is a toy. One question moves hit@1 by 7 points, so this exam can catch a pipeline that is broken (in Move 5 you will watch 13 right answers fall to 1), not prove that one decent model beats another by 2%. For that you need dozens to hundreds of labeled queries, counted separately for the slices you care about: identifiers, languages, question types.

πŸ’‘ Every question you diagnose by hand becomes a new labeled query. The exam then grows exactly where your system is weak.

Move 1: Read the geometry

Symptom: you can't predict what a query will match, or explain why an odd result came back.

"The company it keeps", scaled up

The idea underneath is old. Zellig Harris (1954) argued that words appearing in similar contexts tend to have similar meanings, and J. R. Firth (1957) put it memorably: "You shall know a word by the company it keeps." Word vectors built from that idea place words used in the same contexts near each other.

Sentence embedding models apply it to whole texts, with an explicit teacher: pairs of texts that belong together. E5, for example, was pre-trained on about 270 million naturally occurring pairs, such as Reddit posts and their comments, Stack Exchange questions and their upvoted answers, and titles with their passages, and then fine-tuned on labeled data (Wang et al., 2022). Training pulls the two vectors of each pair together and pushes the other texts in the batch away. A pretrained language model alone does not do this well: Reimers & Gurevych (2019) found that simply averaging BERT's outputs gave sentence embeddings "often worse than averaging GloVe embeddings". Vector Embeddings tells that story. For reasoning about results, keep one sentence:

🧠 "Similar" means "looks like the kind of pair this model was trained to pull together."

Each text becomes a point (384 numbers for MiniLM), and search returns the points nearest the query's point. Nothing in that geometry knows what a policy, a part number or a customer is. It knows which texts the training data put near each other.

Predict first

For the question "How do I reset my password?", which text scores higher: "I forgot my login credentials" or "How do I reset my router?"? The router question shares five words with it; the credentials sentence shares two.

def sim(a: str, b: str) -> float:
    va, vb = model.encode([a, b])        # unit vectors, so the dot product is the cosine
    return float(va @ vb)

query = "How do I reset my password?"
for text in ["I forgot my login credentials", "How do I reset my router?"]:
    print(f"{sim(query, text):.3f}  {text}")
# 0.678  I forgot my login credentials
# 0.530  How do I reset my router?
Check your answer

With MiniLM the credentials sentence wins, 0.678 to 0.530: forgetting a login and resetting a password turn up together in the kind of pairs it was trained on. Now the surprise. Here are the same pairs in two other small English models (e5-small-v2 with its query: prefix on both sides, the setting its card gives for symmetric tasks):

Pair all-MiniLM-L6-v2 bge-small-en-v1.5 e5-small-v2
password / "I forgot my login credentials" 0.678 0.817 0.894
password / "How do I reset my router?" 0.530 0.775 0.934
"The meeting is on Monday." / "The capital of France is Paris." 0.028 0.497 0.720

e5-small-v2 puts the router question first. No model is broken: "alike" is whatever the training pairs made it.

The last row is the one to remember. Two sentences with nothing in common score 0.028 in one model and 0.720 in another, higher than MiniLM's best match above. A similarity score is not a probability of relevance, and scores from different models do not share a scale. A rule like "above 0.8 means relevant" belongs to one model, one kind of text and the labeled test that justified it. Cosine Similarity & Distance Metrics shows how to pick a threshold from data.

Directions, and why opposites sit close

The geometry has more structure than distance. word2vec made "king βˆ’ man + woman β‰ˆ queen" famous, but published analogy results exclude the three input words from the possible answers, and without that rule the nearest vector is often one of the inputs (Nissim, van Noord & van der Goot, 2020). Treat the geometry as a useful map, not an algebra of meaning.

The distributional idea also explains a result that surprises people: opposites sit close. "Hot" and "cold" appear in the same sentences ("the coffee is hot", "the coffee is cold"), so MiniLM scores hot/cold at 0.519, not far below hot/warm at 0.717. Keep that in mind for the next move.

πŸ’‘ Measure it. Before you choose any threshold, embed 50 known matches and 50 unrelated pairs from your own data with your model and look at both score ranges. If they overlap, no threshold separates them.

Move 2: Know the blind spots

Symptom: the top hit is about the right thing but has the wrong number, the wrong part, the opposite meaning or an outdated version.

Predict: which of these pairs scores highest with MiniLM, and which lowest?

pairs = [
    ("Opened items can be returned within 30 days.", "Opened items cannot be returned."),  # negation
    ("Refunds are issued within 30 days.", "Refunds are issued within 90 days."),       # a number
    ("Replacement filter XK-4471", "Replacement filter XK-4417"),                       # a part number
    ("I love this phone", "I hate this phone"),                                         # opposite opinion
    ("The meeting is on Monday.", "The capital of France is Paris."),                   # unrelated
]
for a, b in pairs:
    print(f"{sim(a, b):.3f}  {a} | {b}")
Check your answer
0.720  Opened items can be returned within 30 days. | Opened items cannot be returned.
0.852  Refunds are issued within 30 days. | Refunds are issued within 90 days.
0.905  Replacement filter XK-4471 | Replacement filter XK-4417
0.813  I love this phone | I hate this phone
0.028  The meeting is on Monday. | The capital of France is Paris.

The two part numbers, which name different products, form the most "similar" pair of all. Every pair that flips the meaning scores far above the unrelated control.

Why one word barely moves the vector

A sentence vector summarises the whole text. "Can" versus "cannot", 30 versus 90, 4471 versus 4417: each is one small change in the input, and the training pairs rarely taught the model that such a small change reverses what the text says. So the vector barely moves. Researchers see the same pattern at scale. NevIR built pairs of documents that differ only by a negation and found that "most information retrieval models (including SOTA ones) do not consider negation, performing the same or worse than a random ranking"; cross-encoders did best and bi-encoders worst (Weller, Lawrie & Van Durme, EACL 2024). Move 4 explains why the architecture matters.

The part number that was almost found

Identifiers deserve their own measurement. Take a catalogue of ten filters whose part numbers are the same four digits in different orders: XK-4471 fits the AquaPure 300 jug, XK-4417 the AquaPure 500 tap unit, XK-4147 the AquaPure 700 fridge, and so on. Ask for each part twice, once as the bare number ("XK-4417") and once as a question ("Which appliance does the XK-4417 filter fit?"). How often does the right filter come first?

Retriever Right part at rank 1 (of 20 queries)
all-MiniLM-L6-v2 3
bge-small-en-v1.5 15
e5-small-v2 (with its prefixes) 18
BM25 over the same ten chunks 20
How this was measured
import re
from rank_bm25 import BM25Okapi

codes = ["4471", "4417", "4147", "4174", "4714", "4741", "1447", "1474", "7144", "7414"]
fits = ["AquaPure 300 jug", "AquaPure 500 tap unit", "AquaPure 700 fridge", "ClearFlow 2 jug",
        "ClearFlow 5 under-sink unit", "AquaPure 350 jug", "AquaPure 550 tap unit", "ClearFlow 3 jug",
        "AquaPure 900 fridge", "ClearFlow 7 under-sink unit"]
parts = [f"The XK-{c} replacement filter fits the {f}." for c, f in zip(codes, fits)]
queries = [(q, i) for i, c in enumerate(codes)
           for q in (f"XK-{c}", f"Which appliance does the XK-{c} filter fit?")]

def top1_dense(m, q_prefix="", d_prefix=""):
    D = m.encode([d_prefix + p for p in parts], normalize_embeddings=True)
    Q = m.encode([q_prefix + q for q, _ in queries], normalize_embeddings=True)
    return sum(int((Q[j] @ D.T).argmax()) == i for j, (_, i) in enumerate(queries))

tokens = lambda s: re.findall(r"[a-z0-9]+", s.lower())
bm25 = BM25Okapi([tokens(p) for p in parts])
print("all-MiniLM-L6-v2 ", top1_dense(model))
print("bge-small-en-v1.5", top1_dense(SentenceTransformer("BAAI/bge-small-en-v1.5")))
print("e5-small-v2      ", top1_dense(SentenceTransformer("intfloat/e5-small-v2"), "query: ", "passage: "))
print("BM25             ", sum(int(bm25.get_scores(tokens(q)).argmax()) == i for q, i in queries))
# all-MiniLM-L6-v2  3
# bge-small-en-v1.5 15
# e5-small-v2       18
# BM25              20

Two lessons in one table. First, models differ enormously on identifiers, so test yours. Second, even the best dense model here missed two in twenty, while the keyword index found all twenty, because the token "4417" either occurs in a chunk or it does not. When users type identifiers (part numbers, error codes, ticket IDs, clause numbers), give them an exact path: a lexical or learned-sparse index next to the dense one, fused (Hybrid Retrieval Systems, Sparse vs Dense Retrieval), or a structured field you filter on when the identifier can be parsed out of the query (Metadata & Filtering). The cost is a second index to keep in sync and a fusion rule to tune.

Similar, relevant, true

The hook's misses sit at different levels, and each level needs a different tool:

Level Question it answers What can judge it The hook's example
Similar Are these texts alike? Embedding similarity The forum post is very like the password question
Relevant Does this passage help answer this query? A model that reads both together; human labels Only the reset instructions answer it
True Is it correct, and in force, for this user now? Metadata and sources: status, effective dates, authority The 2023 policy is alike and on point, and no longer true

Embedding search measures only the first row. The 45-day failure lives in the third, and no embedding model fixes it, because nothing in the vector marks "replaced" as more important than every other word. Record versions at ingestion (a status field, valid_from and valid_to dates) and filter to the rules in force before ranking; Metadata & Filtering shows how. A reranker works on the second row, not the third. In our run the cross-encoder cross-encoder/ms-marco-MiniLM-L6-v2 happened to put the current rule first (raw scores 3.38 versus 2.81), but the replaced rule stayed right behind it and would still land in a three-chunk prompt.

Your turn. A laptop shop's search receives "laptops without a touchscreen". What will dense retrieval return, and what would you do instead?

Check your answer

Touchscreen laptops, near the top. In a quick test on six laptop descriptions, MiniLM ranked a touchscreen convertible first (0.636) and another touchscreen model second (0.530), above both non-touch laptops. "Without" is one word; "laptop" and "touchscreen" dominate the vector. "Without X" is a constraint on a product attribute, not a topic: parse it into a filter (touchscreen = false) and search the rest of the query semantically. Query Understanding & Intent covers extracting filters from queries safely. To measure it, collect twenty negated queries from your logs and count results that violate the constraint, before and after.

Move 3: Questions are not answers

Symptom: a question retrieves other questions (or a title retrieves other titles) instead of the passages that answer it.

In symmetric search both sides are the same kind of text: finding duplicate questions, matching a sentence to its paraphrase or its translation. In asymmetric search they are not: an eight-word question against a paragraph that answers it. Most help-center and RAG retrieval is asymmetric, and many models trained for it want to be told which side a text is on:

Model family (examples as of 2026) Query side Document side
E5, including multilingual-e5 query: passage:
BGE English v1.5 optional Represent this sentence for searching relevant passages: nothing
Nomic Embed Text v1.5 search_query: search_document:
Qwen3-Embedding Instruct: {task}\nQuery:{query} nothing
Cohere Embed v3 and later (API) input_type="search_query" input_type="search_document"
Voyage (API) input_type="query" input_type="document"

The prefix is part of the model's input format, like a field name the model was trained to read. Here is what it does to the hook's password question with e5-small-v2:

e5 = SentenceTransformer("intfloat/e5-small-v2")      # trained with "query: " and "passage: " prefixes

question = "How do I reset my password?"
answer, forum_post = CHUNKS[9], CHUNKS[10]
for label, q_prefix, d_prefix in [("no prefixes", "", ""),
                                  ("query: on both", "query: ", "query: "),
                                  ("as documented", "query: ", "passage: ")]:
    q = e5.encode(question, prompt=q_prefix, normalize_embeddings=True)
    a, f = e5.encode([answer, forum_post], prompt=d_prefix, normalize_embeddings=True)
    print(f"{label:15s} answer {q @ a:.3f}   forum post {q @ f:.3f}")
# no prefixes     answer 0.895   forum post 0.895
# query: on both  answer 0.901   forum post 0.950
# as documented   answer 0.907   forum post 0.866

Without prefixes the answer and the forum post tie. With query: on both sides, the symmetric setting, the forum post wins clearly. With the documented asymmetric setup the answer wins. The prefixes tell the model to look for the other half of a question-and-answer pair rather than for a twin.

How much does it matter? Measure

The card's setting is the right default, but the size of the effect depends on your data:

e5-small-v2 setup Help-center exam, hit@1 (14 questions) SciFact, nDCG@10 (300 queries)
query: / passage: (documented) 10/14 0.687
no prefixes 11/14 0.680
query: on both 11/14 0.690
passage: on both 11/14 0.636

Read it honestly. On fourteen questions, 10 versus 11 is one question: noise. SciFact asks 300 scientific claims against 5,183 paper abstracts; its queries are declarative sentences shaped much like passages. There the documented setup and no prefixes land within one point of each other, while putting passage: on the queries costs five. (nDCG@10 is a rank-weighted score from 0 to 1; Retrieval Metrics defines it.) The Qwen3-Embedding card reports that leaving out its query instruction can cost roughly 1% to 5% of retrieval performance.

So the rule has two parts. Use exactly what the model card specifies, and use it identically at indexing and at query time. Which mistake costs what depends on your data: query: on the documents let the forum post beat the answer, a risk wherever the corpus holds question-shaped text (FAQs, forum posts, ticket titles), while passage: on the queries cost five points on SciFact. Measure it on your own slices.

⚠️ A library trap. sentence-transformers has encode_query() and encode_document(), which add the prompts stored in the model's configuration. For intfloat/e5-small-v2 those stored prompts are empty strings (checked with sentence-transformers 6.1.0), so encode_query() adds nothing. Pass prompt="query: " yourself, or print model.prompts before you trust the helper.

Your turn. To "keep things consistent", a teammate re-indexes the help center with query: on the documents too. Which row of the table does the system move to, and where would users notice?

Check your answer

It moves to "query: on both", the symmetric setting. On SciFact-like data it would barely register (0.690 versus 0.687). In the help center the forum post now beats the answer, 0.950 to 0.901, so people asking how-to questions start getting other people's unanswered questions. You catch that with a slice of labeled how-to questions in a corpus that also holds question-shaped posts, not with an average over everything.

Move 4: Two towers, one trade-off

Symptom: relevance depends on how one detail of the query meets one detail of the passage (a negation, a condition, which product), and the vectors keep missing it.

Every model in this lesson so far is a bi-encoder: the query and the chunk go through the encoder separately, and only their vectors meet.

OFFLINE: once per chunk                      ONLINE: once per query

chunk 1 ──► encoder ──► vector ──┐           question ──► encoder ──► query vector
chunk 2 ──► encoder ──► vector ───                                        β”‚
  ...                            β”œβ”€β”€β–Ί stored vectors ◄──── compare β”€β”€β”€β”€β”€β”€β”€β”˜
chunk n ──► encoder ──► vector β”€β”€β”˜            β”‚
                                              β–Ό
                                   top-k nearest chunks

Why it is fast. A chunk's vector does not depend on any query, so you compute it once, at indexing time. A query then costs one encoder pass plus a comparison with the stored vectors. For 100,000 chunks of 384 dimensions that is 38.4 million multiply-adds; numpy did it, top-10 included, in about 5 ms on the laptop used for this lesson. Far beyond that, an approximate index replaces the full comparison (Vector Database Architecture, ANN Algorithms).

What it gives up. The chunk vector is fixed before the question exists. It must summarise, in 384 numbers, everything anyone might ask about that chunk, and the question cannot make the model look harder at "not" or at "4417". A cross-encoder reads the question and the chunk together in one pass, so every word of one can interact with every word of the other. That is why cross-encoders handled negation best in NevIR, and why, in our run, cross-encoder/ms-marco-MiniLM-L6-v2 separated the password answer from the forum post by a wide margin (raw scores 8.24 versus 4.58; these are unbounded logits, not probabilities). The price: nothing can be precomputed. One question against 100,000 chunks means 100,000 forward passes. At an assumed 1,000 pairs per second that is 100 seconds per query; on a pool of 100 candidates, 0.1 seconds.

So production systems combine them: the bi-encoder retrieves a candidate pool, then a cross-encoder or an LLM-based reranker reorders it. Reranking Models owns that stage. Between the two sits late interaction (ColBERT, Khattab & Zaharia, SIGIR 2020): it stores one vector per token and matches the query token by token, buying finer matching with far more storage.

Two towers also means that "one space" need not mean "one set of weights". DPR (Karpukhin et al., EMNLP 2020) trained separate question and passage encoders into a shared space. What matters is that the two sides were trained together. That is the next move.

Your turn. Which of this lesson's failures would a cross-encoder over a candidate pool of 100 likely fix: the forum post beating the answer, "laptops without a touchscreen", part-number queries losing to near-twin codes, the replaced 45-day policy?

Check your answer
  • Forum post vs answer: yes; we measured 8.24 versus 4.58. Relevance is exactly what a cross-encoder is trained on.
  • Negation: likely better, not solved. NevIR found cross-encoders best on negation, with a clear gap to humans. For attribute constraints a filter is still more reliable.
  • Part numbers: it can help when the right part is among the candidates, because it reads both strings side by side. An exact-match path is still the dependable fix, and a reranker cannot rescue a chunk that never made the pool.
  • Replaced policy: no. That is a truth problem; the current and the replaced rule are both relevant. Filter on version metadata.

Move 5: Keep the contract between indexing and querying

Symptom: it worked in the notebook; in production the results are worse, and nothing raised an error.

Indexing and querying run at different times, often in different services written by different people. They meet only in the vector store, so everything that shapes a vector must match on both sides.

OFFLINE: indexing job                               ONLINE: query service

document ─► clean ─► chunk ─► "passage: " + chunk        question ─► "query: " + question
                                    β”‚                                     β”‚
                      e5-small-v2 (pinned revision)          e5-small-v2 (same revision)
                      reads at most 512 word pieces          reads at most 512 word pieces
                                    β”‚                                     β”‚
                            unit vector (384) ──► store ◄── search ── unit vector (384)
                            + manifest: model, prefixes, dim, normalized, max length

The contract:

  1. One embedding space. Normally the same model at the same version on both sides. Mix models only where the vendor documents a shared space: Voyage, for example, states that all embeddings created with its 4 series are compatible with each other (Voyage docs). Confirm such a claim on your own labeled set before relying on it.
  2. The same prefixes or instructions on each side, as in Move 3.
  3. The same text preparation: which fields are joined (title plus body?), casing, cleanup.
  4. A known maximum input length. Everything past it is dropped without a warning.
  5. The metric the model was trained for (see Cosine Similarity & Distance Metrics).

Loud mismatch, silent mismatch

If the new model has a different dimension, you are lucky: the search fails loudly (numpy refuses to multiply the shapes; Qdrant 1.19 answers 400 … Vector dimension error: expected dim: 384, got 768). The dangerous case keeps the same dimension: a new model, a new version of the same model, or a changed prefix. Every query still returns results; only their quality drops, anywhere from a little to all the way down to chance. Some pairs of models are partly aligned, so many questions still come back right and a three-query demo after the deploy looks fine. Re-embedding only the new documents is the same bug in slow motion: chunks embedded with the old model tend to sink down the ranking, below unrelated chunks embedded with the new one.

So a model upgrade means re-embedding everything into a second index built alongside the first, running the exam on it, and switching reads over in one step. Vector Embeddings measures the damage of a same-dimension swap and gives the migration recipe; Embedding Pipeline covers doing it at scale.

The fix for the silent case is to make it loud. Store a manifest next to the vectors and check it before every search:

# Written next to the vectors when the index is built.
INDEX_MANIFEST = {
    "model": "intfloat/e5-small-v2", "dim": 384,
    "doc_prefix": "passage: ", "query_prefix": "query: ",
    "normalized": True, "max_seq_length": 512,
}

def check_query_side(query_config: dict, manifest: dict = INDEX_MANIFEST) -> None:
    """Refuse to search if the query side would land in a different space."""
    problems = [f"{key}: index={manifest[key]!r}, query={query_config.get(key)!r}"
                for key in manifest if query_config.get(key) != manifest[key]]
    if problems:
        raise RuntimeError("embedding contract broken -> " + "; ".join(problems))

check_query_side(dict(INDEX_MANIFEST))                    # passes silently
try:
    check_query_side({**INDEX_MANIFEST, "model": "BAAI/bge-small-en-v1.5", "query_prefix": ""})
except RuntimeError as err:
    print(err)
# embedding contract broken -> model: index='intfloat/e5-small-v2', query='BAAI/bge-small-en-v1.5'; query_prefix: index='query: ', query=''

The check compares configuration, not behaviour, so pair it with the exam: run the labeled set as a smoke test after every deploy. That turns a quiet drop in hit@1 into a failed deploy instead of a support ticket.

The words the model never read

all-MiniLM-L6-v2 reads at most 256 word pieces (sub-word tokens) and silently drops the rest: encode() raises no error, and a sentence that starts past the limit changes the chunk's vector by nothing at all. In the English samples measured for this lesson, text ran about 1.2 to 1.6 word pieces per word, so 256 pieces is roughly 160 to 210 words. Vector Embeddings shows the effect on a real document and where to find each model's limit.

Longer-context models move the limit without removing the trade-off: one vector for a long chunk is a blurrier summary. Count tokens with the model's own tokenizer, keep chunks inside the limit, and see Smart Chunking for where to cut.

Stale vectors

The contract also runs through time. If a policy page is edited but its chunk is not re-embedded, the vector still describes the old wording, and search finds the page by what it used to say. Re-embed on every content change, and record which model and which text version produced each vector; Data Freshness & Lifecycle goes deeper.

Your turn. The manifest check passes, yet the exam drops from 12/14 to 6/14 after a deploy. Name two ways the contract can break that the manifest above does not cover.

Check your answer

Any two of these: the model revision changed under the same name (the manifest stores no revision or weights hash); the text preparation changed (the indexer stopped joining titles, or started indexing raw HTML); a tokenizer or truncation setting changed; normalization moved into a different code path; or a partial re-index left some chunks embedded with the old setup. Add the revision and a hash of the preparation code to the manifest, and keep the exam as the behavioural check.

Move 6: Check the fit: domain and language

Symptom: results are good for English questions in everyday words, and poor for other languages or for your field's jargon.

Language

We asked ten of the exam questions in Norwegian and in Turkish, against the same English help center and the same labeled chunks:

Questions in all-MiniLM-L6-v2 (English-only) multilingual-e5-small
English 8/10 7/10
Norwegian 3/10 8/10
Turkish 1/10 6/10

An English-only model sees a Norwegian question as mostly unfamiliar word pieces. A multilingual retrieval model does far better, though not equally well in every language, and on English it is no better here (7 versus 8 is one question). Cross-lingual matching is trained, not automatic: seeing many languages in pretraining does not by itself put a sentence and its translation at the same point. Models learn that from pairs across languages. Reimers & Gurevych (2020) trained a student model so that "a translated sentence should be mapped to the same location in the vector space as the original sentence"; multilingual-E5 was contrastively pre-trained on about a billion multilingual text pairs (Wang et al., 2024).

Predict: add one Norwegian chunk to the English help center, "Gavekort er gyldige i tre Γ₯r og kan ikke byttes mot kontanter." (Gift cards are valid for three years and cannot be exchanged for cash.) Ask the multilingual model the Norwegian question "Kan jeg returnere en vannkoker jeg har prΓΈvd?" (Can I return a kettle I have tried?). Where does the right English chunk land?

Check your answer

Third. The Norwegian gift-card chunk comes first (0.821), the replaced 2023 returns policy second (0.796), and the right answer third (0.792). Without the Norwegian chunk the right answer was already second, behind the replaced policy (Move 2's version problem again); sharing a language pulled an off-topic chunk above every English chunk and pushed the right answer down one more place. This same-language bias is common in multilingual embedders, and it means that adding a few pages in the user's language can push better answers out. Measure each language as its own slice, including mixed-language corpora.

Domain

Jargon breaks the "company it keeps" assumption, because your field keeps different company from the web text the model learned from. Two MiniLM measurements:

  • "The patient had a heart attack." scores 0.797 against "The patient had a myocardial infarction." (the same condition), 0.798 against "The patient had a cardiac arrest." (a different one: a heart attack is a blocked artery, a cardiac arrest is the heart stopping; see the American Heart Association), and 0.771 against "The patient had a panic attack." The sentence frame dominates; the medical distinction barely registers. bge-small-en-v1.5 gets the order right (0.941, 0.927, 0.898), but the whole spread from the same condition to a panic attack is 0.043, on a scale where our unrelated pair scored 0.497. That is far too thin a margin for any threshold.
  • In an HR help desk, "How do I submit a PTO request?" scores 0.181 against "How do I book paid time off?" and 0.495 against "How do I engage the PTO on the tractor?" (power take-off). bge-small-en-v1.5 makes the same mistake, 0.637 versus 0.761.

The moves, cheapest first: add the failing jargon queries to your exam as their own slice; expand acronyms and attach your glossary terms to chunks at indexing time (or expand the query); give a term an exact-match path only where it appears verbatim in the documents, because when they spell it out, a keyword match on "PTO" finds the tractor, not the leave policy; compare two or three candidate models on the slice; fine-tune on your own pairs once you have enough of them. Public leaderboards such as MTEB average over many tasks and domains. Your slice decides.

Shortlist examples as of 2026, each to be tested and none a default winner: multilingual-e5 (open weights, needs query: and passage: ), BGE-M3 (open weights, 100+ languages, 8,192 tokens, dense, sparse and multi-vector output from one model), Qwen3-Embedding (open weights, 100+ languages, 32K context, an instruction on the query side), and API families such as Cohere Embed and Voyage.

Move 7: Diagnose before you fix

Symptom: "search is bad". A ticket, a thumbs-down, a wrong answer from the RAG bot.

The fastest way to waste a week is to swap the embedding model because one query failed. Find the chunk that should have come back, then walk down this list, which runs from cheap and common to expensive and rare:

1. Is it in the index?        missing, stale, or the answer split across chunks
2. Did the model read it?     word pieces vs the model's max length
3. Where does it rank?        rank 2 by 0.05 is a ranking problem; rank 900 is a recall problem
4. Is the contract intact?    model, revision, prefixes, text preparation
5. Which blind spot?          identifier, number, negation, version, question-shaped text
6. Is the model a poor fit?   compare candidates on a labeled slice

A tiny tool covers steps 2 and 3:

def diagnose(question: str, gold: int, m=model) -> None:
    doc_vecs = m.encode(CHUNKS)
    scores = doc_vecs @ m.encode(question)
    order = np.argsort(-scores)
    rank = int(np.where(order == gold)[0][0]) + 1
    top = int(order[0])
    pieces = len(m.tokenizer(CHUNKS[gold])["input_ids"])
    print(f"gold chunk {gold}: rank {rank}, score {scores[gold]:.3f} | top chunk {top}: {scores[top]:.3f}")
    print(f"gold chunk length: {pieces} word pieces (model reads {m.max_seq_length})")
    print(f"top chunk says: {CHUNKS[top][:60]}...")

diagnose("Can I return it after 45 days?", gold=0)
# gold chunk 0: rank 2, score 0.551 | top chunk 8: 0.600
# gold chunk length: 20 word pieces (model reads 256)
# top chunk says: Returns policy (2023, replaced): items can be returned withi...

The right chunk is in the index, 20 word pieces long, at rank 2, less than 0.05 behind a chunk whose text says "replaced". That points straight at step 5, the version blind spot, and the fix is a metadata filter, not a new model.

Run it on "I want to send back the kettle I tried" (gold chunk 1) and you get rank 2 again, behind the descaling chunk: a shared word outweighed the meaning. The tempting fix is a reranker, so measure before you believe it. Reranking MiniLM's top 10 with the cross-encoder from Move 4 put the right chunk first for the password and 45-day questions (hit@1 from 11 to 13 of 14; the 45-day one by luck, as Move 2 showed) and pushed this one from rank 2 to rank 6: the cross-encoder, too, scored the descaling chunk highest. This is a paraphrase miss ("send back" means return, "tried" means used), so it lands on step 6. Build a slice of paraphrased questions, then compare two or three models on it or test rewriting queries into the shop's own words (Query Understanding & Intent). One fix can help two queries and hurt a third, which is why every change runs the whole exam.

Measure it. Each diagnosed failure becomes a labeled query. Track hit@k (did a right chunk make the top k) and MRR per slice, before and after every change, and read the queries that got worse as carefully as the ones that got better. Evaluation & Quality Metrics takes this further.

Final round: no label on the symptom

Real tickets don't say which move they need. Find the gold chunk, walk the list, then pick.

Challenge 1: the Tuesday deploy

After Tuesday's release, support agents say answers "feel random". The logs show no errors, and a product manager's three test queries look fine. The release notes mention "switched the query service to a faster 768-dimensional model"; the index was built last month with a different 768-dimensional model. What happened, and what do you do first?

Check your answer

A silent mismatch: the queries now land in a different space from the documents. Equal dimensions meant no error, and partly aligned spaces or lucky queries explain the passing spot check. Roll the query service back first. Then either re-embed the whole corpus with the faster model into a new index and switch after the exam passes, or keep the old model. Add the model name and revision to a manifest that the query service checks, and run the labeled exam as a deploy gate.

Challenge 2: the notice clause

A contract-search tool gets the query "clauses that do not require prior written notice". The top five results all require prior written notice. The chunks are 150 words long, well inside the model's limit, and the index is fresh. Which step of the list does this land on, and what would you try, in order?

Check your answer

Step 5: a negation blind spot. The bi-encoder matched the topic ("prior written notice") and ignored "do not". First, if clauses carry structured attributes (a requires_notice field extracted at ingestion), filter on them. Otherwise rerank a candidate pool of about 100 with a cross-encoder, which reads the query and clause together and handles negation better, though not perfectly. Measure both on a slice of negated legal queries: the share of top-5 results that violate the constraint.

Challenge 3: Boss level

Redesign the hook's help center so that all three misses are handled, and say how you would know each fix worked.

Check your answer
  • Replaced policy: add status and valid_to metadata at ingestion and filter to rules in force before ranking. Success: zero replaced chunks in any top-k on the exam, plus a new labeled query about an old rule.
  • Forum post beating the answer: move to a model trained for asymmetric search and use its documented prefixes (or keep the model and rerank a candidate pool with a cross-encoder, which put the answer first in our run, 8.24 versus 4.58). Success: the how-to slice of the exam.
  • "Kettle" beating the returns chunk: not the reranker, which made this one worse in our run. Treat it as a paraphrase problem: build a slice of paraphrased questions ("send back", "tried", "not happy with it"), compare two or three models on it, and test query rewriting. Success: the paraphrase slice's hit@1, with no losses on the rest of the exam.

Then re-run the whole exam and read every query that got worse, not just the three you fixed.

Cheat sheet: symptom β†’ move β†’ cost

Symptom Move What it costs How you know it worked
Can't predict matches; a threshold misbehaves Read "similar" as "like a training pair"; tune thresholds per model on labeled data A labeled set Scores of matches and non-matches separate
Wrong part number or error code Exact-match path (lexical or learned sparse), or parse into a filter A second index, fusion tuning Identifier slice hit@1
"Not" or "without" ignored Filter for attribute constraints; rerank the rest Extraction work; reranker latency Constraint-violation rate on a negation slice
An old version ranks first Status and validity metadata, filtered before ranking Ingestion must maintain status No superseded chunks in results
Questions return questions The model's documented prefixes; a reranker A full re-embed (prefixes or model); reranker latency How-to slice
Worse after a deploy, no errors Manifest check plus exam smoke test; full re-embed on any model change Re-embedding compute, double storage during the switch Exam score matches the pre-deploy run
Answers at the end of long chunks are missed Count word pieces; chunk inside the limit More chunks Long-chunk slice
Other languages or jargon are weak Per-slice exam; multilingual or better-fitting model; glossary expansion Evaluation time, plus a full re-embed on a model switch Per-slice hit@1

Before moving on, take one query your own system gets wrong and explain it aloud from the gold chunk outward: is it in the index, did the model read it, where does it rank, is the contract intact, which blind spot is it, and what single measurement would show your fix worked. If you can do that, the rest of this roadmap is a set of tools for the answers.

Next: Vector Embeddings for how the models are trained, compared and chosen; Cosine Similarity & Distance Metrics for what the scores are; Hybrid Retrieval Systems for giving identifiers and rare terms their exact-match path.