Running E5 Embedding Models Locally
9 min read · updated August 11, 2026
E5 models will embed anything you hand them and give you a vector of the right shape whether or not you followed the convention they were trained under. That is the problem: the failure is silent, it costs retrieval quality rather than raising an error, and it is the single most common mistake made with this family.
The prefixes are part of the model
Every input to an E5 model must begin with either "query: " or "passage: ", including the trailing space. These are not hints. During training every example carried one of them, so the token that begins the sequence is part of what the model conditions on, and omitting it puts the input off the distribution the encoder was fitted to.
The rules, as the intfloat model cards state them:
- Asymmetric tasks — passage retrieval, open-domain question answering, ad-hoc search. Use
"query: "for the short side and"passage: "for the indexed side. - Symmetric tasks — semantic textual similarity, paraphrase retrieval, bitext mining. Use
"query: "for both sides. Not one of each. - Embeddings as features — clustering, linear probes, classification. Use
"query: ".
The symmetric rule is the one people get wrong, because “query” and “passage” sound like they describe the two sides of any comparison. They do not: they describe the two roles the model learned, and a symmetric comparison has only one role in it. Using query: against passage: for a sentence-similarity task asks the model to score a retrieval match between two things that are not a query and a document.
Running it
pip install sentence-transformers
python - <<'PY'
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("intfloat/multilingual-e5-base")
docs = ["passage: " + d for d in [
"The Amsterdam metro opened in 1977.",
"Rotterdam's harbour is the largest in Europe.",
]]
qs = ["query: when did the amsterdam metro open"]
D = model.encode(docs, normalize_embeddings=True)
Q = model.encode(qs, normalize_embeddings=True)
print((Q @ D.T).round(3))
PYThe multilingual variants are built on XLM-RoBERTa: base is 12 layers at 768 dimensions, large is 24 layers at 1024, both capped at 512 tokens and both MIT licensed. The English-only e5-*-v2 models use the same prefixes with a BERT backbone. Everything on this page applies to both; what differs is where a multilingual model spends its parameters.
Masked mean pooling, done properly
E5 pools by averaging the token vectors, not by taking the [CLS] position the way BGE does. If you use sentence-transformers, its pooling module reads the model’s configuration and does this for you. If you drop to raw transformers, you must implement it, and the part that goes wrong is padding.
A batch is padded to its longest member. Those padding positions still produce hidden states, and a plain .mean(dim=1) averages them in — so the vector you get for a short sentence depends on what else happened to be in the batch with it. The same text embedded alone and embedded next to a long document produces two different vectors. That is a genuinely nasty bug, because it makes your index non-reproducible.
import torch
from torch import Tensor
from transformers import AutoTokenizer, AutoModel
def average_pool(last_hidden: Tensor, mask: Tensor) -> Tensor:
h = last_hidden.masked_fill(~mask[..., None].bool(), 0.0)
return h.sum(dim=1) / mask.sum(dim=1)[..., None]
tok = AutoTokenizer.from_pretrained("intfloat/multilingual-e5-base")
enc = AutoModel.from_pretrained("intfloat/multilingual-e5-base").eval()
texts = ["query: how tall is the dom tower",
"passage: The Dom Tower in Utrecht is 112 metres tall."]
batch = tok(texts, max_length=512, padding=True,
truncation=True, return_tensors="pt")
with torch.no_grad():
out = enc(**batch)
vecs = average_pool(out.last_hidden_state, batch["attention_mask"])
vecs = torch.nn.functional.normalize(vecs, p=2, dim=1)
print((vecs[:1] @ vecs[1:].T).item())The masked version zeroes the padded positions and divides by the real token count, which makes the result independent of batch composition. Verify it: encode one sentence alone and again in a batch with a long one, and check the two vectors are identical to floating-point tolerance. If they are not, your pooling is reading padding.
The symmetric-task rule in practice
Deduplication is the case where this bites hardest. You have a million support tickets and you want the near-duplicates. That is a symmetric task: every ticket plays the same role. Prefix all of them with "query: ", not half with each, and not with "passage: " on the grounds that they are documents.
Clustering is the same. So is any use where you feed embeddings to a downstream classifier — the model card is explicit that features should be produced under the query: prefix. The mental model that keeps this straight: passage: means “this text is a thing that might answer a question”, and if no question is involved, nothing is a passage.
The rule also settles a question that comes up in hybrid systems. Suppose you retrieve with query: and passage: and then want to cluster the retrieved documents to diversify the results. Those are two different tasks over the same texts, and the correct answer is two different embeddings: the passage: vectors for retrieval and query: vectors for the clustering step. Reusing the retrieval vectors is cheaper and slightly wrong, and whether that matters is a judgement about your data rather than a rule — but it should be a judgement you made rather than an accident.
One thing the prefixes do not do is disambiguate models. E5 and BGE both accept arbitrary text and both return 768-dimensional unit vectors at base size, so a configuration error that swaps one for the other produces an index that loads, queries and returns results. The only signal is quality. If you serve more than one embedding model, write the model name and revision into the index metadata and check it at query time; the alternative is discovering the mismatch from a user complaint.
What the prefix costs you
Two things, both small and both worth knowing.
The first is token budget. "passage: " tokenises to a few tokens that come out of the 512-token maximum, and truncation is silent. On documents chunked to exactly 512 tokens, adding the prefix pushes the last few tokens of every chunk off the end. Chunk to the limit minus the prefix length, not to the limit.
There is a related trap in chunk overlap. If you chunk with a stride so that consecutive chunks share text, the prefix is added to each chunk, so the shared region appears in both chunks at a different offset from the start of the sequence. That is harmless for retrieval but it does mean two chunks containing identical text produce different vectors, which surprises people running a deduplication check over their own index and finding no exact matches.
The second is that the prefix must be applied at exactly one layer of your system. Applying it in the ingestion script and again in a shared embedding helper gives you "passage: passage: ...", which is a different input producing different vectors — and since the whole corpus gets the same double prefix, the index is internally consistent and nothing looks wrong until queries with a single prefix start matching poorly against it. Put the prefix in one function, assert on it if you like, and keep the raw text out of the encoder entirely.
- Decide whether your task is asymmetric or symmetric before you write any code; that decision sets both prefixes.
- Encode with sentence-transformers, or implement masked mean pooling and verify batch-independence.
- Normalise, then use dot product as cosine similarity.
- Chunk to 512 minus the prefix length, and apply the prefix in exactly one place.