# Hybrid search and reranking: building a RAG pipeline that finds the right contract number or product code

> Embeddings are good at understanding what a question means, but they easily mistake HD-2024-0153 for HD-2024-0135. This guide fixes that in six steps, from an eval set and a tokenizer through to reranking, and shows how to measure the result.

Bản gốc: https://fdetimes.net/en/guides/hybrid-search-rerank-rag-exact-identifiers/

Picture a member of a legal team typing into an internal chatbot: "What is the penalty clause in contract HD-2024-0153?" The RAG system answers fluently, but the answer comes from contract HD-2024-0135. Semantically the two strings are almost identical, and an embedding has no reason to tell them apart.

This kind of error turns up readily in document collections dense with identifiers, and Anthropic gives an example of the same kind: a user searching for "Error code TS-999" in a technical support database.

According to Anthropic, BM25 is particularly effective for queries containing unique identifiers or technical terms, because it looks for that exact text string in the documents.

On Anthropic's dataset, combining embeddings with BM25 cut retrieval failures in the top 20 chunks by 49%, and adding reranking cut them by 67%. This guide walks through building that pipeline yourself and measuring the equivalent numbers on your own data.

## What you will build, and what you need

The pipeline has five stages: the query runs in parallel through a keyword retriever and a vector retriever, the two lists are merged with reciprocal rank fusion (RRF), the result is reranked, and only then does it reach the LLM. You need Python 3, a set of chunks in the form `{chunk_id: text}`, and access to an embedding and reranking service. The examples below refer to Cohere.

The keyword search code in this article has been **simplified** for teaching purposes. In production, use the built-in BM25 in Elasticsearch or Weaviate.

## Step 1: Build the eval before writing any code

Without measurements, you cannot show a client that you have improved anything. Collect real questions that contain codes, link each to its correct chunk, and measure how often the correct chunk fails to make the top 20. This is also the metric Anthropic uses.

```python
def failure_rate(evalset, retrieve, top_n=20):
miss = sum(1 for q, gold in evalset
if gold not in retrieve(q)[:top_n])
return miss / len(evalset)
```

**Check:** run this function against your current vector pipeline and record the number. That is your baseline.

## Step 2: The tokenizer must keep codes intact

When keyword search misses an identifier, a common cause is the tokenizer, not the algorithm. If "HD-2024-0153" is split into "hd", "2024" and "0153", the token "2024" will match hundreds of other contracts and becomes useless.

```python
import re
TOKEN = re.compile(r"\w+(?:[-/]\w+)*")

def tokenize(text):
return [t.lower() for t in TOKEN.findall(text)]

tokenize("Hợp đồng HD-2024-0153 ký ngày")
# ['hợp', 'đồng', 'hd-2024-0153', 'ký', 'ngày']
```

The example input is Vietnamese ("contract HD-2024-0153 signed on"); note that the contract number survives as a single token.

**Check:** print the tokenized output for 20 real codes taken from the client's data. Each code must come out as exactly one token.

## Step 3: The keyword retriever (simplified)

The function below just counts how many times the query's tokens appear in each chunk. It has none of the IDF or length normalisation of real BM25, but it is enough to show how string matching catches codes.

```python
from collections import Counter

def keyword_search(query, chunks, top_n=20):
q = set(tokenize(query))
scored = []
for cid, text in chunks.items():
tf = Counter(tokenize(text))
score = sum(tf[t] for t in q)
if score:
scored.append((score, cid))
scored.sort(reverse=True)
return [cid for _, cid in scored[:top_n]]
```

**Check:** for a query containing HD-2024-0153, that contract's chunk must rank above the chunk for HD-2024-0135. The HD-2024-0135 chunk may still appear in the list, because this simplified function also scores shared words such as "clause", "penalty" and "contract". If the two chunks tie, go back and check the tokenizer from step 2.

## Step 4: The vector retriever, and a parameter that often gets forgotten

Elastic distinguishes between two kinds of search: keyword search matches on words, while semantic search matches on the meaning of the question. Vector search is still needed for questions such as "which contract has the heaviest penalty for late delivery".

A common mistake sits in the embedding step: Cohere's documentation requires queries to be embedded with `input_type="search_query"`, while documents must use a different input_type value reserved for documents.

```python
# giản lược: embed() là wrapper bạn tự viết quanh SDK
q_vec = embed([query], input_type="search_query")
```

(The comment reads: "simplified: embed() is a wrapper you write yourself around the SDK".)

**Check:** grep the whole codebase to make sure the indexing step and the query step do not share the same input_type value.

## Step 5: Merge with RRF, no weight tuning needed

Anthropic's step is to merge the embedding and BM25 results using rank fusion and deduplicate them. The Elasticsearch documentation gives the RRF formula as follows: for each list, a document's score is increased by `1/(k + rank)`. The advantage is that you do not have to find linear combination weights between the two retrievers.

```python
def rrf(result_lists, k):
scores = {}
for results in result_lists:
for rank, cid in enumerate(results, start=1):
scores[cid] = scores.get(cid, 0.0) + 1.0 / (k + rank)
return sorted(scores, key=scores.get, reverse=True)
```

The `scores` dictionary handles deduplication on its own. You can set k to whatever value your engine's documentation recommends.

A worked example by hand, with k = 10 chosen only to keep the numbers readable. Chunk A ranks 1st on the keyword side and 8th on the vector side: 1/11 + 1/18 ≈ 0.146. Chunk B ranks 1st on the vector side only: 1/11 ≈ 0.091. Chunk C ranks 2nd on both sides: 1/12 + 1/12 ≈ 0.167.

C wins despite topping neither list. That is the essence of RRF: it rewards chunks that both retrievers agree on.

**Điểm mấu chốt:** RRF looks only at rank, so a chunk that appears in both lists usually beats a chunk that tops just one.

If the client uses Weaviate, you do not need to write this function yourself. Weaviate has two fusion algorithms, of which rankedFusion keeps only each result's position in each list and discards the scores. The alpha parameter sets the balance between dense and sparse: 0.5 is an even split, and the default is 0.75.

For query sets heavy with codes, try several alpha values on the eval set from step 1 rather than leaving the default.

## Step 6: Rerank, with YAML formatting for structured data

According to Anthropic, reranking is a filtering technique used to ensure that only the most relevant chunks are passed to the model. Take the top chunks after RRF, send them to Cohere Rerank along with the query, and keep only the first few. Cohere's documentation recommends that if documents contain structured data, they should be formatted as YAML strings for best performance.

```python
# giản lược: không xử lý escape ký tự đặc biệt
def to_yaml(rec):
return "\n".join(f"{k}: {v}" for k, v in rec.items())

to_yaml({"so_hop_dong": "HD-2024-0153",
"ben_ban": "...", "dieu_khoan_phat": "..."})
```

(The comment reads: "simplified: does not escape special characters". The keys mean contract number, seller and penalty clause.)

So for contract records pulled from an ERP, follow that recommendation: convert them into one `key: value` line per field before reranking, rather than sending raw JSON.

**Check:** rerun `failure_rate` for three configurations (vector, hybrid, hybrid + rerank) and put the three numbers side by side in a table.

## Mistakes that make a pipeline look right but behave wrongly

The first is a tokenizer that splits codes, as described in step 2. It raises no exception; results simply get steadily worse. The second is indexing documents with the input_type meant for queries, usually because code was copied over from an experimental notebook. The third is sending structured records to the reranker while ignoring the YAML formatting recommendation.

The fourth is in how results are reported. The 49% and 67% figures are Anthropic's results on Anthropic's data, not a promise for your client's data. Do not put them on a slide as if they were your results. Present the measurements from step 1.

## How this skill shows up at a client site

In your first week with a client, the first thing worth asking is: do users often type product codes, contract numbers or error codes, and what format do those codes follow? The answer determines how you write the tokenizer, before any discussion of which embedding model to pick.

When reading job descriptions for FDE or AI engineer roles, phrases such as "retrieval quality", "hybrid search" and "evaluation" refer to exactly this skill. On your CV, avoid vague lines like "built RAG".

Write it in this form instead: "reduced top-20 retrieval failure rate from X% to Y% on a 200-question eval set containing contract codes, using hybrid search and reranking", where X and Y are numbers you measured yourself.

The client will not remember whether you used RRF or what alpha you chose. They will remember the first time the chatbot answered about the right contract, HD-2024-0153.

**Thử ngay tuần này:**

- Take 30 real questions containing product codes or contract numbers, link each to its correct chunk, and measure the top-20 failure rate of your current vector pipeline
- Implement the hyphen-preserving tokenizer and the rrf function from this article, rerun the same eval set, and compare the two numbers
- Convert 5 contract records to YAML and try reranking them, checking whether the correct record rises to the top of the list

## Nguồn

- [Introducing Contextual Retrieval (Anthropic)](https://www.anthropic.com/news/contextual-retrieval)

- [Reciprocal rank fusion (Elasticsearch documentation)](https://www.elastic.co/docs/reference/elasticsearch/rest-apis/reciprocal-rank-fusion)

- [Hybrid Search Explained (Weaviate)](https://weaviate.io/blog/hybrid-search-explained)

- [What is Semantic Search? (Elastic)](https://www.elastic.co/what-is/semantic-search)

- [Introduction to Embeddings at Cohere](https://docs.cohere.com/docs/embeddings)

- [Rerank Overview (Cohere docs)](https://docs.cohere.com/docs/rerank-overview)
