Vector Isn't Enough: Hybrid Search & Retrieval — Jeff Vestal & James Williams, Elastic
Source
youtube.com
Author
AI Engineer
Date
Why it matters
Semantic search misses exact identifiers and keyword search misses paraphrases. The workshop shows how to combine them with RRF or weighted scores and measure with judgment sets before feeding results to an agent.
Key takeaways · AI-distilled
The motivating agent diagnoses exit code 137 as an out-of-memory failure by retrieving documentation and citing it, which makes the quality of its answer depend directly on which documents the search layer returns.
Reciprocal rank fusion merges semantic and BM25 results by rank position, while linear combination normalizes scores and exposes tunable weights. Filters shrink the candidate pool, and search templates can route query types to different strategies.
A bonus section compares pointwise and listwise rerankingA second pass that re-scores retrieved candidates by reading each one against the query, fixing the ordering that fast vector search got approximately right.Full definition → and frames the choice as whether the precision gain justifies another inference step; questions cover chunkingSplitting documents into passages small enough to embed and retrieve individually — the step that quietly determines whether retrieval works at all.Full definition →, embeddingA list of numbers representing a piece of text's meaning, so that similar meanings end up numerically close and can be searched.Full definition → consistency, quantizationShrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.Full definition → and latency.
Terms in this piece · Glossary
reranking — A second pass that re-scores retrieved candidates by reading each one against the query, fixing the ordering that fast vector search got approximately right.
chunking — Splitting documents into passages small enough to embed and retrieve individually — the step that quietly determines whether retrieval works at all.
embedding — A list of numbers representing a piece of text's meaning, so that similar meanings end up numerically close and can be searched.
quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.