Foundations
The architecture of production-grade retrieval-augmented generation.
A RAG prototype can find a document and still give a poor answer. NEXUS follows every layer between a user's question and a grounded response: processing the corpus, finding candidates, reranking evidence, constructing context, and measuring whether the whole system works.
- 8 chapters
- End-to-end RAG
- Worked Helix example
Digital and print purchase links will be added when the book is published.
The idea
About the book
Production RAG is a chain of decisions. The wrong chunk boundary, a weak embedding, a missed identifier, or an unsupported citation can each spoil the final answer. NEXUS makes those layers visible so an engineer can diagnose and improve them one at a time.
Eight chapters build a practical baseline from first principles. A fictional CI/CD platform called Helix carries the examples through the book, so the design choices stay connected to one evolving system.
Inside Volume I
What you'll take from the book
Eight chapters · one connected RAG pipeline
The chapters follow the order in which a production system needs each decision, from shaping documents to validating the answer it returns.
- Chapter 01
What retrieval actually is
Which problem should retrieval solve?
Separate search, retrieval, and grounded generation before choosing an architecture.
- Know when RAG fits, when a small corpus can live in context, and when a task calls for something else.
- Build a real-query baseline and measure retrieval, answer quality, and abstention separately.
- Chapter 02
The embedding space
How do you choose a model for your corpus?
Understand the geometry behind dense retrieval and the trade-offs that shape an embedding choice.
- Compare models on your own queries instead of selecting by a general leaderboard score.
- Reason about dimensions, asymmetric queries, and quantization as quality and storage decisions.
- Chapter 03
Chunking & document processing
What should become a retrievable unit?
Turn mixed documents into chunks that can be found and still contain enough context to answer.
- Process Markdown, PDFs, code, tables, and conversations according to their structure.
- Tune chunk size and overlap against answer quality, with metadata that preserves source context.
- Chapter 04
Retrieval algorithms & indexes
How do you find the right candidates?
Combine lexical and semantic search, then make the index fast without losing measurable recall.
- Use BM25 and dense retrieval together, with reciprocal rank fusion to merge their results.
- Choose and calibrate an ANN index against exact search; account for filtering before it harms recall.
- Chapter 05
Reranking
Which retrieved chunks deserve the model’s attention?
Add a precise second pass between fast candidate retrieval and the final prompt.
- Use a cross-encoder to rescore candidates before sending a smaller set to the language model.
- Measure precision, latency, and cost against a retrieval-only baseline.
- Chapter 06
Query understanding
What if the user’s question is a poor search query?
Rewrite or expand queries only where the pattern helps the query in front of you.
- Recognize when a conversational follow-up, ambiguous question, or metadata filter needs different handling.
- Route rewriting, HyDE, and other patterns conditionally; track both their lift and their harm rate.
- Chapter 07
Context construction
How do retrieved chunks become a trustworthy answer?
Build the boundary between retrieved evidence and generation, where citations and refusal are enforced.
- Ground citations in retrieved chunk IDs and validate them before returning an answer.
- Make abstention explicit and treat retrieved documents as data, never as instructions.
- Chapter 08
The production baseline
How do all the layers work as one system?
Bring the seven disciplines together into a RAG pipeline with clear contracts and measurable outputs.
- Follow one query through retrieval, reranking, context construction, generation, and evaluation.
- Know what the foundations cover and what later volumes reserve for operations and advanced patterns.
The running example
One question, traced through the stack.
In the fictional Helix platform, a user asks why a runner pod restarts every four hours on Node 22. The book follows that question from retrieval to a cited answer, showing what each layer contributes and what remains uncertain.
- 01
Understand the question
Keep the specific “Node 22” identifier while preparing the query for retrieval.
- 02
Find and rank evidence
Fuse retrieval results, then rerank candidates to select the most useful runbook and policy chunks.
- 03
Answer with receipts
Cite the restart policy, validate the cited sources, and flag a separate Node 22 issue whose cause is still unconfirmed.
Who it's for
Engineers and architects who have a RAG prototype and need to make it dependable under real queries. It assumes basic familiarity with Python, embeddings, and language models, then builds the measurement and architecture decisions needed for a production baseline.
Coming soon
Know when NEXUS is available
The book is being prepared for release. Follow the newsletter for publication news; digital and print links will appear here once the editions are available.
Get release updates →