AI · 5 min read
Build RAG From Scratch in Python: Test Retrieval With LangChain
Build a tested LangChain retrieval-to-prompt pipeline with source IDs, access scopes and repeatable indexing, before adding model generation.
If a retrieved document is wrong, a polished answer can make the mistake harder to see. The first artifact worth inspecting in a retrieval-augmented generation pipeline is the context that reaches the model, with its source IDs still attached.
This Python experiment uses actual LangChain documents, a recursive text splitter, a runnable chain and a prompt template. It checks retrieval and prompt construction against three local documents. It deliberately stops before model generation: there is no paid API, downloaded embedding model or claim about generated-answer quality.
The original RAG paper combines retrieval with generation. The lab here isolates an application boundary within that broader idea. It is useful precisely because a test can say which document should be present before a nondeterministic model response enters the picture.
Start with questions that have known sources
The supplied corpus contains a public refund policy, a public API retry policy and a staff-only payroll fixture. The refund document states a fourteen-day window. The payroll text contains an obviously synthetic code; it is present to test access scope, not to supply real private data.
Each document has a stable source ID and an explicit scope. Those fields survive splitting. A request for refund window should retrieve refund-policy; a public request for payroll secret should retrieve nothing. A question about orbital velocity has no matching source in the corpus.
These expectations are more actionable than “the chatbot sounded correct.” A failure tells us whether to inspect tokenization, chunking, ranking or access filtering. The tests do not infer correctness from a generated sentence that happens to contain the number fourteen.
Use the real framework without hiding the retrieval rule
The example pins LangChain Core 0.3.75 and LangChain Text Splitters 0.3.9. It uses Document, RecursiveCharacterTextSplitter, RunnableLambda and PromptTemplate. The splitter documentation describes recursive separator-based splitting; the test adds a longer document so that splitting is exercised rather than only instantiated.
Chunks are limited by character count in this example, with an overlap and starting offset recorded. Character count is not a language model's token budget. Before sending real prompts, measure tokens with the selected model's tokenizer and reserve space for the question, instructions and answer.
Retrieval uses normalized term counts implemented in the downloadable source. It is a lexical cosine baseline, not semantic embedding search. A small stop-word list reduces obvious filler terms; documents with no token overlap are excluded. This makes fixture behavior reproducible, but it will miss paraphrases that share no terms and can retrieve irrelevant text that shares a word.
That limitation is intentional and inspectable. Replacing the baseline with an embedding service should preserve the surrounding tests and add representative paraphrase cases. It should not silently redefine “zero matches” as a guarantee that no answer exists.
Filter access before ranking
The retriever first selects documents in the allowed scopes, then scores those candidates and limits the result count. The public chain has only the public scope; the staff test explicitly supplies the staff scope. In a real application, those permissions must come from trusted authentication and authorization state, not a scope field provided by the person asking the question.
The pipeline joins the selected chunks with their source IDs and constructs the prompt:
def build_pipeline(retriever, allowed_scopes=('public',)): def prepare(question): docs=retriever.retrieve(question,allowed_scopes) context='\n'.join(f"[{d.metadata['source']}] {d.page_content}" for d in docs) return {'question':question,'context':context or '(no matching context)'} return RunnableLambda(prepare) | PROMPTA source label is not evidence that a model will cite it accurately. It is a handle that lets later checks trace an answer back to the supplied context. The prompt also says that context is untrusted data. That instruction is useful documentation of intent, but this lab contains no prompt-injection resistance claim and gives retrieved text no tool-execution privileges.
The empty-result path includes an explicit (no matching context) marker. It does not fabricate a fallback document. A generation layer could then abstain without calling a model, or request an answer that acknowledges insufficient evidence. Which behavior is appropriate needs a product decision and a test at that additional boundary.
Re-indexing must remove old content
Each chunk ID is derived from source ID, starting offset and text. Repeating the same import produces the same IDs. Changing a source changes its chunk identity. The in-memory replace operation replaces the complete corpus, so a deleted source does not survive as an old chunk.
This is a deliberate small-corpus design. It avoids pretending that an append-only import is a synchronization strategy. A persistent vector store needs corresponding upserts and deletion tracking, plus a way to publish a consistent index version. Stable IDs alone do not delete obsolete records.
The test suite runs replacement twice, checks identical IDs, changes a source and checks a changed ID, then removes the refund source and verifies that the old refund text is no longer retrievable. Those are useful checks even if the eventual embedding model or database changes.
Run the boundary before adding generation
Download the implementation, corpus, tests, requirements and setup instructions. With the pinned dependencies installed:
python3 -B -m unittest -v test_example.pyAll six tests pass on Python 3.11.5. The recorded demonstration returns refund-policy for the refund query, zero matches for the out-of-corpus query and zero staff-document matches for a public caller. It then prints the actual LangChain prompt containing the refund text and its source ID.
The next experiment should attach one explicitly configured model at the prompt boundary, preserve the retrieved IDs in the response record and evaluate answers separately from retrieval. Include questions whose correct answer is absent, documents containing misleading instructions and revisions that contradict old text. The local checks here establish what context reaches that boundary. They leave the model's faithfulness, semantic retrieval quality, latency and remote service behavior unverified.
Found a mistake or tried a different approach?
Send Alex a note ↗