Build lab

Experiment 03 · RAG & evaluation

I gave my AI two manuals that disagree.

Build a document assistant that can show its sources, notice contradictions and stop when the evidence runs out.

Free project brief · Video planned · Your first LLM feature

The experiment

Build a small question-answering system over two synthetic manuals. Give them one deliberate disagreement and see what happens. This is a project brief for an upcoming experiment, not a claim about completed results.

Prepare the evidence

Create short manuals for a fictional product. Put a matching policy in both, a conflicting policy in each, and a question neither manual answers. Give every document a stable ID, version and date. Define whether a newer version supersedes an older one; if no policy resolves a conflict, the correct result should say so.

The practice documents and evaluation cases in the free toolkit are a starting point. Add your own conflict cases before tuning retrieval.

Build the simplest baseline

Start by placing the two short manuals in context. Ask for an answer plus source IDs. Compare that with a retrieval step only when the collection is large enough to need it. Keep document access checks before retrieval.

Test four situations

SituationWhat a useful answer should do
Both sources agreeAnswer with the supporting passage
One source is explicitly supersededApply the documented version rule
Two current sources conflictName the conflict and request resolution
Neither source answersState the limitation without inventing a policy

Change the wording of the question without changing its meaning. Include a passage that looks relevant but does not support the answer. Check the text of the cited passage, not just whether a citation appears.

What to publish

Record the naive answer, the failure, and the improvement you can reproduce on held-out questions. Share the synthetic manuals, evaluation cases and a small results table. Explain which conflicts remain unresolved.

Use the classroom guides on traceable retrieval and evaluation as the technical companion.