PeterParker8991Short answer: RAG hallucination persists when an ask-your-docs chatbot receives stale, incomplete, or...
Short answer: RAG hallucination persists when an ask-your-docs chatbot receives stale, incomplete, or conflicting evidence, so good embeddings and a larger context window still produce wrong answers. For an edtech assistant that reviews code changes, the practical fix is a versioned evidence pipeline: pin the course and repository revision, retrieve requirements and changed code separately, reject weak evidence, validate a small finding schema, and replay a fixed evaluation set before release.
| Choice | Wrong-answer risk | Provider portability | Operating burden |
|---|---|---|---|
| Send the nearest chunks directly to generation | High when revisions or document types conflict | Low if prompts depend on one provider's response shape | Low |
| Add reranking but no evidence gate | Lower ranking risk; unsupported output can still escape | Medium | Medium |
| Use a typed evidence contract plus abstention | Explicitly controlled at the application boundary | High | Medium |
Recommendation: own the evidence contract in Node.js and make the model provider an adapter. This costs more engineering time than a single prompt, but it protects the scarce resource in a one-person SaaS: hours that should go into the weekly product release, not manual review of plausible findings.
Embeddings answer a narrow question: which indexed passages appear related to the query? A code-review assistant needs a harder answer: which passages are authoritative for this exact change, repository revision, assignment, language, and policy version?
Those are different jobs.
Imagine a learner changes gradeSubmission() in revision 8f21c4a. The index contains the current rubric, last term's rubric, an instructor exception, generated API documentation, and a discussion that quotes obsolete behavior. A semantically close chunk can be irrelevant because its authority or effective revision is wrong. Increasing the context window may make matters worse: both rules now fit, so generation receives more contradiction rather than more truth. My rule is blunt: freshness is a hard filter, not a similarity hint. Every indexed unit should carry fields such as repositoryId, revision, documentKind, effectiveFrom, and supersedesId, plus stable source coordinates. A finding that cites “somewhere in the handbook” isn't reviewable; one tied to a file, symbol, rubric item, and immutable content digest is. I won't trade that traceability for a marginally simpler ingestion job because the later debugging cost lands on the same person trying to ship the next feature.
Version first.
This changes the debugging question. Stop asking only, “Was the right chunk in the top five?” Ask, “Was every eligible chunk valid for the reviewed revision, and did an older rule survive ingestion?” The second question catches a class of failures that embedding swaps and larger windows cannot.
A diff and a rubric play different roles. The diff describes what changed. The rubric defines what should be true. Mixing both into one vector query produces a convenient bag of text, but it erases that distinction before the model has to reason about it.
Use two retrieval lanes. The first resolves code evidence from the changed files, surrounding symbols, tests, and pinned base revision. The second resolves policy evidence from the active assignment rubric and applicable instructor rules. Join them only after each lane passes metadata filters. This is intentionally boring infrastructure. Boring ships.
The join should create explicit claim candidates such as “the new branch can return an ungraded state” paired with code coordinates and the rubric clause that makes the state invalid. If either side is missing, there is no finding yet. There is only a search lead.
Chunk boundaries matter here, but fixed token counts are a blunt default. Keep a function with its signature, keep a rubric requirement with its exceptions, and preserve parent headings as metadata. Overlap can help a boundary case, though excessive overlap duplicates evidence and can make several retrieved chunks look like independent support. Deduplicate by source identity and content digest before generation.
For a weekly release cadence, I would outsource parsing where a maintained parser already exists and keep the domain rules in application code. The differentiator is not splitting Markdown. It is knowing that “late submission exception” overrides the general deadline only for a named assignment cohort.
Similarity scores are not calibrated truth probabilities. Do not print a universal cutoff into a blog post and pretend it transfers across embedding models, corpora, and distance functions. Choose thresholds from labeled queries in your own corpus, then version those thresholds alongside the retriever.
The gate still needs deterministic rules. Here is a compact contract for an initial implementation:
type Evidence = {
sourceId: string;
digest: string;
revision: string;
kind: "diff" | "code" | "rubric" | "policy" | "test";
locator: string;
score: number;
text: string;
};
type ReviewRequest = {
repositoryId: string;
revision: string;
rubricRevision: string;
};
type EvidenceBundle = {
code: Evidence[];
rules: Evidence[];
};
function admitEvidence(
request: ReviewRequest,
candidates: Evidence[],
minimumScore: number,
): EvidenceBundle | null {
const eligible = candidates.filter((item) =>
item.score >= minimumScore &&
(item.kind === "rubric" || item.kind === "policy"
? item.revision === request.rubricRevision
: item.revision === request.revision),
);
const unique = [...new Map(
eligible.map((item) => [`${item.sourceId}:${item.digest}`, item]),
).values()];
const bundle = {
code: unique.filter((item) =>
item.kind === "diff" || item.kind === "code" || item.kind === "test"),
rules: unique.filter((item) =>
item.kind === "rubric" || item.kind === "policy"),
};
return bundle.code.length > 0 && bundle.rules.length > 0 ? bundle : null;
}
The minimumScore is configuration derived from evaluation, not a magic number hidden in the function. The revision equality checks are more important than another decimal place of similarity. If admitEvidence returns null, the correct response is a typed insufficient_evidence result. Do not ask the generator to improvise.
That refusal path is product behavior, so design it deliberately. It can request the missing rubric, wait for indexing, or route the change to human review. OWASP's guidance for LLM applications treats excessive agency and overreliance as risk areas; keeping authorization and consequential actions outside generated prose follows the same defensive boundary.
Free-form review text hides retrieval defects. A typed result exposes them.
Require each finding to name the code location, the violated rule, supporting source identifiers, and a confidence category defined by your application. Then validate the object before it reaches a pull request or learner dashboard. JSON Schema is useful here because validation stays outside the model and does not depend on a provider-specific SDK.
type Finding = {
summary: string;
severity: "info" | "warning" | "error";
codeLocator: string;
ruleLocator: string;
evidenceIds: string[];
};
type ReviewResult =
| { status: "complete"; findings: Finding[] }
| { status: "insufficient_evidence"; missing: string[] };
function evidenceIsClosed(
result: ReviewResult,
admittedIds: ReadonlySet<string>,
): boolean {
if (result.status === "insufficient_evidence") return true;
return result.findings.every((finding) =>
finding.codeLocator.length > 0 &&
finding.ruleLocator.length > 0 &&
finding.evidenceIds.length >= 2 &&
finding.evidenceIds.every((id) => admittedIds.has(id)),
);
}
Requiring two evidence identifiers is an application policy in this example, not a universal law: one should resolve to code and one to a rule. In production, validate those kinds as well as membership. Also verify that quoted spans actually occur at the stored locator. A syntactically valid citation can still point at evidence that does not support the claim.
Keep the portable boundary small: messages in, typed result out, usage and timing metadata beside it. Provider-specific tool calls, token accounting, safety fields, and retry semantics belong inside adapters. The core review service should not know which SDK produced the candidate object.
Portability has a limit. Different models may interpret the same evidence differently, so swapping an adapter is not proof of equivalent behavior. Run the same evaluation before promoting any provider or model change.
An ask-your-docs system has at least four independently changing artifacts: source documents, the chunker and metadata extractor, the retrieval configuration, and the generation configuration. Record their versions on every trace. Without that tuple, a wrong answer cannot be reproduced after the index changes.
Build a small regression set from real domain shapes without inventing “golden” prose. Each case should contain a pinned diff, active rules, expected evidence locators, allowed finding categories, and explicit cases that must abstain. Include adversarial cases: an obsolete rubric with better lexical overlap, two assignments that share a function name, a deleted test, and a policy excerpt embedded inside untrusted source text. OWASP identifies prompt injection as a core LLM application risk, so retrieved content must be treated as data rather than instructions.
Measure stages separately. Retrieval recall asks whether required evidence survived the pipeline. Citation precision asks whether cited material supports a finding. Schema validity asks whether downstream code can safely parse the result. Abstention tests ask whether missing or conflicting evidence stops publication. A single “answer quality” score cannot tell you which component to repair.
Start with 30 to 50 carefully reviewed cases if that is what one person can maintain; the number is a workflow choice, not a statistical guarantee. Add a case whenever a new failure shape appears. Run the set when documents are re-indexed, retrieval logic changes, prompts change, or a provider adapter changes. Block publication on deterministic failures, and review semantic differences before release.
This is the revenue-per-hour decision: spend effort once on replayable evidence and avoid repeatedly inspecting opaque chatbot output. Ship weekly, but promote the index, retriever, and generator as one tested release unit. Bigger context is an input budget. It is not a correctness mechanism.