Ask a billing and coding policy question in plain English and receive an answer based exclusively on the provided source manuals. Every claim includes a citation to the specific passage that supports it. This system uses retrieval-augmented generation (RAG) over reference materials, with an evaluation harness designed to measure retrieval and answer quality.
Billers and coders work extensively with dense policy manuals and reference documents. The corpus includes the NCCI Policy Manual, the Claims Processing Manual, and modifier guidance. Finding the specific passage that resolves a policy question can be time-consuming. SourceFetch turns this reference corpus into a question-answering service. It retrieves relevant passages from a vector store and generates answers based strictly on those passages. Each claim includes an inline citation to its supporting source. SourceFetch complements ScrubCheck. ScrubCheck identifies that a claim is likely to deny. SourceFetch explains why, using the applicable policy manual.
- Retrieves the most relevant passages for a question from a ChromaDB vector store using cosine similarity over embeddings.
- Answers questions using a model constrained to the retrieved passages. Each claim is cited with a
[n]marker, and the system states when the available passages do not contain an answer rather than inventing one. - Provides source context by returning the retrieved passages and their similarity scores alongside the answer. Citation markers link directly to the supporting passages.
- Runs without an API key or model download by default. An offline embedder and extractive answer mode provide a complete, self-contained pipeline out of the box.
GET /api/status What is loaded, which embedder is in use, and which answer mode is active.
POST /api/search Retrieval only: passages and similarity scores.
POST /api/ask Retrieval plus a grounded, cited answer.
The same principle guides the entire system: the model is only allowed to generate answers from retrieved facts. This constraint shapes every layer of the architecture.
- Retrieval provides the substance; generation provides the phrasing. Vector search determines what information the answer can use. The model determines only how to present it. The system prompt prohibits the model from adding any policy, code, or fact that does not appear in the retrieved passages. If the passages do not answer the question, the system reports that the passages do not cover it rather than generating a guess.
- Citations are a first-class feature. Every claim includes a
[n]marker that maps to a specific retrieved passage. This makes the RAG workflow verifiable: users can check each claim against its source, with the interface providing direct access to the supporting passage. - The system degrades gracefully. Without an API key, it returns the top retrieved passages verbatim using extractive mode. Without
sentence-transformers, it falls back to an offline hashing embedder. The service remains useful without requiring a paid API call or a model download.
question ─▶ embed ─▶ Chroma top-k ─▶ assemble cited context ─▶ Claude (grounded) ─▶ answer + citations
└────────────── or ──────────────▶ top passages (extractive)
EMBEDDINGS_BACKEND determines how source text is converted into vectors:
hashing(default): Uses aHashingVectorizerwith word unigrams and bigrams. It is offline, deterministic, and requires no model download. The approach is primarily lexical rather than semantic. It is sufficient for demonstrating the retrieval pipeline and producing useful rankings, while keeping startup immediate and dependency-free.sentence-transformers: Usesall-MiniLM-L6-v2from the Hugging Face ecosystem for semantic retrieval that can better handle paraphrased questions. Install the optional dependency and run ingestion again after enabling it.
Changing the embedding backend changes the vector space. Run the ingestion process again after switching backends. The ingestion script rebuilds the collection.
scripts/evaluate.py evaluates the pipeline against a small gold-standard dataset (eval/qa.jsonl):
- Retrieval: Measures
hit@kand MRR for the expected source document. These metrics are available entirely offline. - Generation: When an API key is available, checks whether the answer contains the expected fact and uses an LLM judge to verify that the answer is supported by the retrieved passages without introducing unsupported claims.
On the included sample corpus using the default offline lexical backend:
hit@3: 14/14 = 1.00
MRR: 0.952
The evaluation set deliberately includes paraphrased questions with limited vocabulary overlap. These cases are expected to challenge the lexical backend and prevent the retrieval metrics from remaining artificially high. This provides a more meaningful evaluation range than a benchmark that consistently produces perfect scores.
- Python 3.10+
- An Anthropic API key is optional but recommended. It enables synthesized answers and grounding evaluation.
git clone https://github.com/rajeshnandipaty/sourcefetch.git
cd sourcefetch
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python scripts/ingest.py # Chunk and embed the sample corpus into Chroma.
uvicorn app.main:app --reload # Serve at http://localhost:8000.Open http://localhost:8000 and ask a question.
cp .env.example .env # Add ANTHROPIC_API_KEY.
pip install -r requirements-semantic.txt # Install optional Hugging Face embeddings.
# In .env: EMBEDDINGS_BACKEND=sentence-transformers.
python scripts/ingest.py # Re-ingest the corpus with the new embedder.
uvicorn app.main:app --reloadpython scripts/evaluate.py --k 5The sample corpus in corpus/ is illustrative and is not authoritative. The included documents cover public CMS topics such as PTP edits, MUEs, modifiers 25, 59, and X, global package rules, and add-on codes. To answer questions using the actual policy manuals, place the relevant CMS PDFs in corpus/ and run the ingestion process again:
cp ~/Downloads/ncci_policy_manual_chapters/*.pdf corpus/
python scripts/ingest.pyThe ingestion pipeline supports .md, .txt, and .pdf files through pypdf, so PDF documents can be used directly without conversion. The CMS NCCI policy manual and related guidance are publicly available.
- Keep retrieval separate from generation. It is tempting to rely on the model's knowledge of billing policy, but approximate recall of a manual is risky in a domain where the exact wording of a rule can determine whether a claim is paid. Placing vector retrieval before generation and restricting the model to the retrieved passages provides a stronger basis for trustworthy answers. Explicitly stating when the source material does not contain an answer is an important part of that design.
- Make citations part of the answer. When every claim points to a specific source passage, users can verify the answer rather than relying on the model's authority. This design decision contributes more to answer credibility than prompt tuning alone.
- Keep embeddings interchangeable. The offline lexical backend makes the application easy to run without external model downloads. The Hugging Face semantic backend provides a stronger path for handling paraphrased queries. Keeping these implementations behind the same interface makes the tradeoff explicit without presenting the lightweight default as a production-grade semantic retriever.
- Use evaluation to measure retrieval quality. A small gold-standard question set and a pair of retrieval metrics turn qualitative confidence into measurable performance. The evaluation set can also identify which types of paraphrased questions the lexical backend handles poorly.
SourceFetch can run as a public demo on Google Cloud Run in retrieval-only mode. The retrieval layer does not require paid API calls, making it suitable for publicly accessible deployments. The answer-generation layer uses a paid API and therefore remains gated behind an API key. Enable answer generation only in environments where API access can be appropriately controlled. See DEPLOY.md for deployment details. The source code is available in the repository, and a short demonstration is available on my portfolio.
SourceFetch is an educational tool that operates on ingested sample and public reference materials. Its answers are limited to the content available in the corpus and may not reflect the most current guidance. It is not a substitute for a certified coding professional or official CMS guidance, and it does not guarantee claim payment or reimbursement.
sourcefetch/
├── app/
│ ├── main.py FastAPI application: /api/status, /api/search, /api/ask, and static UI.
│ ├── rag.py Retrieval → cited context → grounded answer, with extractive fallback.
│ ├── embeddings.py Embedding abstraction: hashing (offline) or sentence-transformers (Hugging Face).
│ ├── store.py ChromaDB wrapper using application-provided embeddings.
│ ├── chunking.py Paragraph-aware chunking with overlap.
│ └── static/
│ └── index.html Single-page UI with links from citations to retrieved passages.
├── scripts/
│ ├── ingest.py Corpus → chunks → embeddings → Chroma; rebuilds the collection.
│ └── evaluate.py Retrieval metrics (hit@k, MRR) and optional grounding evaluation.
├── corpus/ Sample policy documents; illustrative and replaceable with real PDFs.
├── eval/
│ └── qa.jsonl Gold-standard questions for evaluation.
├── requirements.txt Core dependencies; supports offline operation.
├── requirements-semantic.txt Optional Hugging Face embedding dependencies.
├── Dockerfile Builds the index during the image build and runs on Cloud Run.
├── DEPLOY.md Local Docker and Cloud Run deployment guide.
├── .dockerignore
├── .env.example
└── .gitignore