A local applied-AI project exploring how internal IT helpdesk documentation can be turned into grounded, cited answers using retrieval, embeddings, and RAG.
The project started from a practical question: how can an assistant answer internal IT questions without inventing policy details? Instead of building a generic chatbot, this system focuses on controlled document ingestion, keyword and semantic retrieval, citation validation, insufficient-evidence fallback, and evaluation.
The main goal is to learn, build, and demonstrate the engineering building blocks behind reliable LLM applications: retrieval quality, grounding, evaluation, and safe fallback behavior.
- Full-stack RAG demo with FastAPI and React
- Keyword and semantic retrieval pipelines
- OpenAI embeddings with local cache
- Grounded answer generation with citation validation
- Retrieval and answer evaluation harnesses
- Insufficient-evidence fallback to reduce hallucinations
Internal helpdesk information is often spread across policies, setup guides, troubleshooting playbooks, and escalation docs. A generic chatbot can answer fluently, but without retrieval discipline it may invent policy details or cite the wrong source.
This project solves a focused version of that problem: answer employee IT questions only from a controlled fictional knowledge base, show the retrieved evidence, validate citations, and return a safe fallback when the knowledge base does not support an answer.
- Building a full-stack RAG application with FastAPI and React.
- Designing document ingestion, metadata extraction, and section-based chunking.
- Creating a keyword retrieval baseline before adding embeddings.
- Improving retrieval with explainable scoring, field boosts, and synonyms.
- Implementing semantic search with OpenAI embeddings and a local cache.
- Generating grounded answers with citations and confidence labels.
- Handling off-topic questions with insufficient-evidence fallback behavior.
- Validating model citations against retrieved evidence.
- Building deterministic retrieval and answer evaluation harnesses.
- Detecting answer-generation flakiness with repeat-run eval support.
- Markdown knowledge base ingestion from
docs/. - Document metadata extraction: ID, title, filename, category, word count, and section headings.
- Section-based chunking by H2 headings.
GET /documentsandPOST /ingestfor inspecting parsed documents/chunks.POST /searchkeyword retrieval baseline with explainable scoring.- Improved keyword scoring with token normalization, stopwords, synonyms, and weighted title/category/section/content signals.
POST /semantic-searchusing OpenAI embeddings and a local JSON embedding cache.- Retrieval evaluation cases and report generation.
POST /askgrounded RAG answer generation.- Citation validation, citation normalization, and safe fallback citations from retrieved evidence only.
- Confidence labels:
high,partial, andinsufficient. - Insufficient-evidence fallback for off-topic questions.
- Deterministic answer evaluation with required term groups, forbidden terms,
citation checks, confidence checks, latency, and
--repeatsupport. - React + TypeScript + Vite frontend demo for asking questions and inspecting answers, citations, limitations, and retrieved chunks.
Markdown Docs
-> Document Loader
-> Section-Based Chunking
-> Keyword or Semantic Retrieval
-> Retrieved Evidence
-> OpenAI Answer Generation
-> Citation Validation
-> Confidence / Fallback Logic
-> FastAPI /ask Response
-> React Demo UI
The frontend never calls OpenAI directly. It calls the FastAPI backend, and the backend owns retrieval, answer generation, validation, and fallback behavior.
The main engineering goal is not to build another chatbot. The goal is to show the control surfaces that matter in applied AI systems:
- Retrieval quality is measured before answer generation is trusted.
- Keyword retrieval and semantic retrieval are evaluated separately because they fail differently.
- Generated answers are grounded in retrieved chunks, not free-form model memory.
- Citations are validated against actual retrieved evidence.
- Unsupported questions return insufficient evidence instead of invented policy.
- Evaluation is deterministic and repeatable enough to expose flaky behavior.
These are the kinds of practical constraints that matter when moving from an LLM prototype toward a reliable internal assistant.
I chose an IT helpdesk scenario because it is close to real workplace knowledge problems: information is spread across documents, answers need to be actionable, and unsupported questions should not be answered confidently.
- FastAPI
- React + TypeScript + Vite
- OpenAI embeddings
- OpenAI answer generation
uvfor Python environment managementnpmfor frontend package management
backend/ FastAPI API, ingestion, retrieval, semantic search, RAG, eval scripts
docs/ Fictional internal IT helpdesk Markdown knowledge base
eval/ Curated retrieval and answer evaluation cases
frontend/ React + TypeScript demo UI
scripts/ Repository utility scripts
Backend and frontend-specific details live in:
Prerequisites:
- Python 3.11+
uv- OpenAI API key for semantic search and answer generation
Install and run:
cd backend
uv sync
uv run uvicorn app.main:app --reloadThe API runs at http://127.0.0.1:8000.
Create a local backend .env file from the example:
cd backend
Copy-Item .env.example .envSet local values:
OPENAI_API_KEY=your_api_key_here
EMBEDDING_MODEL=text-embedding-3-small
EMBEDDING_CACHE_PATH=.cache/embeddings.json
ANSWER_MODEL=gpt-4o-mini
ANSWER_MAX_CONTEXT_CHUNKS=5
ANSWER_MIN_SEMANTIC_SCORE=0.55
Do not commit .env, API keys, cache files, logs, generated reports,
node_modules, or frontend build output.
In a second terminal:
cd frontend
npm install
npm run devOpen http://127.0.0.1:5173.
The backend enables development CORS for:
http://localhost:5173http://127.0.0.1:5173
Ask:
How do I request VPN access?
Expected behavior:
- Returns a
high confidencegrounded answer. - Cites
VPN Access Policy. - Shows recommended steps and retrieved chunks.
Ask:
My authenticator app is gone
Expected behavior:
- Returns a
partiallygrounded answer. - Cites
MFA Recovery Guide. - Shows limitations when specific support/contact details are not available.
Ask:
What is the vacation policy?
Expected behavior:
- Returns
insufficientconfidence. - Does not invent HR policy details.
- Does not attach unrelated citations.
Health check:
Invoke-RestMethod http://127.0.0.1:8000/healthAsk a grounded RAG question:
Invoke-RestMethod -Method Post `
-ContentType "application/json" `
-Body '{"query":"How do I request VPN access?","top_k":5}' `
http://127.0.0.1:8000/askAsk an insufficient-evidence question:
Invoke-RestMethod -Method Post `
-ContentType "application/json" `
-Body '{"query":"What is the vacation policy?","top_k":5}' `
http://127.0.0.1:8000/askRetrieval evaluation compares keyword retrieval and semantic retrieval against curated cases:
cd backend
uv run python scripts/run_retrieval_eval.pyAnswer evaluation checks grounded /ask behavior with deterministic assertions:
cd backend
uv run python scripts/run_answer_eval.pyRepeat-run answer evaluation helps detect flakiness:
cd backend
uv run python scripts/run_answer_eval.py --repeat 3The answer eval checks:
- Expected confidence.
- Expected citation document IDs.
- Required term groups.
- Forbidden off-topic terms.
- Recommended step counts.
- Basic latency.
Generated reports are written under eval/ and ignored by git.
This is a local demo project. It intentionally avoids several production concerns:
- Uses a local JSON embedding cache instead of a vector database.
- No authentication or authorization.
- No deployment setup.
- No CI/CD yet.
- No observability, tracing, or production monitoring.
- No tenant separation or admin interface.
Reasonable production next steps would be:
- Move embeddings to Postgres + pgvector.
- Add authentication and role-aware access controls.
- Add CI checks for backend, frontend, and eval scripts.
- Add structured logging, tracing, and monitoring.
- Add deployment configuration.
- Add document update workflows and cache invalidation.
This is a complete local demo for exploring and demonstrating applied AI/RAG engineering patterns: retrieval, grounded answer generation, citation validation, fallback behavior, evaluation, and a simple React demo UI.
It is not presented as production-ready software.


