Skip to content

Latest commit

 

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

IT Helpdesk Knowledge Assistant

A local applied-AI project exploring how internal IT helpdesk documentation can be turned into grounded, cited answers using retrieval, embeddings, and RAG.

The project started from a practical question: how can an assistant answer internal IT questions without inventing policy details? Instead of building a generic chatbot, this system focuses on controlled document ingestion, keyword and semantic retrieval, citation validation, insufficient-evidence fallback, and evaluation.

The main goal is to learn, build, and demonstrate the engineering building blocks behind reliable LLM applications: retrieval quality, grounding, evaluation, and safe fallback behavior.

Highlights

  • Full-stack RAG demo with FastAPI and React
  • Keyword and semantic retrieval pipelines
  • OpenAI embeddings with local cache
  • Grounded answer generation with citation validation
  • Retrieval and answer evaluation harnesses
  • Insufficient-evidence fallback to reduce hallucinations

Problem

Internal helpdesk information is often spread across policies, setup guides, troubleshooting playbooks, and escalation docs. A generic chatbot can answer fluently, but without retrieval discipline it may invent policy details or cite the wrong source.

This project solves a focused version of that problem: answer employee IT questions only from a controlled fictional knowledge base, show the retrieved evidence, validate citations, and return a safe fallback when the knowledge base does not support an answer.

What This Project Covers

  • Building a full-stack RAG application with FastAPI and React.
  • Designing document ingestion, metadata extraction, and section-based chunking.
  • Creating a keyword retrieval baseline before adding embeddings.
  • Improving retrieval with explainable scoring, field boosts, and synonyms.
  • Implementing semantic search with OpenAI embeddings and a local cache.
  • Generating grounded answers with citations and confidence labels.
  • Handling off-topic questions with insufficient-evidence fallback behavior.
  • Validating model citations against retrieved evidence.
  • Building deterministic retrieval and answer evaluation harnesses.
  • Detecting answer-generation flakiness with repeat-run eval support.

Features

  • Markdown knowledge base ingestion from docs/.
  • Document metadata extraction: ID, title, filename, category, word count, and section headings.
  • Section-based chunking by H2 headings.
  • GET /documents and POST /ingest for inspecting parsed documents/chunks.
  • POST /search keyword retrieval baseline with explainable scoring.
  • Improved keyword scoring with token normalization, stopwords, synonyms, and weighted title/category/section/content signals.
  • POST /semantic-search using OpenAI embeddings and a local JSON embedding cache.
  • Retrieval evaluation cases and report generation.
  • POST /ask grounded RAG answer generation.
  • Citation validation, citation normalization, and safe fallback citations from retrieved evidence only.
  • Confidence labels: high, partial, and insufficient.
  • Insufficient-evidence fallback for off-topic questions.
  • Deterministic answer evaluation with required term groups, forbidden terms, citation checks, confidence checks, latency, and --repeat support.
  • React + TypeScript + Vite frontend demo for asking questions and inspecting answers, citations, limitations, and retrieved chunks.

Architecture

Markdown Docs
  -> Document Loader
  -> Section-Based Chunking
  -> Keyword or Semantic Retrieval
  -> Retrieved Evidence
  -> OpenAI Answer Generation
  -> Citation Validation
  -> Confidence / Fallback Logic
  -> FastAPI /ask Response
  -> React Demo UI

The frontend never calls OpenAI directly. It calls the FastAPI backend, and the backend owns retrieval, answer generation, validation, and fallback behavior.

Why This Project Matters

The main engineering goal is not to build another chatbot. The goal is to show the control surfaces that matter in applied AI systems:

  • Retrieval quality is measured before answer generation is trusted.
  • Keyword retrieval and semantic retrieval are evaluated separately because they fail differently.
  • Generated answers are grounded in retrieved chunks, not free-form model memory.
  • Citations are validated against actual retrieved evidence.
  • Unsupported questions return insufficient evidence instead of invented policy.
  • Evaluation is deterministic and repeatable enough to expose flaky behavior.

These are the kinds of practical constraints that matter when moving from an LLM prototype toward a reliable internal assistant.

I chose an IT helpdesk scenario because it is close to real workplace knowledge problems: information is spread across documents, answers need to be actionable, and unsupported questions should not be answered confidently.

Tech Stack

  • FastAPI
  • React + TypeScript + Vite
  • OpenAI embeddings
  • OpenAI answer generation
  • uv for Python environment management
  • npm for frontend package management

Repository Layout

backend/   FastAPI API, ingestion, retrieval, semantic search, RAG, eval scripts
docs/      Fictional internal IT helpdesk Markdown knowledge base
eval/      Curated retrieval and answer evaluation cases
frontend/  React + TypeScript demo UI
scripts/   Repository utility scripts

Backend and frontend-specific details live in:

Setup

Backend

Prerequisites:

  • Python 3.11+
  • uv
  • OpenAI API key for semantic search and answer generation

Install and run:

cd backend
uv sync
uv run uvicorn app.main:app --reload

The API runs at http://127.0.0.1:8000.

Environment

Create a local backend .env file from the example:

cd backend
Copy-Item .env.example .env

Set local values:

OPENAI_API_KEY=your_api_key_here
EMBEDDING_MODEL=text-embedding-3-small
EMBEDDING_CACHE_PATH=.cache/embeddings.json
ANSWER_MODEL=gpt-4o-mini
ANSWER_MAX_CONTEXT_CHUNKS=5
ANSWER_MIN_SEMANTIC_SCORE=0.55

Do not commit .env, API keys, cache files, logs, generated reports, node_modules, or frontend build output.

Frontend

In a second terminal:

cd frontend
npm install
npm run dev

Open http://127.0.0.1:5173.

The backend enables development CORS for:

  • http://localhost:5173
  • http://127.0.0.1:5173

Demo Examples

Ask:

How do I request VPN access?

Expected behavior:

  • Returns a high confidence grounded answer.
  • Cites VPN Access Policy.
  • Shows recommended steps and retrieved chunks.

Ask:

My authenticator app is gone

Expected behavior:

  • Returns a partially grounded answer.
  • Cites MFA Recovery Guide.
  • Shows limitations when specific support/contact details are not available.

Ask:

What is the vacation policy?

Expected behavior:

  • Returns insufficient confidence.
  • Does not invent HR policy details.
  • Does not attach unrelated citations.

Demo Screenshots

High Confidence: Grounded VPN Answer

High confidence VPN answer with citation

Partial Confidence: MFA Recovery

Partial confidence MFA recovery answer

Insufficient Evidence: Unsupported HR Question

Insufficient evidence fallback

Useful API Commands

Health check:

Invoke-RestMethod http://127.0.0.1:8000/health

Ask a grounded RAG question:

Invoke-RestMethod -Method Post `
  -ContentType "application/json" `
  -Body '{"query":"How do I request VPN access?","top_k":5}' `
  http://127.0.0.1:8000/ask

Ask an insufficient-evidence question:

Invoke-RestMethod -Method Post `
  -ContentType "application/json" `
  -Body '{"query":"What is the vacation policy?","top_k":5}' `
  http://127.0.0.1:8000/ask

Evaluation

Retrieval evaluation compares keyword retrieval and semantic retrieval against curated cases:

cd backend
uv run python scripts/run_retrieval_eval.py

Answer evaluation checks grounded /ask behavior with deterministic assertions:

cd backend
uv run python scripts/run_answer_eval.py

Repeat-run answer evaluation helps detect flakiness:

cd backend
uv run python scripts/run_answer_eval.py --repeat 3

The answer eval checks:

  • Expected confidence.
  • Expected citation document IDs.
  • Required term groups.
  • Forbidden off-topic terms.
  • Recommended step counts.
  • Basic latency.

Generated reports are written under eval/ and ignored by git.

Limitations

This is a local demo project. It intentionally avoids several production concerns:

  • Uses a local JSON embedding cache instead of a vector database.
  • No authentication or authorization.
  • No deployment setup.
  • No CI/CD yet.
  • No observability, tracing, or production monitoring.
  • No tenant separation or admin interface.

Reasonable production next steps would be:

  • Move embeddings to Postgres + pgvector.
  • Add authentication and role-aware access controls.
  • Add CI checks for backend, frontend, and eval scripts.
  • Add structured logging, tracing, and monitoring.
  • Add deployment configuration.
  • Add document update workflows and cache invalidation.

Status

This is a complete local demo for exploring and demonstrating applied AI/RAG engineering patterns: retrieval, grounded answer generation, citation validation, fallback behavior, evaluation, and a simple React demo UI.

It is not presented as production-ready software.

About

Applied AI/RAG helpdesk knowledge assistant with FastAPI, React, semantic search, citations, fallback behavior, and evals.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages