Student Name: Chidiebele Benjamin Amechi
Index Number: 10022200117
AcaIntel AI is a custom Retrieval-Augmented Generation (RAG) system built for the CS4241 project exam.
It answers user questions using two local sources:
data/Ghana_Election_Result.csv(structured election data)data/2025_budget.pdf(long-form budget document)
This project is implemented from scratch in src/ and does not rely on end-to-end RAG frameworks like LangChain or LlamaIndex.
- Hybrid retrieval: FAISS (dense) + BM25 (lexical) with score fusion
- Query rewrite: expands short or ambiguous phrasing before retrieval
- Domain-aware ranking: election-like queries favor CSV chunks; budget-like queries favor PDF chunks
- Grounded offline QA: no API key required; still returns synthesized answers
- Transparent UI: shows retrieved chunks, scores, effective query, and prompt
- Mode flexibility: supports Offline, OpenAI, Ollama Local, and Ollama Cloud generation
- Multi-chunk answer synthesis: weighted fact voting across retrieved chunks for numeric/comparison answers
- Reliability guardrails: contradiction detection with confidence penalty when evidence conflicts
src/ingestion/: load CSV/PDF source filessrc/preprocessing/: cleaning and chunkingsrc/retrieval/: embeddings, vector store, BM25, hybrid scoring and rerankingsrc/generation/: prompt builder, model clients, offline answer synthesissrc/pipeline/: end-to-end orchestration (RAGPipeline)src/evaluation/: evaluation query sets and runnersrc/utils/: logging and path helpers
User query
-> query rewrite + alias normalization
-> query classification (election / budget / mixed)
-> intent router (structured numeric path vs narrative RAG path)
-> FAISS retrieval + BM25 retrieval
-> metadata-aware filtering boost (source/year/region/party)
-> weighted fusion + domain bonus + lexical rerank
-> cross-encoder rerank (top pool -> requested top-k)
-> dynamic top-k by intent (comparison/numeric pull deeper evidence)
-> dedupe + top-k chunk selection
-> answer composer (direct answer + evidence + confidence + contradiction penalty)
-> response generation (offline / Ollama / OpenAI)
-> JSONL logging with request_id + stage timings
Fusion rule:
final_score = 0.50 * vector_norm + 0.30 * bm25_norm + 0.20 * domain_bonus
-
Record-based chunking for CSV
- Each row becomes a compact, structured chunk
- Ideal for factual, numeric, and comparison questions
-
Paragraph-aware chunking for PDF
- Splits by semantic paragraph boundaries
- Better context continuity for policy/fiscal explanations
Note: fixed-window PDF chunks are retained for experimentation, but live retrieval uses paragraph-aware PDF chunks plus election chunks.
The app supports three answer modes:
-
Offline mode (free, default fallback)
- Works when
OPENAI_API_KEYis empty orOFFLINE_MODE=1 - Produces grounded synthesized answers from retrieved chunks
- Includes numeric refinement logic (vote/percentage extraction and comparative phrasing)
- Works when
-
Ollama local mode (free local model)
- Uses local model served at
OLLAMA_BASE_URL(orOLLAMA_LOCAL_BASE_URL) - Supports full generation and token streaming
- Uses local model served at
-
Ollama cloud mode
- Uses cloud endpoint at
https://ollama.com(orOLLAMA_CLOUD_BASE_URL) - Requires
OLLAMA_API_KEYorOLLAMA_CLOUD_API_KEY - Can be toggled at runtime from the Streamlit sidebar (
Ollama Local/Ollama Cloud)
- Uses cloud endpoint at
-
OpenAI API mode
- Uses
OPENAI_API_KEY - Supports standard chat generation and streaming
- Uses
Run:
streamlit run app.pyMain features:
- Chat-based question answering
- New chat + past conversation archive
- Export current thread (
.mdand.json) - Evidence panel with retrieved chunks and scoring details
- Confidence indicator (low/medium/high)
- Sources used section with chunk previews
- A/B panel toggle (RAG answer vs pure LLM answer)
- Answer mode toggle (
concise,detailed,examiner) - Ollama target toggle (
Ollama Local/Ollama Cloud) in sidebar - Prompt visibility for explainability
- Light/dark UI and response rendering options
The app includes user-safe error handling for:
- pipeline initialization failures
- retrieval failures
- generation failures
python -m venv .venvWindows:
.venv\Scripts\activatemacOS/Linux:
source .venv/bin/activatepip install -r requirements.txtCopy env template:
- Windows PowerShell:
Copy-Item .env.example .env
- macOS/Linux:
cp .env.example .env
Edit .env according to your preferred mode.
OPENAI_API_KEY=
# OFFLINE_MODE=1USE_OLLAMA=1
OLLAMA_MODEL=llama3.1:8b
OLLAMA_BASE_URL=http://127.0.0.1:11434
OLLAMA_API_KEY=
OPENAI_API_KEY=USE_OLLAMA=1
OLLAMA_MODEL=gpt-oss:120b-cloud
OLLAMA_BASE_URL=https://ollama.com
OLLAMA_API_KEY=your_ollama_api_key
OPENAI_API_KEY=USE_OLLAMA=1
OLLAMA_LOCAL_BASE_URL=http://127.0.0.1:11434
OLLAMA_LOCAL_MODEL=llama3.1:8b
OLLAMA_LOCAL_API_KEY=
OLLAMA_CLOUD_BASE_URL=https://ollama.com
OLLAMA_CLOUD_MODEL=gpt-oss:120b-cloud
OLLAMA_CLOUD_API_KEY=your_ollama_api_key
# Optional default selection (auto-detected from URL)
OLLAMA_BASE_URL=http://127.0.0.1:11434
OLLAMA_MODEL=llama3.1:8bOPENAI_API_KEY=sk-...
OPENAI_MODEL=gpt-4.1-miniInstall Ollama (Windows PowerShell):
irm https://ollama.com/install.ps1 | iexPull model:
ollama pull llama3.1:8bVerify:
ollama listRun:
python -m src.evaluation.run_evaluationOutput: outputs/evaluation_results.json
Benchmark metrics: outputs/benchmark_metrics.json
Key fields include:
rag_answerpure_llm_answerquery_typeretrieved_chunk_idseffective_queryabstention_detectedconfidencecitationstimings_ms
Use manual rubric columns in the output JSON for final project scoring notes.
Regression benchmark:
- Benchmark dataset:
evaluation/benchmarks/core_benchmark.json - Metrics tracked:
- exact match
- numeric accuracy
- abstention correctness
- citation precision
- CI blocks regressions if key thresholds drop.
- Runtime logs:
outputs/logs.jsonl(append-only JSON lines) - Evaluation results:
outputs/evaluation_results.json
Each log entry includes query, effective query, retrieved chunks, prompt, response, confidence, citations, request id, stage timings, and timestamp.
To prebuild retrieval artifacts (FAISS index, embeddings cache, structured election store):
python build_index.pyThis reduces cold-start overhead in deployed environments.
Run all tests:
python -m pytest tests/Project currently includes retrieval, chunking, and pipeline integration tests. Additional tests validate structured answer composition for numeric/comparison synthesis.
Your repo is already on GitHub:
https://github.com/fzSwift/Ai_1002200117
- Go to Streamlit Community Cloud.
- Click New app.
- Select:
- Repository:
fzSwift/Ai_1002200117 - Branch:
main - Main file path:
app.py
- Repository:
- Click Deploy.
In Streamlit app settings, open Secrets and add only what you need.
Offline-only mode:
OFFLINE_MODE = "1"
OPENAI_API_KEY = ""OpenAI mode:
OPENAI_API_KEY = "sk-..."
OPENAI_MODEL = "gpt-4.1-mini"Important:
- Do not use Ollama mode on Streamlit Cloud (
USE_OLLAMA=1) because there is no local Ollama server in that environment. runtime.txtis pinned to Python 3.11 for package compatibility (FAISS/sentence-transformers).
- Push changes to
mainand Streamlit redeploys automatically. - You can also trigger a manual reboot/redeploy from the app settings.
- If build fails, check Manage app -> Logs first.
- If memory is tight, reduce retrieval
top_kin UI defaults and avoid unnecessary large model behavior. - Keep secrets only in Streamlit Secrets, not in
.envcommitted files.
- Answers are constrained by CSV/PDF coverage
- Query rewrite is rule-based and may miss unusual phrasing
- Numeric interpretation may depend on row granularity in source data
- Offline mode is grounded and improved, but still less fluent than larger cloud LLMs
This system is designed to demonstrate:
- grounded retrieval-first QA
- explainability (evidence + prompt transparency)
- robust offline operation
- clean separation of ingestion, retrieval, generation, and evaluation components