Skip to content

Add retrieval for the portfolio assistant - #8

Merged
B-Blarr merged 11 commits into
mainfrom
assistant-retrieval
Sep 26, 2026
Merged

B-Blarr merged 11 commits into
mainfrom
assistant-retrieval

Conversation

@B-Blarr

@B-Blarr B-Blarr commented Sep 26, 2026

Copy link
Copy Markdown
Owner

Stage 3 of the portfolio assistant. The knowledge base is embedded, stored in
PostgreSQL with pgvector, and POST /api/assistant/ returns the closest
sections. No language model yet: retrieval is measured on its own first, so
later answers can be debugged at the right end.

Do not merge yet. Merging deploys, and the first migration runs
CREATE EXTENSION vector. The server does not have postgresql-16-pgvector
yet, and the extension is not trusted, so only a superuser may create it.
The server gets prepared first: install the package, create the extension
once as postgres, and adapt the restore probe to the
COMMENT ON EXTENSION line that pg_dump then writes.

Changes

  • assistant_app becomes a Django app: migration 0001 enables pgvector,
    0002 adds KnowledgeChunk with a 768-dimensional vector and an HNSW index
  • embedding_client.py: httpx client for the embedding service from stage 2
  • knowledge_base.py and manage.py build_index: split the Markdown files at
    every ## heading, embed all sections, replace the index
  • retrieval.py and POST /api/assistant/: the five closest sections with
    their cosine similarity, behind the ASSISTANT_ENABLED switch
  • manage.py evaluate_retrieval with retrieval_questions.json: a fixed
    list of questions to measure retrieval quality
  • The SQLite fallback is removed, PostgreSQL is required
  • Tests for everything above, README updated

Design decisions

  • The extension has its own migration. Django skips CREATE EXTENSION
    when the extension already exists, so the app user on the server never
    needs superuser rights.
  • HNSW with vector_cosine_ops, results ordered by distance. Only that
    exact form uses the index. With 37 rows a sequential scan would do; the
    index is there on purpose, as the part worth learning.
  • build_index embeds everything before it touches the database, then
    swaps old for new chunks in one transaction. A failed run keeps the old
    index, and a visitor never sees an empty table.
  • Two client functions, embed_query and embed_passages, so a question
    can never be embedded with the passage prefix. Timeouts are 5 s and 120 s,
    based on measured 0.3 s per question and 5.3 s for all 37 passages.
  • authentication_classes = [] on the endpoint. DRF authenticates before
    it checks permissions, so a stale token header would otherwise turn
    AllowAny into a 401.
  • ASSISTANT_ENABLED is off by default. The endpoint answers 503 until
    the switch is set, on a server without the embedding service as well.
  • No more SQLite fallback. The HNSW migration cannot run on SQLite, and a
    silent fallback would hide a broken DB_NAME instead of failing at startup.

Retrieval measurement

34 questions on topic (German and English, about Benjamin and addressed to
him directly) and 8 off topic, including a prompt injection:

  • 30 of 34 find an expected section among the first three.
  • The misses are gaps in the content, not in the code: the deployment section
    says "Ausrollen" and never "deploy" (one reworded sentence moved the English
    question to rank 1), and no section gives an overview of projects by stack.
  • Embedding the source label with each passage changed nothing, so the
    simpler form stays.
  • There is no clean threshold. The weakest hit on topic scores 0.788, the
    strongest off-topic question 0.810. A threshold around 0.77 only filters
    obvious noise; borderline questions have to be handled by the system prompt
    in stage 5.

The content fixes follow in a separate pull request.

Verification

  • 84 tests pass against the pgvector image, assistant_app at 100 % coverage
  • flake8, isort and makemigrations --check pass
  • EXPLAIN shows Index Scan using chunk_embedding_hnsw
  • build_index twice gives 37 rows, not 74; with the embedding service down
    it fails with a clear message and the old index stays
  • Endpoint: 503 when switched off or when the service is down, 200 with
    five sections, 400 for a blank question or more than 500 characters,
    200 with a stale token header
  • Tests run with an unreachable service URL as well, so no test depends on
    the real embedding service

@B-Blarr
B-Blarr merged commit 7681139 into main Sep 26, 2026
3 checks passed
@B-Blarr
B-Blarr deleted the assistant-retrieval branch September 26, 2026 18:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant