Add retrieval for the portfolio assistant - #8
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stage 3 of the portfolio assistant. The knowledge base is embedded, stored in
PostgreSQL with pgvector, and
POST /api/assistant/returns the closestsections. No language model yet: retrieval is measured on its own first, so
later answers can be debugged at the right end.
Changes
assistant_appbecomes a Django app: migration0001enables pgvector,0002addsKnowledgeChunkwith a 768-dimensional vector and an HNSW indexembedding_client.py: httpx client for the embedding service from stage 2knowledge_base.pyandmanage.py build_index: split the Markdown files atevery
##heading, embed all sections, replace the indexretrieval.pyandPOST /api/assistant/: the five closest sections withtheir cosine similarity, behind the
ASSISTANT_ENABLEDswitchmanage.py evaluate_retrievalwithretrieval_questions.json: a fixedlist of questions to measure retrieval quality
Design decisions
CREATE EXTENSIONwhen the extension already exists, so the app user on the server never
needs superuser rights.
vector_cosine_ops, results ordered by distance. Only thatexact form uses the index. With 37 rows a sequential scan would do; the
index is there on purpose, as the part worth learning.
build_indexembeds everything before it touches the database, thenswaps old for new chunks in one transaction. A failed run keeps the old
index, and a visitor never sees an empty table.
embed_queryandembed_passages, so a questioncan never be embedded with the passage prefix. Timeouts are 5 s and 120 s,
based on measured 0.3 s per question and 5.3 s for all 37 passages.
authentication_classes = []on the endpoint. DRF authenticates beforeit checks permissions, so a stale token header would otherwise turn
AllowAnyinto a401.ASSISTANT_ENABLEDis off by default. The endpoint answers503untilthe switch is set, on a server without the embedding service as well.
silent fallback would hide a broken
DB_NAMEinstead of failing at startup.Retrieval measurement
34 questions on topic (German and English, about Benjamin and addressed to
him directly) and 8 off topic, including a prompt injection:
says "Ausrollen" and never "deploy" (one reworded sentence moved the English
question to rank 1), and no section gives an overview of projects by stack.
simpler form stays.
strongest off-topic question 0.810. A threshold around 0.77 only filters
obvious noise; borderline questions have to be handled by the system prompt
in stage 5.
The content fixes follow in a separate pull request.
Verification
assistant_appat 100 % coverageflake8,isortandmakemigrations --checkpassEXPLAINshowsIndex Scan using chunk_embedding_hnswbuild_indextwice gives 37 rows, not 74; with the embedding service downit fails with a clear message and the old index stays
503when switched off or when the service is down,200withfive sections,
400for a blank question or more than 500 characters,200with a stale token headerthe real embedding service