Add embedding service for the portfolio assistant - #7
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stage 2 of the portfolio assistant. A small FastAPI service that turns texts
into 768-dimensional vectors with
intfloat/multilingual-e5-base. Nothingcalls it yet, and it is not deployed.
Why a separate service
The model needs about 1.8 GB of memory. Inside Django, each of the three
Gunicorn workers would load its own copy. As its own process it loads once,
and PyTorch stays out of the Django dependencies and the CI.
Changes
embedding_service/main.py:POST /embedandGET /healthembedding_service/requirements.txt: pinned dependencies, CPU-only PyTorchthrough
--extra-index-urlembedding_service/smoke_test.py: checks a running service, standardlibrary only
embedding_service/README.md: setup, API and design notesDesign decisions
kind(queryorpassage). A missing prefix does not raise an error, it only lowersretrieval quality, so the field is required.
definstead ofasync deffor/embed. Encoding is CPU-bound andwould block the event loop, including
/health.setuptoolsfrom PyPI, not from the PyTorch index. The PyTorch indexonly offers 78.1.0, which has known vulnerabilities.
Verification
/embedreturns 768 values per text, each vector with length 1.0kind, blank texts, an empty list, more than 64 texts and textsover 512 tokens return
422unrelated passage 0.754
pip install --dry-run --ignore-installed -r requirements.txtresolves totorch-2.14.0+cpuandsetuptools-84.0.0flake8andisortpass