This repository processes a Wikidata dump with WikidataTextifier, produces text chunks, embeds them with Jina, pushes vectors to Astra DB, and can publish both cleaned Wikidata JSON and cached vectors to Hugging Face datasets.
main.py is the entrypoint and supports:
- Download dump (if missing or
FORCE_DOWNLOAD_DUMP=true) - First pass (optional): save labels for all entities to Textifier MariaDB (
SAVE_LABELS=true) - Second pass (optional): process entities for:
- cleaned Wikidata JSON -> Hugging Face (
SAVE_WD_TO_HF=true) - vectorization -> Astra DB + local vector cache (
SAVE_TO_VECTORDB=true)
- cleaned Wikidata JSON -> Hugging Face (
- Optional post-step: publish local cached vectors to Hugging Face (
SAVE_VECTORS_TO_HF=true) - Optional multi-language loop with
WD_LANGS
When WD_LANGS is set and SAVE_LABELS=true, labels are saved once, then second-pass processing runs once per language.
- has label in
WD_LANG,FALLBACK_LANG, ormul - has claims or description
- is not disambiguation (
P31 != Q4167410) - has at least one sitelink ending in
wiki
The vector DB pass also writes a separate no-sitelink item collection for Q* items that pass the basic item filter, have no sitelink ending in wiki, and are not scholarly articles.
- has label in
WD_LANG,FALLBACK_LANG, ormul - has claims or description
- drop values of datatype:
external-id,commonsMedia,url,geo-shape,tabular-data - keep one value for
monolingualtext - for properties (
P*), drop claim-properties inPROPERTY_CONSTRAINT_PIDS(default:P2302)
- Text is chunked to max tokenizer length (
1024tokens) - Metadata stored with each vector chunk:
QIDorPIDChunkIDLanguageLabelDescriptionInstanceOfPropertiesLastModifiedDumpDate
Credentials are read from environment variables by the entrypoint and passed to
the service classes. For local runs, keep them in .env and let your runner
load them.
JINA_API_KEYASTRA_DB_APPLICATION_TOKENASTRA_DB_API_ENDPOINTASTRA_COLLECTION_PREFIXWD_HF_TOKENandWD_HF_REPO_IDVECTORS_HF_TOKENandVECTORS_HF_REPO_ID
| Variable | Default | Description |
|---|---|---|
SAVE_LABELS |
false |
Run first pass: save all entity labels to Textifier DB |
SAVE_WD_TO_HF |
false |
Run second pass: publish cleaned Wikidata JSON to HF |
SAVE_TO_VECTORDB |
false |
Run second pass: embed and push to Astra DB + local vector cache |
SAVE_VECTORS_TO_HF |
false |
Publish local cached vectors to HF |
SAVE_SITELINK_VECTORS |
true |
Include Wikipedia-sitelink items and properties in vector DB/cache/HF stages |
SAVE_NOSITELINK_VECTORS |
true |
Include non-scholarly items without Wikipedia sitelinks in vector DB/cache/HF stages |
DELETE_STALE_VECTORS |
false |
Prompt to delete vectors absent from the current dump pass |
FORCE_DOWNLOAD_DUMP |
false |
Force re-download of dump |
| Variable | Default | Description |
|---|---|---|
WD_LANG |
en |
Active language for single-language run |
FALLBACK_LANG |
WD_LANG |
Fallback language |
WD_LANGS |
empty | Comma-separated languages for per-language second pass loop |
FALLBACK_LANG_<LANG> |
unset | Per-language fallback override, example: FALLBACK_LANG_DE=en |
| Variable | Default | Description |
|---|---|---|
DUMP_PATH |
data/wd_dump.gz |
Dump file path |
READER_QUEUE_SIZE |
128 |
Reader queue size (in batches) |
READER_BATCH_SIZE |
16 |
Lines per queue batch |
HF_CHUNK_SIZE |
10000 |
Rows per HF upload chunk |
DUMP_DATE |
dump sidecar/file date | Metadata dump date |
PROPERTY_CONSTRAINT_PIDS |
P2302 |
Comma-separated claim-property IDs to drop for P* textification |
| Variable | Default | Description |
|---|---|---|
HF_BRANCH |
dump date (YYYYMMDD) |
Branch for cleaned WD dataset uploads |
VECTOR_HF_BRANCH |
HF_BRANCH |
Branch for vector dataset uploads |
WD_HF_TOKEN |
none | Hugging Face token for cleaned WD dataset uploads |
WD_HF_REPO_ID |
none | Hugging Face repo ID for cleaned WD dataset uploads |
VECTORS_HF_TOKEN |
none | Hugging Face token for vector dataset uploads |
VECTORS_HF_REPO_ID |
none | Hugging Face repo ID for vector dataset uploads |
MERGE_HF_TO_MAIN |
false |
Merge the configured HF branch into main instead of publishing chunks |
All languages use the configured vector dataset branch.
| Variable | Default | Description |
|---|---|---|
WIKIBASE_HOST |
wikibase |
Host checked before dump processing |
WIKIBASE_PORT |
80 |
Wikibase port |
TEXTIFIER_DB_HOST |
db |
Label DB host checked before dump processing |
TEXTIFIER_DB_PORT |
3306 |
Label DB port |
DB_HOST |
db |
Textifier DB host used by Textifier internals |
DB_PORT |
3306 |
Textifier DB port |
DB_NAME |
none | Textifier label DB name |
DB_USER |
none | Textifier DB user |
DB_PASS |
none | Textifier DB password |
| Variable | Default | Description |
|---|---|---|
WDTEXTIFIER_REPO |
https://github.com/wmde/WikidataTextifier.git |
WikidataTextifier Git repo to clone/update |
WDTEXTIFIER_REF |
main |
Branch/tag/commit checked out in WikidataTextifier |
Install dependencies:
uv sync --lockedExample: labels pass only
SAVE_LABELS=true \
SAVE_WD_TO_HF=false \
SAVE_TO_VECTORDB=false \
SAVE_VECTORS_TO_HF=false \
uv run python main.pyExample: second pass to Astra only
SAVE_LABELS=false \
SAVE_WD_TO_HF=false \
SAVE_TO_VECTORDB=true \
SAVE_VECTORS_TO_HF=false \
WD_LANG=en \
FALLBACK_LANG=en \
uv run python main.pyExample: publish cached vectors to HF only
SAVE_LABELS=false \
SAVE_WD_TO_HF=false \
SAVE_TO_VECTORDB=false \
SAVE_VECTORS_TO_HF=true \
WD_LANG=en \
uv run python main.pydocker-compose.yml in this repository is pipeline-only.
It does not start db / wikibase / wdtextifier by itself.
Use the helper script to run everything end-to-end:
- clone/update
WikidataTextifier - checkout configured ref (
WDTEXTIFIER_REF) - start Textifier stack (
db,wikibase,wdtextifier) - wait for
wdtextifierhealth - run this repo's
pipelinecontainer on the same Docker network
./scripts/run_pipeline_with_wdtextifier.shIf you want to run manually:
- Start Textifier stack:
docker compose \
-p wikidatatextifier \
-f /path/to/WikidataTextEmbedding/WikidataTextifier/docker-compose.yml \
--env-file /path/to/WikidataTextEmbedding/.env \
up -d db wikibase wdtextifier- Run pipeline container:
WDTEXTIFIER_COMPOSE_NETWORK=wikidatatextifier_default \
docker compose \
-p wikidatatextembedding-pipeline \
-f /path/to/WikidataTextEmbedding/docker-compose.yml \
--env-file /path/to/WikidataTextEmbedding/.env \
run --rm pipeline- Astra DB collections:
{COLLECTION_PREFIX}_items_<lang>{COLLECTION_PREFIX}_items_nositelinks_<lang>{COLLECTION_PREFIX}_properties_<lang>
- Local vector cache SQLite files:
data/Wikidata/sqlite_wikidata_vectors_items_<lang>.dbdata/Wikidata/sqlite_wikidata_vectors_items_nositelinks_<lang>.dbdata/Wikidata/sqlite_wikidata_vectors_properties_<lang>.db
- Hugging Face dataset uploads:
- cleaned Wikidata rows under
data/(branchHF_BRANCH) - vectors under
data/<lang>/(branchVECTOR_HF_BRANCH)
- cleaned Wikidata rows under
- If all
SAVE_*flags arefalse, the run exits without processing. main.pychecks network reachability to Wikibase and label DB before processing dump passes.DELETE_STALE_VECTORS=trueprompts before deleting AstraDB documents and their local cache entries.- Hugging Face uploads run in a background uploader process and use temporary cache dirs that are cleaned after each chunk.