Agoge Forger is a Python/PyTorch-first post-training research platform for reproducible SFT, evaluation, checkpoints, experiment manifests, and Hugging Face model releases. It is designed for local RTX 5080 development with Transformers, TRL, PEFT, bitsandbytes, and vLLM-compatible model export.
Agoge owns training, evaluation, checkpoints, manifests, and releases. It consumes versioned inputs rather than duplicating other repositories:
| Responsibility | Owner |
|---|---|
| Synthetic generation and public-dataset admission | rmems/synthetic-factory |
| Real engineering trajectories and SFT/DPO datasets | rmems/operation-prometheus |
| Custom CUDA kernels and GPU-performance investigations | rmems/blackwell-kernel-lab |
| Terraform, cloud jobs, costs, and provider runbooks | rmems/Dioscuri-Cloud |
There are no first-party CUDA, Rust, Julia, JAX, or cloud-infrastructure implementation trees in this repository.
uv sync --all-groups
uv run agoge check-torch
uv run agoge train-qlora --config configs/minicpm5_canary.yamlThe canary is openbmb/MiniCPM5-1B-Base; the future measured flagship is ibm-granite/granite-4.1-3b-base. Pin an immutable Hub revision and complete the compatibility, split, and evaluation contracts before inspecting measured results.
To inspect an architecture before choosing LoRA targets:
uv run agoge model-metadata --model-id openbmb/MiniCPM5-1B-Base
uv run agoge inspect-lora-targets --model-id openbmb/MiniCPM5-1B-Basetrust_remote_code defaults to false; enable it only after an explicit architecture review.
checkpoint-* directories are trainer recovery snapshots. The adapter saved at adapters/<run_name> is the LoRA output for continued PEFT work. The merged model under merged/<run_name> is the single final artifact to ship or evaluate as a standalone model.
Every training run records a reproducibility manifest, GPU telemetry, and artifact digests. Safetensors is the default artifact format; path validation rejects traversal before model, dataset, adapter, checkpoint, and output paths are used.
Optional bounded CUDA profile windows are off unless you pass --profile-window or AGOGE_PROFILE_WINDOW. See Bounded CUDA profiling windows and F0 correlation markers. The profiler is not a requirement for ordinary Agoge runs.
uv run agoge serve-vllm --model merged/<run_name>
uv run agoge smoke-vllm --model merged/<run_name> --run-name smoke_<run_name>See vLLM model compatibility for supported artifact forms.
Before resuming or exporting, ask a run directory what it is actually ready for:
agoge run-status adapters/<run_name>
agoge run-status adapters/<run_name> --format table
agoge run-status adapters/<run_name> --merged-dir merged/<custom_name>
agoge run-status adapters/<run_name> --allow-unsafe-serializationThe default --format json report, for a run with two checkpoints, a saved
adapter, and an exported merge (absolute paths shortened, some fields elided):
{
"schema_version": 1,
"run_name": "demo_run",
"checkpoints": {
"valid_count": 2,
"steps": [50, 100],
"latest_step": 100,
"latest_path": "adapters/demo_run/checkpoint-100"
},
"final_adapter": { "present": true, "path": "adapters/demo_run" },
"merged_model": { "present": true, "path": "merged/demo_run" },
"base_model": "openbmb/MiniCPM5-1B-Base",
"resume": {
"ready": true,
"checkpoint_path": "adapters/demo_run/checkpoint-100"
},
"export": {
"ready": true,
"source_path": "adapters/demo_run",
"source_kind": "final_adapter"
}
}How to read it:
resume.checkpoint_pathis exactly the snapshotagoge train-qlorawould pick up withresume_from_latest_checkpoint: true.export.source_pathis exactly whatagoge export-final-modelwould merge, andsource_kindsays which kind it is:final_adaptermeans the run-root adapter,checkpointmeans the latest validcheckpoint-N.checkpoints.valid_countcounts only checkpoints that pass the validity rules — acheckpoint-Ndirectory holding bothtrainer_state.jsonand a safetensors adapter. A half-written snapshot is not counted.merged_modelprobes the conventionalmerged/<run_name>sibling ofadapters/<run_name>, or exactly--merged-dirwhen you pass it."present": falsemeans "not exported yet", not an error.- Legacy
.binadapters read as absent and not ready under the safetensors-only policy, until--allow-unsafe-serializationis passed. - The exit code is
0for any inspectable directory, including one where nothing is ready yet. It is non-zero when the path is missing, is not a directory, contains.., cannot resolve a~userhome, or inspection hits a permission/I/O failure while building the report.
The report does not materialize tensor payload data, so it needs no GPU or
network. It does use the installed PyTorch restricted weights_only unpickler
to inspect bounded, memory-mapped optimizer.pt, scheduler.pt, and
rng_state.pth metadata. Transformers' NumPy compatibility globals are enabled
only while validating rng_state.pth; legacy adapter_model.bin inspection
still requires --allow-unsafe-serialization. Keep the locked PyTorch version
current with upstream security patch releases before inspecting untrusted run
directories, because weights_only reduces pickle risk but is not equivalent
to the non-executable safetensors format.
checkpoint-* trees are trainer recovery snapshots. Once a run has produced a final adapter or a merged model they are dead weight, and save_total_limit only caps them during training. After exporting, prune them:
agoge export-final-model --run-dir adapters/<run_name> --out-dir merged/<run_name>
agoge cleanup-run adapters/<run_name> --dry-run # lists candidates, deletes nothing
agoge cleanup-run adapters/<run_name>
agoge cleanup-run adapters/<run_name> --keep-latest 1 # leave one resume point
agoge cleanup-run adapters/<run_name> --format json | jq .bytes_reclaimedAlways look at --dry-run first: it reports every candidate directory and the estimated bytes reclaimed without touching the filesystem.
How it protects you:
- It refuses to delete anything unless a final adapter or a merged model already exists, because otherwise the checkpoints are the only recoverable artifact.
--forceoverrides that and says plainly what becomes unrecoverable. - Only directories named
checkpoint-<step>are ever candidates, so the final adapter, tokenizer files,artifact_index.json, theruns/<run_name>/manifests, andmerged/<run_name>are out of reach by construction. --keep-latest Nkeeps the N newest valid checkpoints. A half-written snapshot from a crashed run never occupies one of those slots — it is reclaimable garbage, not a resume point.- A symlinked run directory is refused — every path component is checked — and a symlinked or bind-mounted
checkpoint-*entry is skipped rather than followed.
Run cleanup before publishing an evaluation contract, not after. The artifact index written at the end of training hashes every file in the run directory, checkpoints included, and the evaluation contract requires the index to match the files actually present. So when a checkpoint was actually removed, an existing artifact_index.json is rewritten over the survivors — which changes its sha256 and invalidates any contract already pinning it. A run where nothing was deleted leaves the index byte-identical. artifact_index_rewritten in the report says whether that happened. A run with no index is cleaned without inventing one. An index that already exists is rebuilt whether or not it carried sealed provenance, so it never keeps listing files cleanup just deleted. A rewrite that fails is reported under failed with a non-zero exit rather than passing silently.
After a frozen split and an SFT adapter or merged artifact exist, compare base and SFT on the same held-out membership:
SPLIT_MANIFEST=/path/to/split_manifest.json
SFT_ARTIFACT=adapters/code-repair-canary
OUTPUT_DIR=eval/code-repair-canary
BASE_MODEL_ID=ibm-granite/granite-4.1-3b-base
BASE_REVISION=dacb9cb9157bec98e99b09f285c92a4d58405c96
uv run agoge held-out-eval \
--split-manifest "$SPLIT_MANIFEST" \
--sft-artifact "$SFT_ARTIFACT" \
--output-dir "$OUTPUT_DIR" \
--base-model-id "$BASE_MODEL_ID" \
--base-revision "$BASE_REVISION" \
--context-window 4096See Frozen split and evaluation contracts for the bundle layout and local GPU steps. smoke-eval remains a toy hardcoded-prompt check and is not this harness.
A sealed experiment directory can be checked without network access or loading weights. The command hashes membership, validates the run manifest, locked config, frozen split, artifact indexes, and both evaluation arms, then prints a machine-readable verdict:
uv run agoge verify-bundle /path/to/bundle
uv run agoge verify-bundle /path/to/bundle --format tableUnknown schema versions, truncated JSON, extra or missing files, mutated
digests, duplicate or path-escaping inventory entries, and unsafe symlinks fail
closed with exit 1. See Reproducibility bundle schema.
uv run ruff check .
uv run ruff format --check .
uv run mypy src/agoge_forger
uv run pytest tests/The root Docker image is a CPU/smoke image. It installs from the locked dependency graph, runs as a non-root user, and never bakes in HF_TOKEN.
Ordinary train-qlora does not start a profiler. To capture one bounded
forward/backward/optimizer window and emit joinable markers:
AGOGE_RUN_ID=run_minicpm5_profile_001 \
AGOGE_PROFILE_WINDOW=train:1-1 \
uv run agoge train-qlora --config configs/minicpm5_canary.yamlSee Bounded CUDA profiling windows.
Apache-2.0. See LICENSE.