Skip to content

Latest commit

 

History

87 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Agoge Forger

Agoge Forger is a Python/PyTorch-first post-training research platform for reproducible SFT, evaluation, checkpoints, experiment manifests, and Hugging Face model releases. It is designed for local RTX 5080 development with Transformers, TRL, PEFT, bitsandbytes, and vLLM-compatible model export.

Repository boundaries

Agoge owns training, evaluation, checkpoints, manifests, and releases. It consumes versioned inputs rather than duplicating other repositories:

Responsibility Owner
Synthetic generation and public-dataset admission rmems/synthetic-factory
Real engineering trajectories and SFT/DPO datasets rmems/operation-prometheus
Custom CUDA kernels and GPU-performance investigations rmems/blackwell-kernel-lab
Terraform, cloud jobs, costs, and provider runbooks rmems/Dioscuri-Cloud

There are no first-party CUDA, Rust, Julia, JAX, or cloud-infrastructure implementation trees in this repository.

Quickstart

uv sync --all-groups
uv run agoge check-torch
uv run agoge train-qlora --config configs/minicpm5_canary.yaml

The canary is openbmb/MiniCPM5-1B-Base; the future measured flagship is ibm-granite/granite-4.1-3b-base. Pin an immutable Hub revision and complete the compatibility, split, and evaluation contracts before inspecting measured results.

To inspect an architecture before choosing LoRA targets:

uv run agoge model-metadata --model-id openbmb/MiniCPM5-1B-Base
uv run agoge inspect-lora-targets --model-id openbmb/MiniCPM5-1B-Base

trust_remote_code defaults to false; enable it only after an explicit architecture review.

Artifact model

checkpoint-* directories are trainer recovery snapshots. The adapter saved at adapters/<run_name> is the LoRA output for continued PEFT work. The merged model under merged/<run_name> is the single final artifact to ship or evaluate as a standalone model.

Every training run records a reproducibility manifest, GPU telemetry, and artifact digests. Safetensors is the default artifact format; path validation rejects traversal before model, dataset, adapter, checkpoint, and output paths are used.

Optional bounded CUDA profile windows are off unless you pass --profile-window or AGOGE_PROFILE_WINDOW. See Bounded CUDA profiling windows and F0 correlation markers. The profiler is not a requirement for ordinary Agoge runs.

vLLM compatibility smoke

uv run agoge serve-vllm --model merged/<run_name>
uv run agoge smoke-vllm --model merged/<run_name> --run-name smoke_<run_name>

See vLLM model compatibility for supported artifact forms.

Inspect run readiness

Before resuming or exporting, ask a run directory what it is actually ready for:

agoge run-status adapters/<run_name>
agoge run-status adapters/<run_name> --format table
agoge run-status adapters/<run_name> --merged-dir merged/<custom_name>
agoge run-status adapters/<run_name> --allow-unsafe-serialization

The default --format json report, for a run with two checkpoints, a saved adapter, and an exported merge (absolute paths shortened, some fields elided):

{
  "schema_version": 1,
  "run_name": "demo_run",
  "checkpoints": {
    "valid_count": 2,
    "steps": [50, 100],
    "latest_step": 100,
    "latest_path": "adapters/demo_run/checkpoint-100"
  },
  "final_adapter": { "present": true, "path": "adapters/demo_run" },
  "merged_model": { "present": true, "path": "merged/demo_run" },
  "base_model": "openbmb/MiniCPM5-1B-Base",
  "resume": {
    "ready": true,
    "checkpoint_path": "adapters/demo_run/checkpoint-100"
  },
  "export": {
    "ready": true,
    "source_path": "adapters/demo_run",
    "source_kind": "final_adapter"
  }
}

How to read it:

  • resume.checkpoint_path is exactly the snapshot agoge train-qlora would pick up with resume_from_latest_checkpoint: true.
  • export.source_path is exactly what agoge export-final-model would merge, and source_kind says which kind it is: final_adapter means the run-root adapter, checkpoint means the latest valid checkpoint-N.
  • checkpoints.valid_count counts only checkpoints that pass the validity rules — a checkpoint-N directory holding both trainer_state.json and a safetensors adapter. A half-written snapshot is not counted.
  • merged_model probes the conventional merged/<run_name> sibling of adapters/<run_name>, or exactly --merged-dir when you pass it. "present": false means "not exported yet", not an error.
  • Legacy .bin adapters read as absent and not ready under the safetensors-only policy, until --allow-unsafe-serialization is passed.
  • The exit code is 0 for any inspectable directory, including one where nothing is ready yet. It is non-zero when the path is missing, is not a directory, contains .., cannot resolve a ~user home, or inspection hits a permission/I/O failure while building the report.

The report does not materialize tensor payload data, so it needs no GPU or network. It does use the installed PyTorch restricted weights_only unpickler to inspect bounded, memory-mapped optimizer.pt, scheduler.pt, and rng_state.pth metadata. Transformers' NumPy compatibility globals are enabled only while validating rng_state.pth; legacy adapter_model.bin inspection still requires --allow-unsafe-serialization. Keep the locked PyTorch version current with upstream security patch releases before inspecting untrusted run directories, because weights_only reduces pickle risk but is not equivalent to the non-executable safetensors format.

Reclaim run disk

checkpoint-* trees are trainer recovery snapshots. Once a run has produced a final adapter or a merged model they are dead weight, and save_total_limit only caps them during training. After exporting, prune them:

agoge export-final-model --run-dir adapters/<run_name> --out-dir merged/<run_name>
agoge cleanup-run adapters/<run_name> --dry-run   # lists candidates, deletes nothing
agoge cleanup-run adapters/<run_name>
agoge cleanup-run adapters/<run_name> --keep-latest 1   # leave one resume point
agoge cleanup-run adapters/<run_name> --format json | jq .bytes_reclaimed

Always look at --dry-run first: it reports every candidate directory and the estimated bytes reclaimed without touching the filesystem.

How it protects you:

  • It refuses to delete anything unless a final adapter or a merged model already exists, because otherwise the checkpoints are the only recoverable artifact. --force overrides that and says plainly what becomes unrecoverable.
  • Only directories named checkpoint-<step> are ever candidates, so the final adapter, tokenizer files, artifact_index.json, the runs/<run_name>/ manifests, and merged/<run_name> are out of reach by construction.
  • --keep-latest N keeps the N newest valid checkpoints. A half-written snapshot from a crashed run never occupies one of those slots — it is reclaimable garbage, not a resume point.
  • A symlinked run directory is refused — every path component is checked — and a symlinked or bind-mounted checkpoint-* entry is skipped rather than followed.

Run cleanup before publishing an evaluation contract, not after. The artifact index written at the end of training hashes every file in the run directory, checkpoints included, and the evaluation contract requires the index to match the files actually present. So when a checkpoint was actually removed, an existing artifact_index.json is rewritten over the survivors — which changes its sha256 and invalidates any contract already pinning it. A run where nothing was deleted leaves the index byte-identical. artifact_index_rewritten in the report says whether that happened. A run with no index is cleaned without inventing one. An index that already exists is rebuilt whether or not it carried sealed provenance, so it never keeps listing files cleanup just deleted. A rewrite that fails is reported under failed with a non-zero exit rather than passing silently.

Held-out base-vs-SFT canary

After a frozen split and an SFT adapter or merged artifact exist, compare base and SFT on the same held-out membership:

SPLIT_MANIFEST=/path/to/split_manifest.json
SFT_ARTIFACT=adapters/code-repair-canary
OUTPUT_DIR=eval/code-repair-canary
BASE_MODEL_ID=ibm-granite/granite-4.1-3b-base
BASE_REVISION=dacb9cb9157bec98e99b09f285c92a4d58405c96

uv run agoge held-out-eval \
  --split-manifest "$SPLIT_MANIFEST" \
  --sft-artifact "$SFT_ARTIFACT" \
  --output-dir "$OUTPUT_DIR" \
  --base-model-id "$BASE_MODEL_ID" \
  --base-revision "$BASE_REVISION" \
  --context-window 4096

See Frozen split and evaluation contracts for the bundle layout and local GPU steps. smoke-eval remains a toy hardcoded-prompt check and is not this harness.

Verify a reproducibility bundle offline

A sealed experiment directory can be checked without network access or loading weights. The command hashes membership, validates the run manifest, locked config, frozen split, artifact indexes, and both evaluation arms, then prints a machine-readable verdict:

uv run agoge verify-bundle /path/to/bundle
uv run agoge verify-bundle /path/to/bundle --format table

Unknown schema versions, truncated JSON, extra or missing files, mutated digests, duplicate or path-escaping inventory entries, and unsafe symlinks fail closed with exit 1. See Reproducibility bundle schema.

Validation

uv run ruff check .
uv run ruff format --check .
uv run mypy src/agoge_forger
uv run pytest tests/

The root Docker image is a CPU/smoke image. It installs from the locked dependency graph, runs as a non-root user, and never bakes in HF_TOKEN.

Optional CUDA profile windows

Ordinary train-qlora does not start a profiler. To capture one bounded forward/backward/optimizer window and emit joinable markers:

AGOGE_RUN_ID=run_minicpm5_profile_001 \
AGOGE_PROFILE_WINDOW=train:1-1 \
uv run agoge train-qlora --config configs/minicpm5_canary.yaml

See Bounded CUDA profiling windows.

License

Apache-2.0. See LICENSE.

About

PyTorch fine-tuning, experimenting with public SWE against engineering trajectories and synthetic datasets

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages