Code and artifacts for Same Targets, Different Computation: how post-training divides work across model layers.
language-models reproducibility fine-tuning large-language-models instruction-tuning mechanistic-interpretability activation-patching llm-interpretability model-diffing cross-patching
-
Updated
Aug 11, 2026 - Python