An encoder-only Transformer in PyTorch for binary classification of
HH -> bbWW signal and background events in the CMS Run-3 analysis. The model
operates on low-level particle information and is evaluated with a two-fold
(even/odd) event split.
The repository contains the training code, notebooks, exported ONNX models, configuration files, and the plots produced during evaluation.
The classifier receives 40 features describing:
- two leptons: momentum components, energy, PDG ID, and charge;
- four small-radius jets and one large-radius jet: momentum components, energy, and b-tag information;
- missing transverse energy (MET): momentum components and energy.
The signal class is HH (ggF and VBF samples); the background class combines
the configured TT, DY, and other background samples. Training uses a
Transformer encoder followed by a neural-network classification head.
| Path | Description |
|---|---|
transformer_training.py |
Command-line training and evaluation script |
models.py |
Transformer classifier definition |
datasets.py |
Dataset and event preprocessing utilities |
utils.py |
Training and evaluation helpers |
transformer_training.ipynb |
Interactive training workflow |
*_model/ |
Configurations, plots, and ONNX exports for a trained variant |
Install the Python dependencies:
pip install -r requirements.txtThe training script reads ROOT files from the Bamboo results directory. Supply the split name and adjust the event count, number of epochs, and input directory for the environment being used:
python transformer_training.py \
--split even_noise_N_1e5 \
--n_events 100000 \
--n_epochs 150 \
--noise_level 0.01 \
--bamboo_results_dir /path/to/bamboo/results/The script writes a <split>_model/ directory containing config.json,
training and evaluation plots, and ONNX exports. CUDA, Apple MPS, and CPU
execution are supported; the script selects the first available device.
| Output directory | Event split | Events | Noise |
|---|---|---|---|
even_model |
even | 100,000 | 0.01 |
odd_model |
odd | 100,000 | not configured |
even_full_model |
even | 10,000,000 | 0.0 |
odd_full_model |
odd | 10,000,000 | 0.0 |
All four saved configurations use 150 epochs, learning rate 0.001, and
weight decay 0.0001. See each directory's config.json for the exact run
settings.
The ROC plot for the even split reports an AUC of 0.896.
| Training history | ROC curve |
|---|---|
![]() |
![]() |
Additional even-split evaluation plots:
| Score distribution | Calibration |
|---|---|
![]() |
![]() |
| Confusion matrix | Permutation importance |
|---|---|
![]() |
![]() |
The score distributions after probability calibration are available in
even_model/model_test_dist_calibrated.png
and even_model/calibrated_predicted_scores.png.
| Training history | ROC curve |
|---|---|
![]() |
![]() |
| Training history | ROC curve |
|---|---|
![]() |
![]() |
| Training history | ROC curve |
|---|---|
![]() |
![]() |
Each model directory also contains score distributions, normalized event weights, confusion matrices, permutation-importance plots, and (where available) calibration plots.
Each trained variant includes:
model.onnxandmodel_simplified.onnx;- calibrated counterparts (
calibrated_model.onnxandcalibrated_model_simplified.onnx); - the JSON configuration used for the run.
The ONNX files can be loaded with ONNX Runtime for inference without importing the PyTorch training code.
See LICENSE.











