From 47d60c27901c380e4ce90adf940b599f6141051a Mon Sep 17 00:00:00 2001 From: Junwen Yin <39627445+junwen94@users.noreply.github.com> Date: Fri, 11 Sep 2026 11:28:08 +0100 Subject: [PATCH] docs: organise campaigns around the published records, drop the snapshot page A campaign page now exists to say how the calculations behind a published record were made and how to reproduce them. Two campaigns, three records: the SCF k-point sweep published 52713-55d86 and mcpnq-g1j55, which are two views of one body of calculations and are presented together, and the nscf band campaign published r3byg-xp284, which gets a page of its own. results.md is deleted. A dated snapshot of an in-flight campaign -- workchain counts, ultra rate, regenerated local files -- is our operational record, not something a reader of the site needs. The nscf record is live, so its placeholder link is replaced. Two accuracy fixes while reorganising: the SCF campaign page states that the criterion is the whole remaining tail, which is stricter than the "three consecutive" wording in 52713-55d86's own README; and the convergence reference now says its three thresholds come from campaign.yaml and differ from the package defaults. --- docs/campaigns/index.md | 22 +++++--- docs/campaigns/qe-kpoints.md | 96 +++++++++++++++++++++------------ docs/campaigns/qe-nscf-bands.md | 79 +++++++++++++++++++++++++++ docs/index.md | 23 ++++---- docs/published-records.md | 34 +++++++----- docs/reference/convergence.md | 9 +++- docs/results.md | 44 --------------- mkdocs.yml | 2 +- 8 files changed, 198 insertions(+), 111 deletions(-) create mode 100644 docs/campaigns/qe-nscf-bands.md delete mode 100644 docs/results.md diff --git a/docs/campaigns/index.md b/docs/campaigns/index.md index 0ab5273..5133a39 100644 --- a/docs/campaigns/index.md +++ b/docs/campaigns/index.md @@ -8,14 +8,20 @@ A campaign is one reproducible combination of: - an ordered parameter schedule; - submission and convergence rules. -Each campaign directory keeps its README, machine-readable settings, -submission scripts, analysis notebook, and result manifest together. +Every [published record](../published-records.md) comes out of one, and a +campaign page says what the settings were and how to reproduce the calculations +behind the record. -## Available campaigns +## Campaigns and what they published -| Code | Task | Status | Guide | -| --- | --- | --- | --- | -| Quantum ESPRESSO | No-spin SCF k-point convergence | Active | [Open](qe-kpoints.md) | +| Campaign | Published as | Guide | +| --- | --- | --- | +| QE no-spin SCF k-point convergence | [`52713-55d86`](https://data-collections.psdi.ac.uk/records/52713-55d86) · [`mcpnq-g1j55`](https://data-collections.psdi.ac.uk/records/mcpnq-g1j55) | [Open](qe-kpoints.md) | +| QE nscf band structures | [`r3byg-xp284`](https://data-collections.psdi.ac.uk/records/r3byg-xp284) | [Open](qe-nscf-bands.md) | -Only implemented datasets appear here. A code or task is not listed merely -because support is planned. +The two SCF records are two views of one body of calculations, not two +campaigns: one is the converged mesh per structure, the other the full +per-calculation output. + +Only campaigns that have published something appear here. A code or task is not +listed merely because support is planned. diff --git a/docs/campaigns/qe-kpoints.md b/docs/campaigns/qe-kpoints.md index 3d56e7d..6510a84 100644 --- a/docs/campaigns/qe-kpoints.md +++ b/docs/campaigns/qe-kpoints.md @@ -1,7 +1,13 @@ # QE SCF k-point convergence -This campaign determines the smallest gamma-inclusive k-point mesh at which -the total energy is stable for each structure. +This campaign determines the smallest gamma-inclusive k-point mesh at which the +total energy of a structure is stable. + +Published as two records, from one body of calculations: +[`52713-55d86`](https://data-collections.psdi.ac.uk/records/52713-55d86), the +converged mesh per structure, and +[`mcpnq-g1j55`](https://data-collections.psdi.ac.uk/records/mcpnq-g1j55), the +full per-calculation output behind the paper. ## Campaign definition @@ -9,58 +15,80 @@ the total energy is stable for each structure. | --- | --- | | Code | Quantum ESPRESSO `pw.x` | | Task | No-spin SCF | -| Pseudopotentials | PseudoDojo 0.4 PBEsol SR standard | -| Schedule | Distinct gamma-inclusive meshes indexed by kindex | -| Extension | Three new kindex points per structure | -| Capacity | At most 50 active or newly submitted WorkChains | -| Check interval | 15 minutes | +| Functional | PBEsol | +| Pseudopotentials | SSSP, for both published records | +| Schedule | The 1-based gamma-inclusive k-mesh ladder, floor `min_k_distance = 0.03` Å⁻¹ | +| Criterion | Energy only, per atom, over a tail of at least three meshes | -The machine-readable settings live in -[`campaign.yaml`](https://github.com/stfc/goldilocks-data/blob/main/campaigns/qe/kpoints/campaign.yaml). +The pseudopotential family is the one axis that moves between runs of this +campaign; everything else above defines it. The machine-readable settings of the +current run live in +[`campaign.yaml`](https://github.com/stfc/goldilocks-data/blob/main/campaigns/qe/kpoints/campaign.yaml), +so read the pseudopotential family from the record you are reproducing rather +than from that file. -## How one cycle works +## Reproducing it -1. load the saved analysis snapshot; -2. query AiiDA for active and already-existing WorkChains; -3. select clean structures that reached `well` but not `ultra`; -4. allocate complete groups of three new kindex points within the capacity; -5. submit only when `--execute` is present; -6. wait 15 minutes and query AiiDA again. +1. **Take the structures from the record.** `CIF_files/.cif`, and + `source_db_id` is the join key to every table. -The snapshot selects the candidate cohort. Live AiiDA queries prevent duplicate -submissions and enforce the active-workchain limit. +2. **Rebuild the ladder.** `k_index` is reproducible from the structure alone — + no calculation needed — provided you use the same floor: -## Preview one cycle + ```python + from pymatgen.core import Structure + from goldilocks_data.kmesh import build_gamma_kmesh_entries -```bash -uv run --extra aiida --extra kmesh python \ - campaigns/qe/kpoints/scripts/monitor.py \ - --once \ - --cif-dir /path/to/CIF_files -``` + entries = build_gamma_kmesh_entries(Structure.from_file("100115.cif")) + print([(e.kindex, e.mesh) for e in entries[:5]]) + ``` + + A `k_index` computed under a different floor is comparable only where the + two ranges overlap. See [k-mesh quantities](../reference/kmesh.md). -## Run the controller +3. **Run the sweep.** One SCF per rung, unshifted mesh, holding everything else + fixed. The package submits these through AiiDA with the structure identifier + and the sweep point recorded on each node — see the submit example in the + [repository README](https://github.com/stfc/goldilocks-data#submit-example). + +4. **Label convergence.** `goldilocks_data.analysis.convergence` takes the + finished energies and returns the smallest rung whose remaining tail holds + within the threshold. See [convergence criteria](../reference/convergence.md). + +!!! warning "The criterion is the whole tail, not three points" + + `52713-55d86`'s README describes the converged mesh as the first of *three + consecutive* meshes agreeing within 1 meV per atom. The code is stricter + than that wording: the oscillation is `max - min` over **every** rung from + that point to the end of the sweep, with at least three rungs required. A + reimplementation that stops after checking three will label some structures + converged that this campaign did not. + +## Running it at scale + +A full campaign is a loop: submit a bounded number of WorkChains, wait, query +AiiDA for what finished, extend the structures that have not converged yet. +`monitor.py` does that. Preview one cycle: ```bash uv run --extra aiida --extra kmesh python \ campaigns/qe/kpoints/scripts/monitor.py \ - --execute \ + --once \ --cif-dir /path/to/CIF_files ``` -Stop the loop with `Ctrl-C`. The controller finishes the current synchronous -cycle before waiting for the next one. +Submit for real by replacing `--once` with `--execute`. Stop the loop with +`Ctrl-C`; it finishes the current cycle first. + +Live AiiDA queries, not a local file, are what prevent duplicate submissions and +enforce the active-workchain limit. !!! warning Always run a one-cycle dry run after changing the profile, group, pseudo - family, snapshot, or source structures. + family, or source structures. ## Source files - [Task README](https://github.com/stfc/goldilocks-data/tree/main/campaigns/qe/kpoints) - [Submission scripts](https://github.com/stfc/goldilocks-data/tree/main/campaigns/qe/kpoints/scripts) -- [Analysis notebook](https://github.com/stfc/goldilocks-data/blob/main/campaigns/qe/kpoints/notebooks/analysis.ipynb) -- [Result manifest](https://github.com/stfc/goldilocks-data/blob/main/campaigns/qe/kpoints/results/manifest.json) - -Continue with [Results](../results.md). diff --git a/docs/campaigns/qe-nscf-bands.md b/docs/campaigns/qe-nscf-bands.md new file mode 100644 index 0000000..e997262 --- /dev/null +++ b/docs/campaigns/qe-nscf-bands.md @@ -0,0 +1,79 @@ +# QE nscf band structures + +This campaign computes the band structure of each MC3D structure along its +high-symmetry k-point path, and reads the Fermi level, the metallicity and the +band gap off it. + +Published as +[`r3byg-xp284`](https://data-collections.psdi.ac.uk/records/r3byg-xp284) — +19,405 structures, with the eigenvalues themselves, not only the derived +numbers. + +## Campaign definition + +| Setting | Value | +| --- | --- | +| Code | Quantum ESPRESSO `pw.x` | +| Task | nscf band structure, from the charge density of a self-consistent parent | +| Functional | PBEsol | +| Pseudopotentials | SSSP | +| Spin | None | +| Cell | Primitive, as SeeKpath standardises it | +| k-point path | SeeKpath, through `PwBandsWorkChain` | + +Same settings and the same family of structures as the +[SCF k-point campaign](qe-kpoints.md), so a structure's two records describe the +same material computed the same way. + +## Reproducing it + +The record is self-contained: it carries the cell each calculation ran on, so +you do not have to re-derive the structure from MC3D to compare against it. + +1. **Take the structure from the record.** `CIF_files/.cif` is the + primitive cell the bands were computed on, not a conventional cell that then + needs standardising. + +2. **Generate the path with SeeKpath.** The workflow is the standard + [`PwBandsWorkChain`](https://aiida-quantumespresso.readthedocs.io/en/stable/howto/run_pwbands.html), + which finds the primitive cell and builds the high-symmetry path. Its + `bands_kpoints_distance` is what controls how densely the path is sampled: + the built-in protocols use 0.1 (`fast`), 0.025 (`balanced`) and 0.015 + (`stringent`) Å⁻¹. + +3. **Check the path before comparing eigenvalues.** The published + `bands/.csv` carries `kx, ky, kz` and the high-symmetry + `label`, so a regenerated path can be compared point for point. Do that + first — two paths that differ in sampling density are not comparable row by + row even when both are correct. + +4. **Run nscf on the parent charge density.** Eigenvalues are in eV in the + published tables; QE writes them in eV as well, so no conversion is involved. + +!!! note "The path density is not one number across the record" + + The archive was not produced with a single `bands_kpoints_distance`, so the + density was determined per structure. Take it from the published k-point + coordinates rather than assuming a protocol default. The + [record page](../published-records.md#quantum-espresso-nscf-band-structures) + explains how those coordinates were reconstructed and on what evidence each + one was accepted. + +## Deriving the labels yourself + +`metallicity` and `band_gap_ev` are both derived from the eigenvalues and the +Fermi energy that the record publishes, so neither has to be taken on trust. The +threshold used for `metallicity` is a plain zero — see +[the record page](../published-records.md#metallicity-is-a-zero-threshold-not-a-physical-one) +for why no thresholded label is published, and pick your own cut on +`band_gap_ev` if your work needs one. + +## Start here + +`example.py` in the record loads the summary, the structures and the bands +straight out of the tarballs and plots one band structure: + +```bash +python example.py # the summary, then a plotted band structure +python example.py 100115 # one structure by source_db_id +``` diff --git a/docs/index.md b/docs/index.md index ee409e3..e94c551 100644 --- a/docs/index.md +++ b/docs/index.md @@ -28,18 +28,21 @@ Model training belongs in [Goldilocks ML](https://stfc.github.io/goldilocks-ml/). End-user input generation belongs in Goldilocks Core. -## Current data campaign +## What has been published -The first campaign measures **Quantum ESPRESSO no-spin SCF k-point -convergence** with a gamma-inclusive kindex schedule. It compares PseudoDojo -and SSSP PBEsol pseudopotentials. +Three records, from two campaigns, all on MC3D structures with Quantum ESPRESSO: -The campaign extends an unconverged structure by three meshes at a time, -records every calculation in AiiDA, and exports one summary row per structure. +| Campaign | Records | +| --- | --- | +| [No-spin SCF k-point convergence](campaigns/qe-kpoints.md) | the converged mesh per structure, and the full per-calculation output behind the paper | +| [nscf band structures](campaigns/qe-nscf-bands.md) | eigenvalues along the high-symmetry path, with Fermi level, metallicity and band gap | + +Each [record](published-records.md) says what its columns mean and which +conventions it froze; each campaign page says how to reproduce the calculations +behind it. ## Where to go -[Install the environment](installation/index.md){ .md-button .md-button--primary } -[Run the k-point campaign](campaigns/qe-kpoints.md){ .md-button } -[Use the results](results.md){ .md-button } -[Published records](published-records.md){ .md-button } +[Published records](published-records.md){ .md-button .md-button--primary } +[Install goldilocks-data](installation/package.md){ .md-button } +[Data campaigns](campaigns/index.md){ .md-button } diff --git a/docs/published-records.md b/docs/published-records.md index c0c46e9..ea0828a 100644 --- a/docs/published-records.md +++ b/docs/published-records.md @@ -7,13 +7,18 @@ identifier and can be cited. These are snapshots. The AiiDA database remains the authoritative calculation record; a published dataset is a documented view of it at one point in time. -## Quantum ESPRESSO no-spin SCF calculations (SSSP, k-index) +## Quantum ESPRESSO no-spin SCF k-point convergence + +Two records, one body of calculations. Same structures, same settings, same +`pw.x` runs — they differ in what was extracted. Produced by the +[QE SCF k-point campaign](campaigns/qe-kpoints.md). + +### `52713-55d86` — the converged mesh per structure [`52713-55d86`](https://data-collections.psdi.ac.uk/records/52713-55d86) · v2.0 · CC BY 4.0 -The current SSSP k-index dataset: the converged k-point mesh for 17,757 MC3D -structures, numbered on the **1-based** ladder (rung 1 the Γ-only `(1, 1, 1)` +The converged k-point mesh for 17,757 MC3D structures, numbered on the **1-based** ladder (rung 1 the Γ-only `(1, 1, 1)` mesh) and built with the resolution floor `min_k_distance = 0.03` Å⁻¹ rather than a per-axis k-point cap. No spin polarisation, SSSP PBEsol pseudopotentials, every mesh unshifted and therefore gamma-inclusive. @@ -37,39 +42,42 @@ rather than assuming. See [convergence criteria](reference/convergence.md) for how labels are assigned, and the record's own `README.md` for the full definition and reproduction code. -## Quantum ESPRESSO no-spin SCF calculations (SSSP, k-distance) +### `mcpnq-g1j55` — the full per-calculation dump [`mcpnq-g1j55`](https://data-collections.psdi.ac.uk/records/mcpnq-g1j55) · CC BY 4.0 The complete DFT data behind *Automatic generation of input files with optimised k-point meshes for Quantum ESPRESSO self-consistent field total energy -calculations* — the training set for the paper's machine-learning models. Same -family of calculations as the k-index record above, expressed as a k-distance, -and it is the raw per-calculation dump rather than a convergence-label table. +calculations* — the training set for the paper's machine-learning models. The +same calculations as the record above, expressed as a k-distance, and a raw +per-calculation dump rather than a convergence-label table. Take this one when +you want the QE outputs themselves: cutoffs, Fermi level, symmetry counts, wall +time. | File | Contents | | --- | --- | +| `data.tar.gz` | Unpacks to the two entries below | | `summary.csv` | Per-material k-point convergence: Goldilocks-optimised mesh, MC3D reference mesh, k-distance metrics, and medium / well / ultra levels | | `structure_calc_details/` | One directory per structure: the `.cif` plus the full QE output (energies, cutoffs, Fermi level, symmetry counts, wall time, convergence notes) | This record predates the k-mesh ladder convention and carries meshes and k-distances directly, not a `k_index`. -## Quantum ESPRESSO band structures (MC3D, nscf) +## Quantum ESPRESSO nscf band structures - -*Record link to be added — the deposit is with PSDI and not yet published.* · +[`r3byg-xp284`](https://data-collections.psdi.ac.uk/records/r3byg-xp284) · v1 · CC BY 4.0 +Produced by the [QE nscf band campaign](campaigns/qe-nscf-bands.md). + Non-self-consistent band-structure calculations for **19,405 MC3D structures**: one summary row per material, the primitive cell the bands were computed on, and the eigenvalues along the high-symmetry k-point path. Every structure has all three — there is no row without a structure file and none without a band table. -Same settings and the same family of structures as the k-index record above. -That one answers *which mesh is dense enough*; this one answers *what the band +Same settings and the same family of structures as the SCF records above. +Those answer *which mesh is dense enough*; this one answers *what the band structure says about the material*. | File | Contents | diff --git a/docs/reference/convergence.md b/docs/reference/convergence.md index a812e50..e57a0e8 100644 --- a/docs/reference/convergence.md +++ b/docs/reference/convergence.md @@ -10,7 +10,14 @@ calculations. | Ultra | 1 meV/atom | For each label, the selected kindex is the smallest point whose remaining tail -satisfies the threshold. +satisfies the threshold. The oscillation is `max - min` over the whole tail, so +a longer sweep can only make a label harder to earn, never easier. + +These three numbers are the campaign's, from +[`campaign.yaml`](https://github.com/stfc/goldilocks-data/blob/main/campaigns/qe/kpoints/campaign.yaml). +`ConvergenceThresholds` in the package defaults to 10 / 5 / 1 meV per atom, so +pass the thresholds explicitly when reproducing a record rather than relying on +the defaults. ## Kindex meaning diff --git a/docs/results.md b/docs/results.md deleted file mode 100644 index 4d32317..0000000 --- a/docs/results.md +++ /dev/null @@ -1,44 +0,0 @@ -# Results - -A dated view of the PseudoDojo QE SCF k-point campaign. The AiiDA group remains -the authoritative calculation record. - -## Current snapshot - -| Measure | Value | -| --- | ---: | -| WorkChains | 127,921 | -| Structures | 16,208 | -| Ultra converged | 15,474 | -| Ultra rate | 95.47% | -| Median ultra kindex | 4 | - -Snapshot date: 1 September 2026. These figures are taken from -[`snapshot-metadata.json`](https://github.com/stfc/goldilocks-data/blob/main/campaigns/qe/kpoints/results/snapshot-metadata.json). - -## The snapshot files - -| File | Use it for | Where | -| --- | --- | --- | -| `snapshot-metadata.json` | Profile, AiiDA group, thresholds, ladder convention, timestamp, counts | [in the repo](https://github.com/stfc/goldilocks-data/blob/main/campaigns/qe/kpoints/results/snapshot-metadata.json) | -| `manifest.json` | Dataset identity and the file names | [in the repo](https://github.com/stfc/goldilocks-data/blob/main/campaigns/qe/kpoints/results/manifest.json) | -| `source-summary.csv` | One convergence-summary row per structure | regenerated locally | -| `workchain-records.parquet` | Calculation-level values behind the summary | regenerated locally | -| `analysis.ipynb` | Aggregation, quality checks, SSSP comparison | [in the repo](https://github.com/stfc/goldilocks-data/blob/main/campaigns/qe/kpoints/notebooks/analysis.ipynb) | - -`source-summary.csv` and `workchain-records.parquet` are rebuilt from the AiiDA -group every campaign cycle and are not versioned; see -[`campaigns/qe/kpoints/results/README.md`](https://github.com/stfc/goldilocks-data/blob/main/campaigns/qe/kpoints/results/README.md). -They go to PSDI as a citable record when the campaign finishes. - -## Provenance - -The snapshot metadata identifies the profile and AiiDA group used for export. -For the live PseudoDojo campaign, the group is: - -```text -goldilocks/qe-scf/nospin/pseudodojo -``` - -Return to AiiDA when you need the original inputs, outputs, process state, or -provenance graph. diff --git a/mkdocs.yml b/mkdocs.yml index 51bac31..4fb2068 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -52,7 +52,7 @@ nav: - Data campaigns: - Overview: campaigns/index.md - QE SCF k-points: campaigns/qe-kpoints.md - - Results: results.md + - QE nscf bands: campaigns/qe-nscf-bands.md - Publishing: - Published records: published-records.md - Publish a dataset: publishing.md