Skip to content

Add Optional toml-f Support For Generated Fortran Input Types #26

Description

@MuellerSeb

nml-tools currently generates one Fortran derived type per schema and exposes
methods such as init, from_file, set, is_set, is_valid, and
filled_shape where applicable.

The generated from_file path is currently hard-wired to Fortran namelist I/O.
It initializes the object, searches the input file for the first matching
namelist block named by x-fortran-namelist, reads that block, and then
applies the usual generated assignment and validation behavior.

This works well for namelist-based workflows, but it prevents using the same
schema-generated Fortran type with TOML input files. The missing capability is
an opt-in from_toml method, backed by toml-f, that reads TOML input which
follows the same schema model as the namelist reader.

Requirement

Add an opt-in generator path that emits a from_toml type-bound procedure for
generated namelist types.

The design should:

  • Keep the existing namelist-based from_file behavior unchanged.
  • Make TOML support optional so namelist-only users do not need toml-f.
  • Reuse the existing schema model rather than introducing a second schema
    dialect for TOML.
  • Reuse the existing generated assignment and validation logic as much as
    possible.
  • Define a precise TOML layout that remains compatible with the current
    per-schema generated-type model.

Compatibility Assessment

The current package layout is compatible with this feature, but only if the TOML
mapping is kept deliberately narrow.

What already fits

  • One schema already maps to one generated Fortran type.
  • One schema already owns one input group via x-fortran-namelist.
  • Combined input files already exist at the tooling level:
    • the CLI can validate a file against several schemas,
    • template generation can emit several namelist groups into one output file.
  • Validation is schema-driven and already expects an in-memory mapping of
    fields to values.

What the current generated API actually means

The generated type-level API is not "read an entire configuration document".
It is "read the input group belonging to this schema".

That distinction matters because a Fortran namelist file may contain several
different namelist groups, while one generated type still reads only the group
named by its own x-fortran-namelist.

The current from_file implementation therefore behaves like:

  • open file,
  • find first namelist block with the requested name,
  • read that block,
  • ignore unrelated blocks.

from_toml should preserve that mental model.

Recommended TOML mapping

Use one top-level TOML table per schema-owned input group.

If a schema declares:

x-fortran-namelist: optimization

then the corresponding TOML input should be:

[optimization]
niterations = 10
tolerance = 1.0e-6

For a document containing several schema-owned groups:

[physics]
dt = 60.0

[solver]
max_iter = 50
tolerance = 1.0e-8

This is the cleanest analogue to a namelist file with several differently named
namelist groups.

Multi-namelist compatibility

Compatibility is good for files containing several differently named namelists:

  • namelist side: one physical file may contain several named namelist blocks,
  • TOML side: one physical file may contain several named top-level tables,
  • generated-type side: one generated type still reads only the group/table it
    owns.

This means the correct analogue is not "one TOML file per generated type" but
"one TOML document may hold several generated-type tables".

Repeated same-name namelist blocks

Fortran namelist files can, in practice, contain repeated blocks with the same
name. However, that is not a strong part of the current generated-type contract,
because the helper scans the file and stops at the first matching block.

That means the current generated reader already behaves much more like
"read one matching group" than like "merge all groups with this name".

TOML has no natural equivalent for repeated same-name tables because duplicate
tables are invalid. This is acceptable for a first TOML implementation because
the current generated namelist reader does not provide meaningful repeated-block
semantics either.

Key and table name compatibility

This is the main semantic mismatch.

  • Current schema/property handling is case-insensitive.
  • Current namelist validation normalizes field names case-insensitively.
  • TOML is case-sensitive.
  • TOML forbids duplicate keys and duplicate tables.

To preserve one schema for both formats, TOML lookup should be defined as
case-insensitive relative to schema names and field names.

That implies two additional rules:

  • A TOML table should match x-fortran-namelist ignoring case.
  • TOML keys inside that table should match schema property names ignoring case.

Ambiguous TOML input should be rejected, for example if a table or key set
contains two entries that collide after lowercasing.

Array compatibility

Array support looks feasible, but only under the same restrictions that the
current validation and namelist tooling already assume. The TOML reader should
mirror Fortran namelist buffer-reading semantics as closely as possible:
provided array values fill the generated Fortran array storage sequence, and
unspecified entries keep the values established by init.

  • Arrays must remain arrays of scalars.
  • Nested schema arrays remain unsupported.
  • TOML arrays may be partial fills.
  • A TOML scalar assigned to an array field should set only the first Fortran
    array element.
  • A lower-rank TOML array should be accepted as a partial fill of the higher-rank
    Fortran array, following the same storage-order convention used by namelist
    reads.
  • TOML arrays with nested list structure must be rectangular within the provided
    data.
  • The current validation model interprets nested lists in Fortran order, where
    the outermost list corresponds to the last Fortran index.
  • TOML support should keep exactly that convention rather than introducing a
    second array ordering depending on input format.

This is slightly awkward for TOML users, but it preserves schema semantics,
matches the existing Fortran namelist behavior, and avoids silent
format-specific reshaping.

Ragged TOML arrays, mixed-type arrays, arrays with too many dimensions, or
partial fills that exceed the generated Fortran array extent should be rejected.
Schema validation through is_valid remains responsible for detecting whether a
partially filled array is acceptable, for example for required arrays or
flexible tail dimensions.

Nested tables

Nested TOML tables below the schema-owned top-level table do not fit the current
schema model unless they are only a syntactic alternative for flat keys.

The current schema supports object roots with scalar and array properties, not
nested objects / nested derived types. Therefore:

  • nested TOML tables should stay out of scope for the first implementation,
  • dotted TOML keys that effectively create nested tables should also stay out of
    scope unless they are explicitly mapped back to flat property names.

The first implementation should only support a flat table of scalar/array
entries under [<x-fortran-namelist>].

Proposed Interface

Add an opt-in generator setting per namelist entry, for example:

[[namelists]]
schema = "schema/optimization.yml"
mod_path = "src/generated/nml_optimization.f90"
toml_support = true

Generated Fortran usage:

type(nml_optimization_t) :: cfg
integer :: status
character(len=256) :: errmsg

status = cfg%from_toml("optimization.toml", errmsg)
if (status /= NML_OK) then
  print *, trim(errmsg)
end if

If f2py/Python wrappers are generated for the schema, they should optionally
expose matching from_toml wrappers as well.

Implementation Approach

The current generated from_file implementation should be refactored into a
small shared pipeline rather than copied.

Recommended structure:

  • Keep the existing helper module focused on namelist file operations.
  • Do not make the helper module depend on toml-f.
  • Emit TOML-specific imports and reader routines only when TOML support is
    enabled for a generated Fortran module.

Refactor generated module code into the following conceptual steps:

  • shared local declarations,
  • shared initialization via init,
  • backend-specific read into local variables,
  • shared assignment from local variables into this,
  • shared finalization and status handling.

TOML-specific reader behavior

The generated from_toml path should:

  • parse the TOML document using toml-f,
  • locate the top-level table matching x-fortran-namelist,
  • reject missing table with a deterministic non-OK status,
  • iterate over the table keys and reject unknown keys,
  • detect case-colliding keys/tables after normalization and reject them as
    ambiguous,
  • read only schema-declared scalar and array fields,
  • avoid relying on toml-f default insertion for schema values,
  • reuse the same generated assignment tail as from_file.

Defaults and missing values

The current generated type already encodes defaults and sentinel values via
init and later assignment behavior. TOML parsing should not invent a second,
independent defaulting system.

That means:

  • from_toml should initialize the object first,
  • absent TOML keys should behave like absent namelist entries,
  • required/default/optional behavior should continue to come from the schema and
    generated type logic, not from toml-f table mutation.

Status codes

A minimal initial implementation can likely reuse existing status codes:

  • NML_ERR_FILE_NOT_FOUND for missing file,
  • NML_ERR_OPEN / NML_ERR_CLOSE for file I/O failures,
  • NML_ERR_READ for TOML parse/type/shape/key errors,
  • NML_ERR_NML_NOT_FOUND for missing top-level TOML table, or a new generic
    alias if a clearer name is preferred.

There is no need to widen the public status code surface unless implementation
experience shows a real need.

Acceptance Criteria

  • nml-tools can optionally generate a from_toml type-bound procedure.
  • The feature is opt-in and does not affect namelist-only outputs.
  • The generated namelist-only code remains unchanged when TOML support is not
    enabled.
  • from_toml reads the top-level TOML table named by x-fortran-namelist.
  • A TOML file may contain additional top-level tables for other schemas without
    affecting the current generated type.
  • Missing expected table returns a deterministic non-OK status.
  • Unknown TOML keys return a non-OK status with a useful error message.
  • TOML keys/tables that collide after lowercasing are rejected.
  • Scalar values are read using the same schema types as namelist input.
  • Arrays use the same Fortran-order convention already used by current
    validation.
  • Array fields support namelist-like partial fills.
  • Scalar values assigned to array fields fill the first array element only.
  • Ragged or mixed-type TOML arrays are rejected.
  • TOML array data that exceeds the generated Fortran shape is rejected.
  • Flexible array tail dimensions behave consistently with the current generated
    semantics.
  • Generated f2py/Python wrappers can optionally expose from_toml.
  • README documents the TOML layout, case rules, array ordering, dependency
    requirements, and the distinction between per-type loading and combined input
    documents.
  • Tests cover:
    • single-table TOML input,
    • multi-table TOML input,
    • missing table,
    • unknown key,
    • case-collision rejection,
    • scalar type mismatch,
    • rectangular and ragged arrays,
    • flexible arrays,
    • wrapper generation when enabled.

Out Of Scope

  • Replacing or deprecating namelist input.
  • Automatic format detection inside from_file.
  • Supporting nested-object schemas via nested TOML tables.
  • TOML output generation or TOML template generation.
  • Root-table TOML input without an explicit [<x-fortran-namelist>] table.
  • Representing repeated same-name namelist blocks in TOML.
  • Full build-system orchestration for linking toml-f.
  • Extending CLI validation to .toml input in the same issue unless it falls
    out naturally from the same parsing abstraction.

Notes

This feature fits the current schema-centric design of nml-tools, but only if
the TOML mapping is kept strict.

The right abstraction is not "make the generated type understand arbitrary
TOML". It is "add a second reader for the same schema-owned named input group".

That keeps the current package assumptions intact:

  • one schema owns one generated type,
  • one generated type reads one named input group,
  • several schemas may still share one physical input document,
  • the schema remains the single source of truth for both namelist and TOML
    input.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions