Skip to content

Separate Generated Data And Dimensions Into Namespaces #65

Description

@MuellerSeb

Generated namelist types should separate schema-backed values, runtime
dimensions, and operational procedures into distinct Fortran namespaces.

The generated public layout should become conceptually:

type, public :: nml_run_data_t
  type(period_t) :: period
  type(period_t), allocatable :: periods(:)
  type(station_t) :: station
end type nml_run_data_t

type, public :: nml_run_dims_t
  integer :: n_periods = n_periods__dim_default
end type nml_run_dims_t

type, public :: nml_run_t
  type(nml_run_data_t) :: data
  type(nml_run_dims_t) :: dims
  logical :: is_configured = .false.
contains
  procedure :: init => nml_run_init
  procedure :: init_type => nml_run_init_type
  procedure :: set_dims => nml_run_set_dims
  procedure :: from_file => nml_run_from_file
  procedure :: set => nml_run_set
  procedure :: is_set => nml_run_is_set
  procedure :: is_valid => nml_run_is_valid
end type nml_run_t

Callers would use:

status = config%from_file("run.nml", errmsg=errmsg)
year = config%data%period%start_year
n = config%dims%n_periods

This removes schema values and runtime dimensions from the namespace occupied
by type-bound procedures and generated lifecycle state. Adding a generated
operation should no longer make an otherwise valid schema property or runtime
dimension name illegal.

Problem

The current generated namelist type stores all of these directly on one
Fortran derived type:

  • schema properties;
  • configured runtime dimensions;
  • generated state such as is_configured; and
  • type-bound procedures such as init, set, from_file, is_set, and
    is_valid.

Fortran component and type-bound binding names share a case-insensitive
namespace. The generator must therefore reject schema properties and runtime
dimensions that collide with any generated type member.

This has several undesirable consequences:

  • ordinary domain names such as set, init, or is_valid cannot be used as
    schema properties;
  • the reserved-name list grows whenever another type-bound operation is added;
  • adding an operation can invalidate existing schemas even though their
    namelist representation has not changed;
  • conditional operations such as set_dims and filled_shape still influence
    identifier policy; and
  • properties and runtime dimensions also share the same component namespace,
    despite serving different purposes.

Prefixing generated procedures would reduce the likelihood of collisions but
would not create a stable separation. Moving the values into a nested
container is the direct way to retain the type-bound API while decoupling it
from schema naming.

Proposed Public Layout

Generate up to three related types for each namelist schema:

  1. <namelist>_data_t contains all schema properties in resolved property
    order.
  2. <namelist>_dims_t contains configured runtime dimensions in configuration
    order. It is needed only when the namelist uses runtime dimensions.
  3. The existing <namelist>_t contains data, optional dims, lifecycle
    state, and type-bound procedures.

The exact type-name prefix should follow the existing generated type naming
policy. Companion type names must be checked against other module-level
generated and imported symbols case-insensitively. This is a module-symbol
concern and must not reintroduce restrictions on schema property names.

The generated data and dims components should be public. Their declared
companion types should have sufficient accessibility for portable direct
component access, copying, and use in associate constructs.

For schemas without runtime dimensions, omit the dimensions type and the
outer dims component:

type, public :: nml_report_t
  type(nml_report_data_t) :: data
  logical :: is_configured = .false.
contains
  ! generated bindings
end type nml_report_t

Namespace Policy

After this change:

  • schema properties must remain valid Fortran identifiers and unique
    case-insensitively within data;
  • runtime dimensions must remain valid Fortran identifiers and unique
    case-insensitively within dims;
  • a schema property may use the same name as a generated operation;
  • a runtime dimension may use the same name as a generated operation;
  • a property and a runtime dimension may use the same name because they are
    selected through different containers;
  • schema properties named data or dims are legal and appear as, for
    example, config%data%data and config%data%dims; and
  • errmsg remains reserved case-insensitively for schema properties and
    runtime dimensions because it is the stable public error-message keyword on
    generated type-bound procedures; and
  • existing restrictions unrelated to the outer derived-type namespace remain
    in force, including the reserved internal __ separator and ambiguity
    between constants and runtime dimensions in generated expressions.

This issue does not claim that every Fortran identifier becomes usable in
every generated procedure. Names that collide with setter dummies, error
arguments, local namelist variables, or generated temporaries require
collision-safe local naming or separate targeted restrictions. Those
subprogram-scope collisions should not be confused with the derived-type
member collision addressed here.

Type-Bound Procedure Scope

The namespace split must not simply move identifier collisions into generated
type-bound procedures. In particular, set_dims(...), from_file(...), and
set(...) must accept otherwise valid schema-property and runtime-dimension
names without colliding with generated procedure-scope identifiers.

Use the reserved __ separator for every generated support identifier in
these procedures: dummy arguments which are not schema fields, function
result names, and local variables. Use the nml__ prefix consistently where
possible: nml__status as a function result, nml__obj for the passed-object
dummy, and names such as nml__file, nml__iostat, nml__iomsg,
nml__close_status, and <dimension>__candidate for generated
dummies and locals. This lets a property or dimension be named status,
file, nml, iostat, or close_status. The public error-message dummy
remains errmsg.

The local variables that participate in a Fortran namelist statement are
the deliberate exception: they must keep the schema spelling so that native
namelist input remains unchanged. Implement from_file(...) as a small public
wrapper and an internal reader helper: the wrapper has no schema-spelled
locals, while the helper owns the namelist variables and receives only
collision-safe nml__* dummies. Thus a schema field named file does not
collide with the public reader's source-file argument or the helper's
operational state. errmsg remains rejected by the general procedure-keyword
reservation below.

Generated code calls its small set of required intrinsics directly. Rather
than re-exporting intrinsics through the helper module, reject identifiers
that would shadow them in generated scopes. This documented restriction
includes present, size, shape, allocated, associated, trim, len,
any, all, huge, reshape, achar, char, iachar, ichar, index,
len_trim, and minval, case-insensitively. Apply it to root properties,
runtime dimensions, constants, kind aliases, and derived type names where
they are emitted unqualified.

The helper source itself is generator-owned and required whenever native
Fortran or f2py output is requested. Configured application modules remain
supported for imported derived types and kinds, not as replacements for the
helper API.

Keep errmsg as the public error-message dummy on every generated procedure.
Rather than introducing an artificial err__msg keyword or procedure-specific
spelling, reserve errmsg case-insensitively for schema properties and runtime
dimensions. This deliberately narrow restriction preserves the existing native
Fortran API.

This policy applies to every generated procedure that introduces
procedure-scope identifiers, not only the three procedures above. Add a
focused collision audit whenever a generated procedure gains a dummy or local
name.

Runtime-Dimension Invariants

Runtime dimensions control allocation and validation of fields in data:

allocate(this%data%periods(this%dims%n_periods))

The generated set_dims(...) operation remains the supported way to mutate
them. It must continue to:

  1. calculate candidate dimensions;
  2. validate all candidates and default-extent requirements before mutation;
  3. update this%dims only after validation succeeds;
  4. deallocate affected arrays in this%data; and
  5. mark the outer object unconfigured.

Fortran does not provide read-only public components. Keeping
config%dims%n_periods directly readable therefore also leaves it directly
assignable. This is already possible with the current flat public dimension
members. Initially preserve that accessibility and document that direct
assignment does not resize dependent storage and that callers must use
set_dims(...) for a consistent update.

Making dimensions private and adding a query API can be considered separately;
it is not required for the namespace split.

Array-Bounds Policy

Generated namelist storage is strictly one-based in every array dimension.
Fixed-shape declarations and all generated allocate(...) statements must use
the default lower bound of one. Public index arguments, schema indices, filled
shapes, and native namelist subscripts therefore always refer to the logical
range 1:size(array, dim).

Generated code must not query lower bounds for its owned storage. Replace
lbound(...)/ubound(...) bookkeeping with one-based expressions:

array(1:size(array))
array(1:size(array, 1), 1:size(array, 2))

In particular:

  • index validation checks 1 <= idx(dim) <= size(array, dim);
  • flexible-array scans iterate from size(array, dim) down to 1;
  • the filled length is the final populated index directly;
  • partial setter assignments target 1:size(value, dim) in each dimension;
    and
  • populated-prefix validation uses 1:filled(dim).

Assumed-shape setter dummies are already one-based within the called procedure
unless their declaration explicitly preserves bounds. Assignment is by shape,
so the caller's actual lower bounds do not change the one-based bounds of the
generated destination storage.

Because allocatable data components are publicly accessible, a caller could
manually deallocate one and reallocate it with a non-one lower bound. That is
outside the generated type's supported invariants; callers must not change
array bounds directly. Generated initialization and runtime-dimension
allocation restore one-based storage.

Behavior That Must Not Change

The containers change generated Fortran storage access, not namelist or schema
paths.

Namelist input remains flat at the root:

&run
  period%start_year = 2000
  periods(1)%start_year = 1980
/

It must not become:

&run
  data%period%start_year = 2000
/

The generated reader already uses local variables in the intrinsic namelist
group and transfers the results into instance storage. Only the transfer target
changes to this%data%....

The following interfaces and representations should also remain unchanged:

  • schema and template paths such as period%start_year;
  • string lookup paths accepted by is_set(...) and filled_shape(...);
  • native set(...), set_dims(...), from_file(...), and is_valid(...)
    signatures, except where a separate identifier-collision fix is necessary;
  • generated Markdown field names and requiredness/default descriptions;
  • native namelist templates;
  • f2py opaque-handle APIs and Python-facing field names;
  • default, sentinel, requiredness, enum, bounds, shape, and presence semantics;
    and
  • local and imported derived values stored below a schema property.

Compatibility And Migration

This is a breaking native Fortran source change:

value = config%iterations

becomes:

value = config%data%iterations

and:

n = config%n_items

becomes:

n = config%dims%n_items

There is no robust general way to expose both spellings. Pointer aliases would
complicate intrinsic scalars, allocatable arrays, structure assignment,
allocation ownership, and object lifetime. The generator should not add a
second alias representation of the same values.

Adopt the new layout either in a documented breaking release or first behind a
temporary explicit generation option with a stated migration schedule. Do not
silently mix flat and namespaced storage among namelists generated by the same
configuration mode.

The generated Python API can remain source-compatible because it operates
through opaque handles and wrapper procedures rather than direct Fortran
component selection.

Application code with many reads can reduce repetition using associate:

associate(data => config%data)
  start_year = data%period%start_year
  station_code = data%station%code
end associate

Implementation Notes

  • Add generator context for the data and dimensions companion type names.
  • Render schema field declarations inside the data type instead of the outer
    namelist type.
  • Render runtime-dimension declarations inside the dimensions type.
  • Update all generated storage expressions consistently:
    • sentinel and default initialization;
    • runtime allocation and deallocation;
    • object/item default overlays;
    • local namelist transfers;
    • setter assignments;
    • derived init_type(...) calls;
    • presence and shape queries;
    • enum and bounds checks; and
    • validity checks.
  • Keep lifecycle state and operations on the outer type. Encapsulating or
    renaming is_configured is optional follow-up work rather than part of this
    issue.
  • Let storage-expression construction happen in shared generator helpers or
    context fields rather than scattering literal this%data% and
    this%dims% concatenation across validation branches.
  • Keep local intrinsic namelist variables named according to schema properties
    so input spelling does not change.
  • Update f2py wrapper generation only where it accesses native storage
    internally; do not expose the containers through the Python API.
  • Regenerate committed fixtures and examples after the generator tests define
    the complete new layout.

Acceptance Criteria

  • Every generated namelist type stores schema properties below a public data
    component.
  • Namelists with runtime dimensions store them below a public dims component.
  • Generated operations and lifecycle state remain on the outer namelist type.
  • Properties and runtime dimensions may use names that collide with generated
    outer members, including init, set, from_file, is_set, is_valid,
    filled_shape, and is_configured where otherwise valid.
  • A property and runtime dimension may have the same case-insensitive name
    without a Fortran component collision.
  • Properties named data and dims generate valid, unambiguous component
    paths.
  • set_dims(...), from_file(...), and set(...) compile and operate for
    properties or runtime dimensions named status, with
    collision-safe generated procedure identifiers.
  • Properties and runtime dimensions named errmsg are rejected
    case-insensitively with a targeted diagnostic explaining the stable
    error-message keyword reservation.
  • Generated procedure support identifiers use the reserved __ separator;
    only schema-spelled local namelist variables retain their original names.
  • Properties and runtime dimensions named present, along with other direct
    intrinsic dependency names, are rejected case-insensitively with a focused
    diagnostic.
  • Runtime allocations and all validation logic use the instance's values from
    dims and storage from data.
  • Failed set_dims(...) calls leave both containers unchanged; successful
    calls preserve the existing deallocation and configuration-state policy.
  • Every generated array has lower bound one in every dimension, and generated
    indexing, partial assignment, filled-shape, and validation code uses
    1:size(...) rather than lbound(...)/ubound(...).
  • Native namelist syntax, schema paths, generated templates, Markdown paths,
    and is_set(...) query strings remain unchanged.
  • Existing Python-facing wrapper signatures and behavior remain unchanged.
  • The native Fortran migration from %field and %dimension to %data%field
    and %dims%dimension is documented prominently.
  • No pointer-based compatibility aliases are generated for the old flat
    component paths.

Tests

Add focused generator and compiled-Fortran coverage for:

  • generated companion type declarations with and without runtime dimensions;
  • direct reads of intrinsic scalar, intrinsic array, local derived, and
    imported derived values through config%data;
  • runtime-dimension reads through config%dims;
  • initialization, defaults, sentinels, setters, file input, presence checks,
    shape queries, and validation through the nested storage paths;
  • allocation and reallocation using non-default runtime dimensions;
  • one-based bounds after initialization, set_dims(...), set(...), and
    from_file(...), plus generated output free of storage-bound
    lbound(...)/ubound(...) bookkeeping;
  • setter input arrays whose caller-side lower bounds are not one, verifying
    shape-based assignment into one-based generated storage;
  • failed and successful set_dims(...) transactional behavior;
  • property names matching every generated outer member case-insensitively;
  • runtime-dimension names matching generated outer members;
  • property and runtime-dimension names matching generated procedure results,
    dummies, and locals, including status, file, nml, iostat,
    close_status, and present;
  • targeted rejection of a property and a runtime dimension named errmsg,
    including case variants;
  • from_file(...) forwarding through its internal reader helper without
    changing native namelist field spelling; and
  • a field and a runtime dimension named present, including calls through
    the nml__present(...) intrinsic alias.
  • a property and runtime dimension with the same name;
  • schema properties named data and dims;
  • deterministic collision handling for generated companion type names and
    imported module symbols;
  • unchanged namelist, Markdown, template, f2py, and Python-facing names; and
  • regenerated golden fixtures and compiled examples using the new native
    component paths.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

FortranFortran related issuedocumentationImprovements or additions to documentationenhancementNew feature or request

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions