Skip to content

[rntuple] Read a column's value range as IEEE 754 doubles - #125

Merged
sathabbott merged 2 commits into
nsmith-:mainfrom
sathabbott:fix/75-column-range-doubles
Sep 30, 2026
Merged

sathabbott merged 2 commits into
nsmith-:mainfrom
sathabbott:fix/75-column-range-doubles

Conversation

@sathabbott

@sathabbott sathabbott commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

🤖 AI generated content

Fixes #75.

  • ColumnDescription.fMinValue / fMaxValue are read as "<d" (IEEE 754 little-endian doubles, per the spec's Column Description), typed float | None. They were read as int64.
  • The two checkboxes of [rntuple] ColumnDescription fMinValue/fMaxValue are IEEE doubles but are read as int64 #75: the abbott TODO on fFirstElementIndex is resolved (it is signed as read; a negative value means deferred and suppressed), and field flag 0x08 (SoA) is added to the flags docstring. No parse change for either.

Test: tests/test_column_value_range.py reads scikit-hep-testdata's test_float_types_rntuple_v1-0-0-0.root. Exactly columns 4–10 (Real32Quant, fields quant1–quant32) have a range, and each reads exactly -2.0 / 3.0; every other column has None. On main the same bytes read as -4611686018427387904 / 4613937818241073152, so the test fails there.

Full suite: 355 passed, 53 skipped, 46 xfailed (one more than main).

Breaking change: fMinValue / fMaxValue change type from int | None to float | None. Any caller that decoded the int64 bit patterns itself now gets the doubles directly.

Assisted-by: claude-code:claude-opus-5-5

ColumnDescription.fMinValue / fMaxValue were read as little-endian int64
("<q"). The spec (Column Description) says the range is "specified as IEEE
754 little-endian double precision floats"; ROOT bit-casts a double through
a uint64 when serializing, which is presumably how the integer reading
crept in. Read them as "<d", typed float | None.

Reproduced on scikit-hep-testdata's test_float_types_rntuple_v1-0-0-0.root:
its seven Real32Quant columns (4 to 10) read as -4611686018427387904 /
4613937818241073152, the bit patterns of -2.0 / 3.0, and now read as the
doubles.

Also, from the same section of the spec:
- fFirstElementIndex is signed, as read: a negative value means the column
  is deferred and suppressed. Replace the TODO with that.
- Document field flag 0x08 (a collection stored in a SoA layout). It adds no
  optional field, so parsing is unchanged.

Fixes nsmith-#75

Assisted-by: claude-code:claude-opus-5-5
@codecov-commenter

codecov-commenter commented Sep 24, 2026 •

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 92.16%. Comparing base (3f43287) to head (1c0830a).
⚠️ Report is 2 commits behind head on main.
❗ Your organization needs to install the Codecov GitHub app to enable full functionality.

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #125      +/-   ##
==========================================
+ Coverage   91.82%   92.16%   +0.33%     
==========================================
  Files          37       37              
  Lines        2962     2962              
==========================================
+ Hits         2720     2730      +10     
+ Misses        242      232      -10     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@sathabbott

Copy link
Copy Markdown
Collaborator Author

@nsmith- I've reviewed and approved this, but will let you see it before it merges.

Comment thread tests/test_column_value_range.py
Comment thread tests/test_column_value_range.py
Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
…_page_descriptions.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
…_info_content.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
…ms.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
@sathabbott

Copy link
Copy Markdown
Collaborator Author

@nsmith- i double checked this just now and think it is ready to merge

@sathabbott
sathabbott merged commit 1845a70 into nsmith-:main Sep 30, 2026
9 checks passed
@sathabbott
sathabbott deleted the fix/75-column-range-doubles branch September 30, 2026 14:52
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
…_page_descriptions.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
…_info_content.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit that referenced this pull request Sep 30, 2026
…er items of #87 (#130)

* Fix TKey's compression test, refuse RBlob keys, and two key nits

- A key payload is compressed iff fObjLen > fNbytes - fKeyLen, the test at
  all of ROOT's read sites (root-io-spec Compression §1), in read_object
  and in TKey_header.is_compressed(). A raw payload longer than fObjLen has
  slack; its first fObjLen bytes are the object. The "!=" form read that
  slack as a compression block header.
- read_object refuses a key of class RBlob. Its fObjLen is decorative and
  one blob can hold several pages (root-io-spec RNTuple NOTES 1), so even
  the correct test would silently truncate it; RNTuple bytes are found
  through the anchor's envelope locators and the page lists.
- A key is large when fVersion > 1000, as in TKey.cxx, so is_short() is
  <= 1000 and update_members uses it (root-io-spec Record §2 and its
  erratum 9). Differs only at exactly 1000.
- Key names are uninterpreted bytes (Conventions §5.1): TKeyList decodes
  them as UTF-8 with surrogateescape instead of ASCII, so a non-ASCII name
  no longer raises and every name maps back to exactly its bytes.

Fixes #77
Part of #87

Assisted-by: claude-code:claude-opus-5-5

* Rename tests/test_tkey.py to tests/test_key_compression_and_names.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on #125).

Assisted-by: claude-code:claude-opus-5-5

* Key TKeyList by the name bytes, with no decoding

Key names are uninterpreted bytes (root-io-spec Conventions §5.1), and
rootfilespec is not the user-facing layer, so TKeyList is now a
Mapping[bytes, TKey]: iteration yields the stored bytes and lookup takes
bytes. Callers that indexed with str (the tests and docs/design.md) now
pass bytes.

Assisted-by: claude-code:claude-opus-5-5

* Close the rest of #87: TDatime, VersionInfo, fUnits, compressed_size()

- TDatime_to_datetime returns None when the packed fields are not a
  calendar date (fDatime == 0, month 13, ...), instead of raising
  ValueError; ROOT does not validate them (root-io-spec Record §3.7).
  write_time(), create_time() and modify_time() are typed datetime | None.
- VersionInfo: a file version >= 1000000 is large (FileHeader §3).
- fUnits is read as u8 (">B") in all three file-header classes
  (FileHeader §2.2); the v622 large header had no byte-order marker.
- RCompressed.compressed_size() is the payload's length on disk: 9 plus
  each block's compressed size, which for LZ4 includes the checksum
  (Compression §4-§5). It left out the headers and the LZ4 checksum.

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
…ms.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
…_page_descriptions.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
…_info_content.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
…_page_descriptions.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Sep 30, 2026
…_info_content.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit that referenced this pull request Oct 1, 2026
* Keep and verify page checksums

A page whose fNElements is negative has an XXH3-64 checksum stored
little-endian right after it, over the stored (sealed) bytes. The locator's
size does not include it (spec, Page Locations; ROOT's reader adds the eight
bytes back itself, RPageStorage.cxx:297). rootfilespec neither read nor
checked it.

- RPageDescription keeps fNElements as stored (the sign is the flag, and the
  stored value is what reproduces the bytes) and gains n_elements,
  has_checksum and stored_size. size stays the locator's size, as the spec
  defines it.
- RPageDescription.page_locator is a new RPageLocator covering the stored
  bytes (offset, stored_size). Its read_from verifies the checksum and
  returns an RPage with the raw bytes and the checksum; nothing is
  decompressed. PageListEnvelope.page_locators returns these, and get_page
  uses it.

Fixes #55

Assisted-by: claude-code:claude-opus-5-5

* Ignore the missing xxhash stubs in the pre-commit mypy environment

The pre-commit mypy hook installs only pytest, numpy and tomli, so it cannot
find xxhash. bootstrap/compression.py already imports it the same way.

Assisted-by: claude-code:claude-opus-5-5

* Rename tests/test_rntuple_pages.py to tests/test_page_checksums.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on #125).

Assisted-by: claude-code:claude-opus-5-5

* Verify page checksums in RPageDescription, drop RPageLocator

RPageDescription already is the page's locator. Instead of a second
locator class next to it, it now fetches and verifies the checksum
itself, as suggested on #55:

- size is the number of bytes to fetch: locator.size, plus the 8-byte
  checksum when fNElements is negative. The spec's page size, which
  excludes the checksum (Page Locations), stays in locator.size.
- read_from splits off the trailer, verifies XXH3-64 over the stored
  bytes and returns RPage(page, checksum). get_page goes through it.

RPageLocator flattened the RLocator into an in-file (offset, size) and
left RPageDescription as a second, unverified way to read a page.
Keeping the RLocator leaves room for further locator types (the spec
reserves 0x02-0x7f), and the checksum check depends only on the
fetched bytes. PageListEnvelope.page_locators is unchanged from main,
so this PR no longer changes any API.

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Oct 1, 2026
…ms.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Oct 1, 2026
Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Oct 1, 2026
…_page_descriptions.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Oct 1, 2026
…_info_content.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit that referenced this pull request Oct 1, 2026
…#127)

* Read RNTuple strings as plain bytes

docs/design.md promises that where a builtin Python type captures a ROOT
type, the deserialized member is that builtin: bytes for strings. RNTuple
strings were RString wrappers instead.

Add CountedString, a MemberSerDe for a string stored as its length (in a
given struct format) followed by that many bytes, read as plain bytes. An
RNTuple string is Annotated[bytes, CountedString("<I")], exported from
rntuple.schema as RNTupleString, and the eight RNTuple string members (the
header's name, description and library; a field's name, type name, type
alias and description; extra type information's type name) use it.
RString is removed.

This is the RNTuple row of #68. The TString, std::string and char* rows
stay open there; char* (#20) can use CountedString(">i").

Part of #68

Assisted-by: claude-code:claude-opus-5-5

* Rename tests/test_rntuple_strings.py to tests/test_counted_string.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on #125).

Assisted-by: claude-code:claude-opus-5-5

* Compare header names with the key names as bytes, after #130

Assisted-by: claude-code:claude-opus-5-5

* Read every ROOT string type as plain bytes with ROOTString (rest of #68, #20)

Replace CountedString with ROOTString(encoding), one Literal value per
on-disk form (root-io-spec Conventions §5), as suggested in review:

- "RNTuple": u32 little-endian length, then the bytes
- "TString": the counted string, one length byte or 255 then a u32
- "std::string": a std::string data member, the counted string inside a
  byte count and version word (§5.3); the byte count is checked
- "charstar": a char* member, i32 length then the bytes, no 255 escape;
  a length <= 0 reads as b"" (§5.4, #20)

TString is now the alias Annotated[bytes, ROOTString("TString")], so
every TString member is bytes and .fString is gone. STLString is
removed. read_string(buffer, encoding) reads one string by hand (TKey,
TList options).

Generated classes write the annotation out: TStreamerString emits
ROOTString('TString'), TStreamerSTLstring ROOTString('std::string'), a
kCharStar basic type ROOTString('charstar'), and cpptype maps string,
std::string and TString elements of containers and pointees to
ROOTString('TString').

A string stored as an object of its own (a key, or through a pointer) is
looked up by class name, so TString, TStringLong and string stay
resolvable by name, and serializable.read_value reads either a class or
an annotated builtin. TStringLong is the "charstar" encoding (§5.1.1).
As a result, root-io-spec's serialization/stringlong.root now reads, and
serialization/unframed-records.root reads its TDatime, TString and
TStringLong records and stops at the TObject record (#123).

Assisted-by: claude-code:claude-opus-5-5

* Make a std::string member's frame a ROOTString option, not an encoding

A std::string data member is the TString encoding inside the usual byte
count and version word (root-io-spec Conventions §5.3). The frame is
not a fourth string encoding, so:

- StringEncoding is back to the three length formats Nick listed on
  #127: "RNTuple", "TString", "charstar".
- ROOTString(encoding, framed=False). With framed=True it reads the
  frame with StreamHeader, as StdVector and the other framed members
  do, and checks the byte count, instead of parsing the frame itself
  with a second copy of kByteCountMask.
- read_string() is gone. ROOTString.read(buffer) is the one string
  reader: the member reader calls it, and TKey and TList call it for
  their hand-read strings.
- read_value() takes type[MemberType] | type[ROOTSerializable], as Nick
  suggested, and goes through _build_read for both kinds of type.
- docs/design.md states the design assumption Nick raised: the bytes
  don't record their encoding; the annotation, key class name or class
  tag that holds them does, and writing a string back needs it.
- test_tdatime_record is dropped: the TDatime record reads as a side
  effect of read_value, but it belongs to #123, not #68.

Assisted-by: claude-code:claude-opus-5-5

* Read a TypedTKey's object with read_value, like an untyped key's

TypedTKey.read stores the looked-up type as objtype, and after #68 a
lookup of "TString", "TStringLong" or "string" returns an annotated
alias, not a class. read_object called objtype.read() on it and raised
AttributeError, so a TypedTKey of a string record could not be read,
although fetching the same key untyped returned its bytes. Both
branches now read through read_value.

Test: the "tstring" record of root-io-spec serialization/unframed-records
through TypedTKey reads b"hello".

Assisted-by: claude-code:claude-opus-5-5

* Fix the update_members typo in docs/design.md

"read_mupdate_membersembers" was a garbled "update_members", as Nick
noted on #127.

Assisted-by: claude-code:claude-opus-5-5

* Say that a string read through a pointer loses its ROOT type

The design note claimed a streamed object's class tag keeps a string's
type, like an annotation or a key's class name does. It doesn't:
read_streamed_item reads the tag to look the type up and then returns
bare bytes. Members, container elements and keyed records keep the
type; a string read through a pointer can't be written back as read.
Tracked in #135, as Nick noted on #127 and in #134's AGENTS.md.

Assisted-by: claude-code:claude-opus-5-5

* Compare key class names as bytes in the page checksum tests, after #128

#128's tests index key.fClassName.fString; after #68 the class name is
plain bytes.

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Oct 1, 2026
…ms.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Oct 1, 2026
…_page_descriptions.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Oct 1, 2026
…_info_content.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit that referenced this pull request Oct 2, 2026
* Verify envelope checksums

Every RNTuple envelope ends with an XXH3-64 checksum. rootfilespec read it
and compared stored checksums with each other, but never hashed the bytes,
so a corrupted header, footer or page list parsed without complaint.

REnvelope.read now checks the checksum over [0, length - 8) of the
uncompressed envelope: the length includes the checksum, and the checksum
covers everything before it (root-io-spec ERRATA 5, checked against ROOT's
VerifyXxHash3, RNTupleSerialize.cxx:934). Unknown trailing bytes are
covered too, as the spec's compatibility notes require.

REnvelopeLocator.read_from also refuses a stored size larger than the
uncompressed length. RNTuple decompression branches on equality (root-io-spec
NOTES 2, RNTupleZip.hxx:106-113), so that case is an error, where the code
used to try to decompress it.

Fixes #117

Assisted-by: claude-code:claude-opus-5-5

* Ignore the missing xxhash stubs in the pre-commit mypy environment

The pre-commit mypy hook installs only pytest, numpy and tomli, so it cannot
find xxhash. bootstrap/compression.py already imports it the same way.

Assisted-by: claude-code:claude-opus-5-5

* Rename tests/test_rntuple_envelopes.py to tests/test_envelope_checksums.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on #125).

Assisted-by: claude-code:claude-opus-5-5

* Verify envelope checksums before parsing, and check the fetched size

Two fixes from review:

- REnvelope.read now verifies the XXH3-64 right after the type and
  length checks, before the payload is parsed, as ROOT's
  DeserializeEnvelope does (RNTupleSerialize.cxx:909-939). Before, the
  payload was parsed first, so 136 of the 224 single-byte corruptions
  of rntuple/anchor.root's header envelope failed with an unrelated
  parse error instead of the checksum mismatch. Envelopes shorter than
  16 bytes are refused, as in ROOT.
- REnvelopeLocator.read_from checks that it was given the locator's
  size, then applies NOTES 2's raw/compressed rule to the locator's
  stored size rather than to the buffer's length. A short read (a
  truncated file) said "Unknown compression algorithm".

Tests: every corrupted byte of that envelope reports the checksum
mismatch; short reads of 0, 200 and 239 of 240 bytes raise.

Assisted-by: claude-code:claude-opus-5-5

* Read the envelope checksum through ReadBuffer, not struct

Slice the buffer to its last 8 bytes and unpack there, as the rest of
the library does, instead of importing struct into envelope.py.

Assisted-by: claude-code:claude-opus-5-5

* Compare key class names as bytes in the envelope tests, after #127

#127 made key class names plain bytes.

Assisted-by: claude-code:claude-opus-5-5

* Split the envelope once into the checksummed bytes and the checksum

REnvelope.read mixed three views of the envelope: the checksum was read
from the buffer after the preamble at length - 16, hashed against a
separate envelope_bytes slice at length - 8, and the payload was parsed
from the buffer. It now splits the envelope once, into the bytes the
checksum covers and the 8-byte checksum, verifies one against the
other, and parses the payload from the covered bytes after the
preamble, so nothing reads into the checksum (review on #129). The
preamble is still read from the whole buffer first, because the split
needs its length.

Every envelope that read before reads the same. Two errors change:
the length mismatch now reports the full envelope length and buffer
length (before, both minus the 8-byte preamble), and a frame that runs
past the payload raises IndexError ("Cannot get slice") where it raised
ValueError ("Cannot consume a negative number of bytes").

Assisted-by: claude-code:claude-opus-5-5

* Test the envelope length checks and the pinned header checksum

- rntuple/anchor.root's case.toml pins the header envelope's checksum at
  500 and the footer's copy at 854; assert the parsed values equal them.
- An envelope whose length is under 16 bytes (the preamble and checksum
  alone) raises, as in ROOT's DeserializeEnvelope.
- anchor.root's header with the preamble's length changed from 240 to
  248 raises the length mismatch.

Also drop the count of failures before the reordering from a test
docstring; the commit that reordered the checks records it.

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Oct 2, 2026
…_page_descriptions.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Oct 2, 2026
…_info_content.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit that referenced this pull request Oct 2, 2026
…infos (#132)

* Read the extra type information's content

ExtraTypeInformation ended at its type name. ROOT's serializer writes one
more string, the content (RNTupleSerialize.cxx:389-391; root-io-spec
ERRATA 9), so its length and bytes were left in the record frame's
_unknown. For content identifier 0 that content is the streamed TList of
TStreamerInfo that streamed (role 0x04) fields need.

Add fContent, an RNTuple string like the type name.

Fixes #118

Assisted-by: claude-code:claude-opus-5-5

* Rename tests/test_rntuple_extra_type_info.py to tests/test_extra_type_info_content.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on #125).

Assisted-by: claude-code:claude-opus-5-5

* Compare key class names as bytes, after the rest of #68 in #127

Assisted-by: claude-code:claude-opus-5-5

* Pin the streamed.root bytes exactly in the extra type info test

AGENTS.md asks tests to assert the values a case.toml pins. Compare the
content with the file's bytes at offset 1240 (length 438 at 1236), the
TStreamerInfo name's length byte at 1319 and the name after it, and the
frame size 462, instead of only checking that the name occurs somewhere.

Say in fContent's docstring that it has an RNTuple string's layout but
holds binary bytes, not UTF-8 text.

Assisted-by: claude-code:claude-opus-5-5

* Decode the streamer info content: RNTuple.streamer_infos()

The extra type information with content identifier 0 holds the
TStreamerInfo of every class that streamed (role 0x04) fields need, as a
TList written as if through a pointer (RNTupleSerializer::
SerializeStreamerInfos, ROOT 6.40.04). For an RNTuple read without its
TFile, it is the only description of those classes.

streamer_infos() reads it with the bootstrap classes, by class name like
FileReader.streamerinfos(). It is derived: fContent keeps the bytes. A list
item that comes back as a Ref is unwrapped, so the result stays the same
when #105 makes every pointee a Ref. Other content identifiers are
ignored, as the spec asks; anything else in the content, bytes after the
TList, or the same class twice is an error rather than a guess.

Assisted-by: claude-code:claude-opus-5-5
sathabbott added a commit to sathabbott/rootfilespec that referenced this pull request Oct 2, 2026
…_page_descriptions.py

Name the module after what it tests, so it does not read as the home of a
broad set of tests (review on nsmith-#125).

Assisted-by: claude-code:claude-opus-5-5
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[rntuple] ColumnDescription fMinValue/fMaxValue are IEEE doubles but are read as int64

3 participants