Skip to content

4.0.1-alpha1 - hashing API; Wide-string CI - #42

Merged
synesissoftware merged 8 commits into
masterfrom
dev
Oct 6, 2026
Merged

synesissoftware merged 8 commits into
masterfrom
dev

Conversation

@synesissoftware

Copy link
Copy Markdown
Owner

No description provided.

synesissoftware and others added 2 commits October 4, 2026 23:14
* fix(performance tests): "assign_len_grow" now more realistic

* fix(performance tests): "create_destroy_empty" now more realistic

* fix(performance tests): "append_len_growth_x" now more realistic

* fix(performance tests): "append_len_reserved" now more realistic

* Revert "fix(performance tests): "create_destroy_empty" now more realistic"

This reverts commit 5998491.

* fix(performance tests): "create_destroy_empty" std::string now constructs ""

The cstring row calls cstring_create(&s, "") and raw_realloc assigns an
empty C string. Default-constructing std::string skipped that step, so the
row compared a materialised empty string with an untouched default object.
Construct std::string from "" so all three sides perform the same creation.
Small-string storage can still make std::string cheaper; that gap is the
library.

* fix(performance tests): assign and copy anchors read the first byte

assign_len_grow and copy returned only the length, which the compiler can
treat as invariant, so it deleted the visible raw_realloc memcpy. Rows from
256 bytes upward stayed near 12 ns with p50 of 0. Each implementation now
adds the first payload byte to the anchor, as create_destroy_len already
does, including the Windows arena lambdas.

* fix(performance tests): append and insert anchors read the first byte

append_len_growth, append_len_reserved, and insert_len_mid returned only
the length, which the compiler can treat as invariant, so it can delete the
visible raw_realloc copies. Each implementation now adds the first payload
byte to the anchor, as assign_len_grow and copy already do, including the
Windows arena lambdas. Every size in these scenarios leaves a non-empty
string, so the first byte is in range.

* fix(unit tests): COM initialisation

* fix(performance tests): append_batch anchor reads stored payload

append_batch returned only the vector length, which is the constant element
count, so the compiler can delete the visible vector<string> insert. The
payload-16 row printed 0 ns/op. Each implementation now takes the address
of the first and last stored string through a volatile pointer and adds the
first byte of each. That address is the copy, not the source, so those two
strings have to be materialised.

* fix(performance tests): one-by-one vector anchors read stored payload

append_one_by_one and prepend_one_by_one returned only the vector length,
which is the constant element count, so the compiler can delete the visible
vector<string> updates. Each implementation now takes the address of the
first and last stored string through a volatile pointer and adds the first
byte of each, as append_batch already does. create_destroy is unchanged:
its elements are empty, so there is no stored payload byte to read.

* fix(performance tests): Windows assign/append ratios use createEx realloc

On Windows, assign_len_grow, append_len_growth, and append_len_reserved
arena rows call createEx("") before the mutation. The printed cstring row
does not. Those arena ratios now use cstring_win_realloc, which runs the
same body with CSTRING_F_USE_REALLOC. std::string and raw_realloc stay
against the cstring row that skips the empty create. create, insert, and
copy are unchanged: their cstring row already matches the arena body.

* fix(performance tests): borrowed rows time a prebuilt buffer

borrowed_fixed_construct and borrowed_fixed_assign built the caller buffer
inside the timed loop, so sizes above 511 bytes included a heap allocation
in the borrowed and fixed_char_buf times. The buffer is now built once per
size and reused. std::string on those rows is an owning construct, and its
ratio column is "-" rather than a comparison with cstring_borrowed.
fixed_char_buf remains the like-for-like peer.

* fix(performance tests): file-line peers share an owning anchor

file_read_all_lines compared an owning fgetc reader with a getline loop
that discarded its buffer on every move, and with platformstl::file_lines,
which stores views over one copy of the file. getline now copies each line
and clear()s the scratch string, and the stream is opened binary. All three
anchors are the line count plus the sum of the line lengths, and a mismatch
drops the rows. file_lines is still timed, and its ratio column is "-"
because it does not own each line. ifstream+getline remains the owning
peer.

* fix(performance tests): dash the ratio when a row is elided

When p50 is 0 and ns/op has not risen against the previous smaller size of
the same scenario and implementation, vs cstr is "-". A deleted raw_realloc
copy was printing a falling ratio (down to 0.02) while the time stayed
flat.  The first such size is dashed as well.

* fix(performance tests): Windows copy calls cstring_copy

The Windows copy rows built the destination with createLenEx of the source
bytes, so they never called cstring_copy. They now createEx("") to select
the arena, then cstring_copy, which allocates with the destination flags.
The printed cstring row does not createEx(""), so those arena ratios use
cstring_win_realloc. std::string and raw_realloc stay on the cstring row.

* fix(performance tests): raw vector create leaves slots empty

create_destroy's raw floor reserved the pointer array, then calloc'd a byte
per slot. cstring_vector_create and vector<string>(n) only construct empty
elements, with no character buffer. The raw slots are now null. Destroy
still frees them, and free(NULL) is a no-op.

* fix(performance tests): fgets honours the readline instance mode

read_all_fgets always used one buffer of line_len+4, so fresh and
pre-reserved did not differ and reuse never split a long line. Fresh now
allocates that buffer per line, pre-reserved keeps one, and reuse keeps a
4096-byte buffer and stitches a short read, including a CR held across a
chunk boundary. Anchors stay the line length plus the first content byte.

* fix(performance tests): time ns/op as one interval

When p99 is linked, ns/op was the sum of per-call samples, and each sample
included start/stop. On short rows that overhead flattened the ratio. ns/op
and vs cstr now use one start/stop around the iteration loop. Percentiles
still come from a per-call pass on the recorded warmup, and that pass is
not added to ns/op.

* fix(performance tests): label percentiles as per iteration

p50, p90, p99, and max are the median and tails of one iteration. ns/op
divides the batch by #acts, so an append or readline percentile was the
whole trial while ns/op was one append or one line. The columns are now
p50/iter, p90/iter, p99/iter, and max/iter. The values stay one iteration.

* fix(performance tests): prepend at payload 256 as well as 16

prepend_one_by_one ran only at payload 16, where vector<string> is still
inside small-string storage. It now also runs at payload 256 and still
prepends 32 elements, beside the append rows. The 16-byte row remains.
* general boilerplate / project-support improvements (#16)

* squash-commit

* 4.0.14

* fix

* .vscode/settings.json

* Add cstring_hash_djb2() and cstring_hash_fnv1a() hashing API and test suite

* Added 64-bit djb2 and FNV-1a hash functions for cstring_t instances
  and slices (cstring_hash_djb2, cstring_hash_djb2_ci,
  cstring_hash_djb2_len, cstring_hash_djb2_len_ci, cstring_hash_fnv1a,
  cstring_hash_fnv1a_ci, cstring_hash_fnv1a_len,
  cstring_hash_fnv1a_len_ci);
* Added cstring_hash_t typedef (uint64_t) with fallback for older MSVC;
* Added C++ inline access shims for instances and buffer slices;
* Added comprehensive xTests unit test suite test.unit.cstring.hash;
* Verified both narrow-character and wide-character build paths;
* Documented hashing API, types, and C++ access shims in README.md;
* Bumped version to 4.1.0-alpha1 and updated CHANGES.md and NEWS.md;

* 4.0.15 (#20)

* Updated all `\param` and `\retval` Doxygen documentation comments in
  `cstring.h`, `cstring.vector.h`, and `cstring.core.c` to obey the 76-rule
  and terminate with a semicolon;
* Bumped version to 4.0.15 in `cstring.h`, `Doxyfile`, `CHANGES.md`, and
  `NEWS.md`;
* Updated `Updated:` header banners to 6th September 2026;

* squash-commit

* NEWS.md

* chore: renamed **test/unit/test.unit.cstring.hash/entry.cpp** => **test/unit/hash/entry.cpp**

* chore

---------

Co-authored-by: synesissoftware <matthew@synesis.com.au>
@synesissoftware synesissoftware self-assigned this Oct 4, 2026

@GarthJL1965 GarthJL1965 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Magic numbers being used as error returns ?

Comment thread include/cstring/cstring.h Outdated
Comment thread include/cstring/cstring.h Outdated
Comment thread src/cstring.hash.c Outdated
@synesissoftware synesissoftware changed the title 4.0.1-alpha1 - hashing API 4.0.1-alpha1 - hashing API; Wide-string CI Oct 5, 2026
* chore(misc): language discrimination; CI branch(es)

* chore(misc): ensuring `\file` designators are correct

* chore(mist): removed unused `CSTRING_VECTOR_ASSERT()`

* chore(misc): tidying

* fix(cstring): discriminate cstring_hash_t and use it for hashing

Gate <stdint.h> on C99, on C++ with that header or C++11, and on MSVC 7.1
or later, and typedef cstring_hash_t to uint64_t when it is included; use
cstring_hash_t for the djb2 and FNV-1a API and for the hash unit tests.

* chore(performance tests): prefer `uint64_t` to `std::uint64_t`

* test(cstring): exercise the hashing API from a C unit test

Move the eight C hashing functions into **test.unit.hash** (`entry.c`) so
the suite builds when `NO_CSTRING_CPP_API` is on. Keep the C++ shim checks
in **test.unit.hash.cxx**, which stays behind that option.

* test(cstring): lock high-byte and embedded-NUL hash results

Cover octet 0xFF and an interior NUL in test.unit.hash. Exact narrow
vectors stay off the wide build so a later wchar_t fix is not frozen.

* feat(cstring): publish the djb2 and FNV-1a basis macros

Add CSTRING_HASH_DJB2_SEED, CSTRING_HASH_FNV1A_OFFSET, and
CSTRING_HASH_FNV1A_PRIME beside cstring_hash_t, and use them as the only
copies of those constants in the hash implementation. Cite Yigit's djb2
page and the Fowler, Noll, and Vo notes and RFC 9923, and lock each macro
to its historic literal in the C unit test.

* fix

* refactor(cstring): rename the case-folded hash suffix to _case

Rename the case-folded hash functions from the `_ci` suffix to `_case`
(`cstring_hash_djb2_case`, `cstring_hash_djb2_len_case`, and the FNV-1a
twins), including the C++ shims and the unit tests.

* refactor(cstring): rename the counted hash suffix from _len to _buf

Rename the pointer-and-length hash functions from `_len` to `_buf`
(`cstring_hash_djb2_buf`, `cstring_hash_djb2_buf_case`, and the FNV-1a
twins), including the C++ shims and the unit tests.
synesissoftware and others added 4 commits October 6, 2026 16:52
* chore(misc): language discrimination; CI branch(es)

* chore(misc): ensuring `\file` designators are correct

* chore(mist): removed unused `CSTRING_VECTOR_ASSERT()`

* chore(misc): tidying

* fix(cstring): discriminate cstring_hash_t and use it for hashing

Gate <stdint.h> on C99, on C++ with that header or C++11, and on MSVC 7.1
or later, and typedef cstring_hash_t to uint64_t when it is included; use
cstring_hash_t for the djb2 and FNV-1a API and for the hash unit tests.

* chore(performance tests): prefer `uint64_t` to `std::uint64_t`

* test(cstring): exercise the hashing API from a C unit test

Move the eight C hashing functions into **test.unit.hash** (`entry.c`) so
the suite builds when `NO_CSTRING_CPP_API` is on. Keep the C++ shim checks
in **test.unit.hash.cxx**, which stays behind that option.

* test(cstring): lock high-byte and embedded-NUL hash results

Cover octet 0xFF and an interior NUL in test.unit.hash. Exact narrow
vectors stay off the wide build so a later wchar_t fix is not frozen.

* feat(cstring): publish the djb2 and FNV-1a basis macros

Add CSTRING_HASH_DJB2_SEED, CSTRING_HASH_FNV1A_OFFSET, and
CSTRING_HASH_FNV1A_PRIME beside cstring_hash_t, and use them as the only
copies of those constants in the hash implementation. Cite Yigit's djb2
page and the Fowler, Noll, and Vo notes and RFC 9923, and lock each macro
to its historic literal in the C unit test.

* fix

* refactor(cstring): rename the case-folded hash suffix to _case

Rename the case-folded hash functions from the `_ci` suffix to `_case`
(`cstring_hash_djb2_case`, `cstring_hash_djb2_len_case`, and the FNV-1a
twins), including the C++ shims and the unit tests.

* refactor(cstring): rename the counted hash suffix from _len to _buf

Rename the pointer-and-length hash functions from `_len` to `_buf`
(`cstring_hash_djb2_buf`, `cstring_hash_djb2_buf_case`, and the FNV-1a
twins), including the C++ shims and the unit tests.

* refactor(cstring): .DEF file renames of hash functions

* feat(cstring): add always-available multibyte and wide hash API

Move the hash declarations into hash.h and the stdint.h discrimination into
common.h. cstring.h includes both. Add mbs, wcs, mbuf, and wbuf entry
points, including the case forms, for djb2 and FNV-1a. Each code unit
contributes its low 8 bits, so ASCII text has one hash in both encodings.
C++ shims overload on char const*, wchar_t const*, and both buffer forms.

* docs(cstring): document hash groups, contract, and consumers

- Document the hash contract, the djb2 and FNV-1a steps, and the published
  vectors in hash.h, README.md, and doc/mainpage.md;
- Define group__cstring_api__hashing__djb2 and
  group__cstring_api__hashing__fnv1a, and keep the published references
  on those groups;
- Place the C++ hash shims in namespace cstring, and import them in
  test.unit.hash.cxx;
- Name the projects that call or link cstring, including the sistools
  programs;
- Parenthesise the hash basis macros and note the FNV-1a prime in decimal;
- Record the header-split and auto-buffer provenance items in TODO.md;
- Spell NUL in the hash and status-code comments;
* chore(misc): language discrimination; CI branch(es)

* chore(misc): ensuring `\file` designators are correct

* chore(mist): removed unused `CSTRING_VECTOR_ASSERT()`

* chore(misc): tidying

* fix(cstring): discriminate cstring_hash_t and use it for hashing

Gate <stdint.h> on C99, on C++ with that header or C++11, and on MSVC 7.1
or later, and typedef cstring_hash_t to uint64_t when it is included; use
cstring_hash_t for the djb2 and FNV-1a API and for the hash unit tests.

* chore(performance tests): prefer `uint64_t` to `std::uint64_t`

* test(cstring): exercise the hashing API from a C unit test

Move the eight C hashing functions into **test.unit.hash** (`entry.c`) so
the suite builds when `NO_CSTRING_CPP_API` is on. Keep the C++ shim checks
in **test.unit.hash.cxx**, which stays behind that option.

* test(cstring): lock high-byte and embedded-NUL hash results

Cover octet 0xFF and an interior NUL in test.unit.hash. Exact narrow
vectors stay off the wide build so a later wchar_t fix is not frozen.

* feat(cstring): publish the djb2 and FNV-1a basis macros

Add CSTRING_HASH_DJB2_SEED, CSTRING_HASH_FNV1A_OFFSET, and
CSTRING_HASH_FNV1A_PRIME beside cstring_hash_t, and use them as the only
copies of those constants in the hash implementation. Cite Yigit's djb2
page and the Fowler, Noll, and Vo notes and RFC 9923, and lock each macro
to its historic literal in the C unit test.

* fix

* refactor(cstring): rename the case-folded hash suffix to _case

Rename the case-folded hash functions from the `_ci` suffix to `_case`
(`cstring_hash_djb2_case`, `cstring_hash_djb2_len_case`, and the FNV-1a
twins), including the C++ shims and the unit tests.

* refactor(cstring): rename the counted hash suffix from _len to _buf

Rename the pointer-and-length hash functions from `_len` to `_buf`
(`cstring_hash_djb2_buf`, `cstring_hash_djb2_buf_case`, and the FNV-1a
twins), including the C++ shims and the unit tests.

* refactor(cstring): .DEF file renames of hash functions

* feat(cstring): add always-available multibyte and wide hash API

Move the hash declarations into hash.h and the stdint.h discrimination into
common.h. cstring.h includes both. Add mbs, wcs, mbuf, and wbuf entry
points, including the case forms, for djb2 and FNV-1a. Each code unit
contributes its low 8 bits, so ASCII text has one hash in both encodings.
C++ shims overload on char const*, wchar_t const*, and both buffer forms.

* docs(cstring): document hash groups, contract, and consumers

- Document the hash contract, the djb2 and FNV-1a steps, and the published
  vectors in hash.h, README.md, and doc/mainpage.md;
- Define group__cstring_api__hashing__djb2 and
  group__cstring_api__hashing__fnv1a, and keep the published references
  on those groups;
- Place the C++ hash shims in namespace cstring, and import them in
  test.unit.hash.cxx;
- Name the projects that call or link cstring, including the sistools
  programs;
- Parenthesise the hash basis macros and note the FNV-1a prime in decimal;
- Record the header-split and auto-buffer provenance items in TODO.md;
- Spell NUL in the hash and status-code comments;

* test(cstring): time djb2 and FNV-1a against a lose-lose baseline

Add test.performance.hash. Each group is one scenario and length, with
lose-lose, djb2, and FNV-1a, across cstring_t, multibyte and wide strings,
and counted buffers, including case folding, embedded NULs, and high
octets. A further algorithm is one extra row on each table.

Stop cstring_strlcpy_safe_() from scanning past lim. createLen, assign,
insert, and append pass a slice of cch characters that need not contain a
NUL. Cover that with a 3-character createLen case in test.unit.cstring.

* test(cstring): report hash spread from a scratch program

Add test.scratch.hash. It builds 10000 distinct pseudo-random strings and
prints full-hash uniqueness, per-octet occupancy, and residues for
power-of-two hashtable sizes and a prime just below each, for lose-lose,
djb2, and FNV-1a.

* feat(cstring): add a 64-bit SDBM hash

Add `cstring_hash_sdbm()` and the multibyte, wide, counted-buffer, and
`_case` forms. The step is `hash * 65599 + octet` from seed 0, in a 64-bit
accumulator. C++ shims are `hash_sdbm()` and `hash_sdbm_case()`. Unit tests
lock the published vectors; the performance and scratch programs append an
SDBM row after djb2 and FNV-1a.

* fix

* fix(cstring): hash every little-endian octet of a wide code unit

Step once per octet of each wchar_t, low byte first, in the wide djb2,
FNV-1a, and SDBM functions and in the lose-lose timing baseline. A
multibyte string and its wide equivalent now differ; the published
multibyte vectors stay as they are, and 64-bit djb2 is documented to match
32-bit djb2 only below 2^32.

* fix: corrected old-VC++ compatibility

* fix(cstring): recognise stdint.h on Visual C++ from VS 2010

Require _MSC_VER 1600 before treating Visual C++ as having <stdint.h>.
Older MSVC keeps the unsigned __int64 fallback for cstring_hash_t.

* test(cstring): assert _case hashes ignore setlocale

Compare C-locale and tr_TR.UTF-8 _case results for octet 0xDD and U+00C0.
Those inputs change under tolower and towlower, so the case fails until the
fold stops following the process locale.

* fix(cstring): fold _case hashes on ASCII A-Z only

Map A-Z to a-z by adding 0x20 and leave every other code unit unchanged, in
the djb2, FNV-1a, and SDBM functions and in the lose-lose baseline.
setlocale no longer changes a _case result.

* feat(cstring): specialise std::hash<cstring_t> on FNV-1a

Provide std::hash<cstring_t> from C++11 as the FNV-1a hash of ptr and len,
converted to size_t. djb2 and SDBM stay available as explicit functions.

* test(cstring): rename hash tests off the old _ci and _len suffixes

Rename the case and counted-buffer unit-test functions to _case and _buf,
matching the hash API. Local bases in the case test follow the same suffix.

* feat(cstring): add cstring_equal and cstring_compare

Add cstring_equal(), returning cstring_truthy_t, and cstring_compare(),
returning cstring_sint_t. Equal returns before reading when the lengths
differ, and memcmp's the payloads only when they match. Compare returns a
signed order of len code units. The multibyte build memcmp's the common
prefix. C++ ==, !=, and < call those functions. Drop the cstring_vector
example's private comparator. Its qsort adapters call cstring_compare().

* test(cstring): tolerate a host with no Turkish ctype locale

Select tr_TR.UTF-8 or another Turkish ctype name when the host has one. A
missing Turkish locale is not a failure. Linux CI generates tr_TR.UTF-8 so
that comparison still runs there.
* refactor(cstring): keep the character helper under test/

Move cstring.helpers.h from include/cstring/ to test/. Examples include it
by relative path, and the test suites include it from cstring_testing.h. It
is not part of the installed library contract.

* fix

* fix(cstring): check the wctomb() shift-state result

GCC rejects a discarded wctomb() result, including a (void) cast. Test the
NULL call that restores the initial shift state.

* ci(cstring): run wide-string Windows cells without performance tests

Add CI cells windows-cl-wide and windows-mingw-wide. They call the shared
cell with --wide-strings, so unit tests, component tests, scratch tests,
and examples run. Performance tests are not run.

* fix(cstring): convert wide output without deprecated wctomb

MSVC rejects wctomb as a deprecated CRT call, and a null destination is not
a valid reset. Wide writes now keep a local mbstate_t and call wcrtomb_s
there, and wcrtomb elsewhere.
Label the hash, comparison, and wide-string work as 4.1.0-alpha1 dated 6th
October 2026. Drop the duplicate hash-shim and unit-test bullets, and move
the --wide-strings note out of the released 4.0.19 section. Point NEWS and
the Doxyfile project number at the same tag.
@synesissoftware
synesissoftware merged commit 810307c into master Oct 6, 2026
46 checks passed
@synesissoftware
synesissoftware deleted the dev branch October 6, 2026 20:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants