Repository navigation
Hash review - #44
Merged
Merged
Hash review#44
Conversation
Gate <stdint.h> on C99, on C++ with that header or C++11, and on MSVC 7.1 or later, and typedef cstring_hash_t to uint64_t when it is included; use cstring_hash_t for the djb2 and FNV-1a API and for the hash unit tests.
Move the eight C hashing functions into **test.unit.hash** (`entry.c`) so the suite builds when `NO_CSTRING_CPP_API` is on. Keep the C++ shim checks in **test.unit.hash.cxx**, which stays behind that option.
Cover octet 0xFF and an interior NUL in test.unit.hash. Exact narrow vectors stay off the wide build so a later wchar_t fix is not frozen.
Add CSTRING_HASH_DJB2_SEED, CSTRING_HASH_FNV1A_OFFSET, and CSTRING_HASH_FNV1A_PRIME beside cstring_hash_t, and use them as the only copies of those constants in the hash implementation. Cite Yigit's djb2 page and the Fowler, Noll, and Vo notes and RFC 9923, and lock each macro to its historic literal in the C unit test.
Rename the case-folded hash functions from the `_ci` suffix to `_case` (`cstring_hash_djb2_case`, `cstring_hash_djb2_len_case`, and the FNV-1a twins), including the C++ shims and the unit tests.
Rename the pointer-and-length hash functions from `_len` to `_buf` (`cstring_hash_djb2_buf`, `cstring_hash_djb2_buf_case`, and the FNV-1a twins), including the C++ shims and the unit tests.
mwsis
approved these changes
Oct 6, 2026
synesissoftware
added a commit
that referenced
this pull request
Oct 6, 2026
* Performance tests (#41) * fix(performance tests): "assign_len_grow" now more realistic * fix(performance tests): "create_destroy_empty" now more realistic * fix(performance tests): "append_len_growth_x" now more realistic * fix(performance tests): "append_len_reserved" now more realistic * Revert "fix(performance tests): "create_destroy_empty" now more realistic" This reverts commit 5998491. * fix(performance tests): "create_destroy_empty" std::string now constructs "" The cstring row calls cstring_create(&s, "") and raw_realloc assigns an empty C string. Default-constructing std::string skipped that step, so the row compared a materialised empty string with an untouched default object. Construct std::string from "" so all three sides perform the same creation. Small-string storage can still make std::string cheaper; that gap is the library. * fix(performance tests): assign and copy anchors read the first byte assign_len_grow and copy returned only the length, which the compiler can treat as invariant, so it deleted the visible raw_realloc memcpy. Rows from 256 bytes upward stayed near 12 ns with p50 of 0. Each implementation now adds the first payload byte to the anchor, as create_destroy_len already does, including the Windows arena lambdas. * fix(performance tests): append and insert anchors read the first byte append_len_growth, append_len_reserved, and insert_len_mid returned only the length, which the compiler can treat as invariant, so it can delete the visible raw_realloc copies. Each implementation now adds the first payload byte to the anchor, as assign_len_grow and copy already do, including the Windows arena lambdas. Every size in these scenarios leaves a non-empty string, so the first byte is in range. * fix(unit tests): COM initialisation * fix(performance tests): append_batch anchor reads stored payload append_batch returned only the vector length, which is the constant element count, so the compiler can delete the visible vector<string> insert. The payload-16 row printed 0 ns/op. Each implementation now takes the address of the first and last stored string through a volatile pointer and adds the first byte of each. That address is the copy, not the source, so those two strings have to be materialised. * fix(performance tests): one-by-one vector anchors read stored payload append_one_by_one and prepend_one_by_one returned only the vector length, which is the constant element count, so the compiler can delete the visible vector<string> updates. Each implementation now takes the address of the first and last stored string through a volatile pointer and adds the first byte of each, as append_batch already does. create_destroy is unchanged: its elements are empty, so there is no stored payload byte to read. * fix(performance tests): Windows assign/append ratios use createEx realloc On Windows, assign_len_grow, append_len_growth, and append_len_reserved arena rows call createEx("") before the mutation. The printed cstring row does not. Those arena ratios now use cstring_win_realloc, which runs the same body with CSTRING_F_USE_REALLOC. std::string and raw_realloc stay against the cstring row that skips the empty create. create, insert, and copy are unchanged: their cstring row already matches the arena body. * fix(performance tests): borrowed rows time a prebuilt buffer borrowed_fixed_construct and borrowed_fixed_assign built the caller buffer inside the timed loop, so sizes above 511 bytes included a heap allocation in the borrowed and fixed_char_buf times. The buffer is now built once per size and reused. std::string on those rows is an owning construct, and its ratio column is "-" rather than a comparison with cstring_borrowed. fixed_char_buf remains the like-for-like peer. * fix(performance tests): file-line peers share an owning anchor file_read_all_lines compared an owning fgetc reader with a getline loop that discarded its buffer on every move, and with platformstl::file_lines, which stores views over one copy of the file. getline now copies each line and clear()s the scratch string, and the stream is opened binary. All three anchors are the line count plus the sum of the line lengths, and a mismatch drops the rows. file_lines is still timed, and its ratio column is "-" because it does not own each line. ifstream+getline remains the owning peer. * fix(performance tests): dash the ratio when a row is elided When p50 is 0 and ns/op has not risen against the previous smaller size of the same scenario and implementation, vs cstr is "-". A deleted raw_realloc copy was printing a falling ratio (down to 0.02) while the time stayed flat. The first such size is dashed as well. * fix(performance tests): Windows copy calls cstring_copy The Windows copy rows built the destination with createLenEx of the source bytes, so they never called cstring_copy. They now createEx("") to select the arena, then cstring_copy, which allocates with the destination flags. The printed cstring row does not createEx(""), so those arena ratios use cstring_win_realloc. std::string and raw_realloc stay on the cstring row. * fix(performance tests): raw vector create leaves slots empty create_destroy's raw floor reserved the pointer array, then calloc'd a byte per slot. cstring_vector_create and vector<string>(n) only construct empty elements, with no character buffer. The raw slots are now null. Destroy still frees them, and free(NULL) is a no-op. * fix(performance tests): fgets honours the readline instance mode read_all_fgets always used one buffer of line_len+4, so fresh and pre-reserved did not differ and reuse never split a long line. Fresh now allocates that buffer per line, pre-reserved keeps one, and reuse keeps a 4096-byte buffer and stitches a short read, including a CR held across a chunk boundary. Anchors stay the line length plus the first content byte. * fix(performance tests): time ns/op as one interval When p99 is linked, ns/op was the sum of per-call samples, and each sample included start/stop. On short rows that overhead flattened the ratio. ns/op and vs cstr now use one start/stop around the iteration loop. Percentiles still come from a per-call pass on the recorded warmup, and that pass is not added to ns/op. * fix(performance tests): label percentiles as per iteration p50, p90, p99, and max are the median and tails of one iteration. ns/op divides the batch by #acts, so an append or readline percentile was the whole trial while ns/op was one append or one line. The columns are now p50/iter, p90/iter, p99/iter, and max/iter. The values stay one iteration. * fix(performance tests): prepend at payload 256 as well as 16 prepend_one_by_one ran only at payload 16, where vector<string> is still inside small-string storage. It now also runs at payload 256 and still prepends 32 elements, beside the append rows. The 16-byte row remains. * Hashing API (#19) * general boilerplate / project-support improvements (#16) * squash-commit * 4.0.14 * fix * .vscode/settings.json * Add cstring_hash_djb2() and cstring_hash_fnv1a() hashing API and test suite * Added 64-bit djb2 and FNV-1a hash functions for cstring_t instances and slices (cstring_hash_djb2, cstring_hash_djb2_ci, cstring_hash_djb2_len, cstring_hash_djb2_len_ci, cstring_hash_fnv1a, cstring_hash_fnv1a_ci, cstring_hash_fnv1a_len, cstring_hash_fnv1a_len_ci); * Added cstring_hash_t typedef (uint64_t) with fallback for older MSVC; * Added C++ inline access shims for instances and buffer slices; * Added comprehensive xTests unit test suite test.unit.cstring.hash; * Verified both narrow-character and wide-character build paths; * Documented hashing API, types, and C++ access shims in README.md; * Bumped version to 4.1.0-alpha1 and updated CHANGES.md and NEWS.md; * 4.0.15 (#20) * Updated all `\param` and `\retval` Doxygen documentation comments in `cstring.h`, `cstring.vector.h`, and `cstring.core.c` to obey the 76-rule and terminate with a semicolon; * Bumped version to 4.0.15 in `cstring.h`, `Doxyfile`, `CHANGES.md`, and `NEWS.md`; * Updated `Updated:` header banners to 6th September 2026; * squash-commit * NEWS.md * chore: renamed **test/unit/test.unit.cstring.hash/entry.cpp** => **test/unit/hash/entry.cpp** * chore --------- Co-authored-by: synesissoftware <matthew@synesis.com.au> * Hash review (#44) * chore(misc): language discrimination; CI branch(es) * chore(misc): ensuring `\file` designators are correct * chore(mist): removed unused `CSTRING_VECTOR_ASSERT()` * chore(misc): tidying * fix(cstring): discriminate cstring_hash_t and use it for hashing Gate <stdint.h> on C99, on C++ with that header or C++11, and on MSVC 7.1 or later, and typedef cstring_hash_t to uint64_t when it is included; use cstring_hash_t for the djb2 and FNV-1a API and for the hash unit tests. * chore(performance tests): prefer `uint64_t` to `std::uint64_t` * test(cstring): exercise the hashing API from a C unit test Move the eight C hashing functions into **test.unit.hash** (`entry.c`) so the suite builds when `NO_CSTRING_CPP_API` is on. Keep the C++ shim checks in **test.unit.hash.cxx**, which stays behind that option. * test(cstring): lock high-byte and embedded-NUL hash results Cover octet 0xFF and an interior NUL in test.unit.hash. Exact narrow vectors stay off the wide build so a later wchar_t fix is not frozen. * feat(cstring): publish the djb2 and FNV-1a basis macros Add CSTRING_HASH_DJB2_SEED, CSTRING_HASH_FNV1A_OFFSET, and CSTRING_HASH_FNV1A_PRIME beside cstring_hash_t, and use them as the only copies of those constants in the hash implementation. Cite Yigit's djb2 page and the Fowler, Noll, and Vo notes and RFC 9923, and lock each macro to its historic literal in the C unit test. * fix * refactor(cstring): rename the case-folded hash suffix to _case Rename the case-folded hash functions from the `_ci` suffix to `_case` (`cstring_hash_djb2_case`, `cstring_hash_djb2_len_case`, and the FNV-1a twins), including the C++ shims and the unit tests. * refactor(cstring): rename the counted hash suffix from _len to _buf Rename the pointer-and-length hash functions from `_len` to `_buf` (`cstring_hash_djb2_buf`, `cstring_hash_djb2_buf_case`, and the FNV-1a twins), including the C++ shims and the unit tests. * fix * Hash API (redux) (#45) * chore(misc): language discrimination; CI branch(es) * chore(misc): ensuring `\file` designators are correct * chore(mist): removed unused `CSTRING_VECTOR_ASSERT()` * chore(misc): tidying * fix(cstring): discriminate cstring_hash_t and use it for hashing Gate <stdint.h> on C99, on C++ with that header or C++11, and on MSVC 7.1 or later, and typedef cstring_hash_t to uint64_t when it is included; use cstring_hash_t for the djb2 and FNV-1a API and for the hash unit tests. * chore(performance tests): prefer `uint64_t` to `std::uint64_t` * test(cstring): exercise the hashing API from a C unit test Move the eight C hashing functions into **test.unit.hash** (`entry.c`) so the suite builds when `NO_CSTRING_CPP_API` is on. Keep the C++ shim checks in **test.unit.hash.cxx**, which stays behind that option. * test(cstring): lock high-byte and embedded-NUL hash results Cover octet 0xFF and an interior NUL in test.unit.hash. Exact narrow vectors stay off the wide build so a later wchar_t fix is not frozen. * feat(cstring): publish the djb2 and FNV-1a basis macros Add CSTRING_HASH_DJB2_SEED, CSTRING_HASH_FNV1A_OFFSET, and CSTRING_HASH_FNV1A_PRIME beside cstring_hash_t, and use them as the only copies of those constants in the hash implementation. Cite Yigit's djb2 page and the Fowler, Noll, and Vo notes and RFC 9923, and lock each macro to its historic literal in the C unit test. * fix * refactor(cstring): rename the case-folded hash suffix to _case Rename the case-folded hash functions from the `_ci` suffix to `_case` (`cstring_hash_djb2_case`, `cstring_hash_djb2_len_case`, and the FNV-1a twins), including the C++ shims and the unit tests. * refactor(cstring): rename the counted hash suffix from _len to _buf Rename the pointer-and-length hash functions from `_len` to `_buf` (`cstring_hash_djb2_buf`, `cstring_hash_djb2_buf_case`, and the FNV-1a twins), including the C++ shims and the unit tests. * refactor(cstring): .DEF file renames of hash functions * feat(cstring): add always-available multibyte and wide hash API Move the hash declarations into hash.h and the stdint.h discrimination into common.h. cstring.h includes both. Add mbs, wcs, mbuf, and wbuf entry points, including the case forms, for djb2 and FNV-1a. Each code unit contributes its low 8 bits, so ASCII text has one hash in both encodings. C++ shims overload on char const*, wchar_t const*, and both buffer forms. * docs(cstring): document hash groups, contract, and consumers - Document the hash contract, the djb2 and FNV-1a steps, and the published vectors in hash.h, README.md, and doc/mainpage.md; - Define group__cstring_api__hashing__djb2 and group__cstring_api__hashing__fnv1a, and keep the published references on those groups; - Place the C++ hash shims in namespace cstring, and import them in test.unit.hash.cxx; - Name the projects that call or link cstring, including the sistools programs; - Parenthesise the hash basis macros and note the FNV-1a prime in decimal; - Record the header-split and auto-buffer provenance items in TODO.md; - Spell NUL in the hash and status-code comments; * Hash (SDBM) (#46) * chore(misc): language discrimination; CI branch(es) * chore(misc): ensuring `\file` designators are correct * chore(mist): removed unused `CSTRING_VECTOR_ASSERT()` * chore(misc): tidying * fix(cstring): discriminate cstring_hash_t and use it for hashing Gate <stdint.h> on C99, on C++ with that header or C++11, and on MSVC 7.1 or later, and typedef cstring_hash_t to uint64_t when it is included; use cstring_hash_t for the djb2 and FNV-1a API and for the hash unit tests. * chore(performance tests): prefer `uint64_t` to `std::uint64_t` * test(cstring): exercise the hashing API from a C unit test Move the eight C hashing functions into **test.unit.hash** (`entry.c`) so the suite builds when `NO_CSTRING_CPP_API` is on. Keep the C++ shim checks in **test.unit.hash.cxx**, which stays behind that option. * test(cstring): lock high-byte and embedded-NUL hash results Cover octet 0xFF and an interior NUL in test.unit.hash. Exact narrow vectors stay off the wide build so a later wchar_t fix is not frozen. * feat(cstring): publish the djb2 and FNV-1a basis macros Add CSTRING_HASH_DJB2_SEED, CSTRING_HASH_FNV1A_OFFSET, and CSTRING_HASH_FNV1A_PRIME beside cstring_hash_t, and use them as the only copies of those constants in the hash implementation. Cite Yigit's djb2 page and the Fowler, Noll, and Vo notes and RFC 9923, and lock each macro to its historic literal in the C unit test. * fix * refactor(cstring): rename the case-folded hash suffix to _case Rename the case-folded hash functions from the `_ci` suffix to `_case` (`cstring_hash_djb2_case`, `cstring_hash_djb2_len_case`, and the FNV-1a twins), including the C++ shims and the unit tests. * refactor(cstring): rename the counted hash suffix from _len to _buf Rename the pointer-and-length hash functions from `_len` to `_buf` (`cstring_hash_djb2_buf`, `cstring_hash_djb2_buf_case`, and the FNV-1a twins), including the C++ shims and the unit tests. * refactor(cstring): .DEF file renames of hash functions * feat(cstring): add always-available multibyte and wide hash API Move the hash declarations into hash.h and the stdint.h discrimination into common.h. cstring.h includes both. Add mbs, wcs, mbuf, and wbuf entry points, including the case forms, for djb2 and FNV-1a. Each code unit contributes its low 8 bits, so ASCII text has one hash in both encodings. C++ shims overload on char const*, wchar_t const*, and both buffer forms. * docs(cstring): document hash groups, contract, and consumers - Document the hash contract, the djb2 and FNV-1a steps, and the published vectors in hash.h, README.md, and doc/mainpage.md; - Define group__cstring_api__hashing__djb2 and group__cstring_api__hashing__fnv1a, and keep the published references on those groups; - Place the C++ hash shims in namespace cstring, and import them in test.unit.hash.cxx; - Name the projects that call or link cstring, including the sistools programs; - Parenthesise the hash basis macros and note the FNV-1a prime in decimal; - Record the header-split and auto-buffer provenance items in TODO.md; - Spell NUL in the hash and status-code comments; * test(cstring): time djb2 and FNV-1a against a lose-lose baseline Add test.performance.hash. Each group is one scenario and length, with lose-lose, djb2, and FNV-1a, across cstring_t, multibyte and wide strings, and counted buffers, including case folding, embedded NULs, and high octets. A further algorithm is one extra row on each table. Stop cstring_strlcpy_safe_() from scanning past lim. createLen, assign, insert, and append pass a slice of cch characters that need not contain a NUL. Cover that with a 3-character createLen case in test.unit.cstring. * test(cstring): report hash spread from a scratch program Add test.scratch.hash. It builds 10000 distinct pseudo-random strings and prints full-hash uniqueness, per-octet occupancy, and residues for power-of-two hashtable sizes and a prime just below each, for lose-lose, djb2, and FNV-1a. * feat(cstring): add a 64-bit SDBM hash Add `cstring_hash_sdbm()` and the multibyte, wide, counted-buffer, and `_case` forms. The step is `hash * 65599 + octet` from seed 0, in a 64-bit accumulator. C++ shims are `hash_sdbm()` and `hash_sdbm_case()`. Unit tests lock the published vectors; the performance and scratch programs append an SDBM row after djb2 and FNV-1a. * fix * fix(cstring): hash every little-endian octet of a wide code unit Step once per octet of each wchar_t, low byte first, in the wide djb2, FNV-1a, and SDBM functions and in the lose-lose timing baseline. A multibyte string and its wide equivalent now differ; the published multibyte vectors stay as they are, and 64-bit djb2 is documented to match 32-bit djb2 only below 2^32. * fix: corrected old-VC++ compatibility * fix(cstring): recognise stdint.h on Visual C++ from VS 2010 Require _MSC_VER 1600 before treating Visual C++ as having <stdint.h>. Older MSVC keeps the unsigned __int64 fallback for cstring_hash_t. * test(cstring): assert _case hashes ignore setlocale Compare C-locale and tr_TR.UTF-8 _case results for octet 0xDD and U+00C0. Those inputs change under tolower and towlower, so the case fails until the fold stops following the process locale. * fix(cstring): fold _case hashes on ASCII A-Z only Map A-Z to a-z by adding 0x20 and leave every other code unit unchanged, in the djb2, FNV-1a, and SDBM functions and in the lose-lose baseline. setlocale no longer changes a _case result. * feat(cstring): specialise std::hash<cstring_t> on FNV-1a Provide std::hash<cstring_t> from C++11 as the FNV-1a hash of ptr and len, converted to size_t. djb2 and SDBM stay available as explicit functions. * test(cstring): rename hash tests off the old _ci and _len suffixes Rename the case and counted-buffer unit-test functions to _case and _buf, matching the hash API. Local bases in the case test follow the same suffix. * feat(cstring): add cstring_equal and cstring_compare Add cstring_equal(), returning cstring_truthy_t, and cstring_compare(), returning cstring_sint_t. Equal returns before reading when the lengths differ, and memcmp's the payloads only when they match. Compare returns a signed order of len code units. The multibyte build memcmp's the common prefix. C++ ==, !=, and < call those functions. Drop the cstring_vector example's private comparator. Its qsort adapters call cstring_compare(). * test(cstring): tolerate a host with no Turkish ctype locale Select tr_TR.UTF-8 or another Turkish ctype name when the host has one. A missing Turkish locale is not a failure. Linux CI generates tr_TR.UTF-8 so that comparison still runs there. * feat(cstring): build and test `cstring_char_t` as `wchar_t` (#43) * refactor(cstring): keep the character helper under test/ Move cstring.helpers.h from include/cstring/ to test/. Examples include it by relative path, and the test suites include it from cstring_testing.h. It is not part of the installed library contract. * fix * fix(cstring): check the wctomb() shift-state result GCC rejects a discarded wctomb() result, including a (void) cast. Test the NULL call that restores the initial shift state. * ci(cstring): run wide-string Windows cells without performance tests Add CI cells windows-cl-wide and windows-mingw-wide. They call the shared cell with --wide-strings, so unit tests, component tests, scratch tests, and examples run. Performance tests are not run. * fix(cstring): convert wide output without deprecated wctomb MSVC rejects wctomb as a deprecated CRT call, and a null destination is not a valid reset. Wide writes now keep a local mbstate_t and call wcrtomb_s there, and wcrtomb elsewhere. * docs(cstring): record the 4.1.0-alpha1 release notes Label the hash, comparison, and wide-string work as 4.1.0-alpha1 dated 6th October 2026. Drop the duplicate hash-shim and unit-test bullets, and move the --wide-strings note out of the released 4.0.19 section. Point NEWS and the Doxyfile project number at the same tag. --------- Co-authored-by: Matt Wilson <152443343+mwsis@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Numerous Hash API related feedback improvements.