Add knowledge language policy and preserve Markdown discourse - #70
Open
tenphi wants to merge 77 commits into
Open
Add knowledge language policy and preserve Markdown discourse#70tenphi wants to merge 77 commits into
tenphi wants to merge 77 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
An English knowledge base could follow the language of incoming sources, and ordinary Markdown could lose the context that made a claim hypothetical. This change makes generated knowledge language explicit and preserves qualification through retention, retrieval, answers and inference. Addresses #61 and #62.
Validation remains in progress. The completed V53 exposed evaluation scored 57/80 independently useful answers: 32/48 selected and 25/32 built-package, with twenty-one writable abstentions and two accepted qualification/action-specificity errors. This regresses from V52's 65/80 and zero accepted errors. Complete useful retention is 8/10 cases; useful retrieval is 72/80 answer rows. Source bytes and case availability pass. All reports, source-only grades, initial/corrected forensic reviews and the technical-paraphrase convention audit are preserved.
Current frozen runtime 9fd704d (V54) passed 2,430 tests across 139 files, build/typecheck, lint, knip, formatting, documentation doctor/build, smoke, installed-package smoke and repository safety. Two separate Sol code review/fix rounds are complete. Build/restart/socket deployment, compiled changed-behavior controls and actual-provider schema transport controls passed.
V54 adds private original-record readings before answer drafting and exact-source comparisons of actors, objects/mechanisms and qualifications inside existing verifier calls. Readings are discarded before independent verification; every existing semantic and retained-excerpt selection verdict remains mandatory. Bounded source-clock/hypothesis wording fixes address reproduced false holds, while retention preserves the original coverage object. No model change, new call, semantic retry or source-authority shortcut is introduced.
The answer-role default ceiling is now 2,400 output tokens, with possible latency/cost increases for fuller audits. Both declared exposed probes are running once at that explicit isolated trial setting under the V54 validation plan. Resolved ceilings are recorded and bound to review fingerprints. The existing service overlay remains 1,024 tokens; results at 2,400 will not establish reliability at its lower limit. No service configuration is overwritten.
Fresh approved V20 held-out remains unexecuted because accepted errors and substantial exposed coverage losses defer full validation. The latest complete V48 repeated trial scored 255/320, failing the unchanged gate. Every started evaluation and review correction remains in the evaluation history.
Acceptance still requires at least 90% independently useful answers in every split/run, at least 80% useful retention/retrieval, zero accepted source-entailment/qualification/language/promotion/source-byte errors and at most 5% case availability failures. No unchanged-runtime semantic retry, model change or threshold weakening is introduced. PR #70 remains open and unmerged until that gate is met. Finite invented English/Russian scenarios do not establish arbitrary-language or longitudinal reliability.