Skip to content

Add knowledge language policy and preserve Markdown discourse - #70

Open
tenphi wants to merge 77 commits into
mainfrom
feat/knowledge-language-discourse
Open

Add knowledge language policy and preserve Markdown discourse#70
tenphi wants to merge 77 commits into
mainfrom
feat/knowledge-language-discourse

Conversation

@tenphi

@tenphi tenphi commented Sep 7, 2026

Copy link
Copy Markdown
Owner

An English knowledge base could follow the language of incoming sources, and ordinary Markdown could lose the context that made a claim hypothetical. This change makes generated knowledge language explicit and preserves qualification through retention, retrieval, answers and inference. Addresses #61 and #62.

  • Configure English knowledge language independently of English/Russian answer language. Preserve original source bytes, exact provided-candidate attestations and policy receipts.
  • Preserve speakers, nested reports, hypotheses, fiction, counterfactuals, questions, lifecycle status and source-relative clocks. Keep qualified content out of factual derivation and automatic graph inference.
  • Require consistent proposition, action-argument and qualification-scope verdicts. Bind retention to original source spans and answer interpretation to cited records, with a separately required selection verdict for private original-source frames. Missing or negative dimensions withhold output; semantic rejection is never retried.
  • Apply language policy to generated maintenance, corrections and curator revisions, with typed failures. Preserve exact authorized moves and source references.
  • Record frozen invented corpora, independent input/output judgments and fingerprinted acceptance gates. Runtime is GPT-5.6 Luna; separate grading and code/forensic reviews use GPT-5.6 Sol.

Validation remains in progress. The completed V53 exposed evaluation scored 57/80 independently useful answers: 32/48 selected and 25/32 built-package, with twenty-one writable abstentions and two accepted qualification/action-specificity errors. This regresses from V52's 65/80 and zero accepted errors. Complete useful retention is 8/10 cases; useful retrieval is 72/80 answer rows. Source bytes and case availability pass. All reports, source-only grades, initial/corrected forensic reviews and the technical-paraphrase convention audit are preserved.

Current frozen runtime 9fd704d (V54) passed 2,430 tests across 139 files, build/typecheck, lint, knip, formatting, documentation doctor/build, smoke, installed-package smoke and repository safety. Two separate Sol code review/fix rounds are complete. Build/restart/socket deployment, compiled changed-behavior controls and actual-provider schema transport controls passed.

V54 adds private original-record readings before answer drafting and exact-source comparisons of actors, objects/mechanisms and qualifications inside existing verifier calls. Readings are discarded before independent verification; every existing semantic and retained-excerpt selection verdict remains mandatory. Bounded source-clock/hypothesis wording fixes address reproduced false holds, while retention preserves the original coverage object. No model change, new call, semantic retry or source-authority shortcut is introduced.

The answer-role default ceiling is now 2,400 output tokens, with possible latency/cost increases for fuller audits. Both declared exposed probes are running once at that explicit isolated trial setting under the V54 validation plan. Resolved ceilings are recorded and bound to review fingerprints. The existing service overlay remains 1,024 tokens; results at 2,400 will not establish reliability at its lower limit. No service configuration is overwritten.

Fresh approved V20 held-out remains unexecuted because accepted errors and substantial exposed coverage losses defer full validation. The latest complete V48 repeated trial scored 255/320, failing the unchanged gate. Every started evaluation and review correction remains in the evaluation history.

Acceptance still requires at least 90% independently useful answers in every split/run, at least 80% useful retention/retrieval, zero accepted source-entailment/qualification/language/promotion/source-byte errors and at most 5% case availability failures. No unchanged-runtime semantic retry, model change or threshold weakening is introduced. PR #70 remains open and unmerged until that gate is met. Finite invented English/Russian scenarios do not establish arbitrary-language or longitudinal reliability.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant