Add reproducible eight-model architecture research with evidence-linked comparison - #52
Open
zbzbdzb wants to merge 2 commits into
Open
Add reproducible eight-model architecture research with evidence-linked comparison#52zbzbdzb wants to merge 2 commits into
zbzbdzb wants to merge 2 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This adds the research packet for #5 using eight actual local model generations: Spark, Qwen, Phi, SmolLM, Granite, Falcon, TinyLlama and Danube. It preserves the pinned revisions, exact prompts, token IDs, generation settings and unedited outputs, with 88 source-linked comparison rows and a concrete synthesis tied to the current runtime.
The focus is compact models on a 12 GB workstation. These are eight named model projects, not eight independent lineages or frontier services. Three outputs hit the recorded token cap. The analysis keeps their omissions and unsupported claims visible—including invented benchmark numbers—and does not present those claims as results. The synthesis turns the useful parts into bounded planning, versioned authority and observable file effects. An isolated 14-case process-crash experiment illustrates the recovery decisions; it is not a test of Cognitive-OS's production recovery.
Validation on final commit
12d2bccb5c93686c3053838660b6605b8d7510af:resourcemodule; this is documented. No tests were disabled, and the Linux checks above provide the required repository validation.Research-only change under
research/ai_generated_agi_architectures/; no runtime APIs, adapter dependencies or existing workflows change. No model weights, credentials or account screenshots are included.