Skip to content

Add reproducible eight-model architecture research with evidence-linked comparison - #52

Open
zbzbdzb wants to merge 2 commits into
aLexzzz430:mainfrom
zbzbdzb:research/local-architecture-comparison
Open

Add reproducible eight-model architecture research with evidence-linked comparison#52
zbzbdzb wants to merge 2 commits into
aLexzzz430:mainfrom
zbzbdzb:research/local-architecture-comparison

Conversation

@zbzbdzb

@zbzbdzb zbzbdzb commented Sep 9, 2026

Copy link
Copy Markdown

This adds the research packet for #5 using eight actual local model generations: Spark, Qwen, Phi, SmolLM, Granite, Falcon, TinyLlama and Danube. It preserves the pinned revisions, exact prompts, token IDs, generation settings and unedited outputs, with 88 source-linked comparison rows and a concrete synthesis tied to the current runtime.

The focus is compact models on a 12 GB workstation. These are eight named model projects, not eight independent lineages or frontier services. Three outputs hit the recorded token cap. The analysis keeps their omissions and unsupported claims visible—including invented benchmark numbers—and does not present those claims as results. The synthesis turns the useful parts into bounded planning, versioned authority and observable file effects. An isolated 14-case process-crash experiment illustrates the recovery decisions; it is not a test of Cognitive-OS's production recovery.

Validation on final commit 12d2bccb5c93686c3053838660b6605b8d7510af:

  • Original CI: repository boundary check and all 10 public smoke tests pass on both Linux Python 3.10 and 3.11.
  • Additional pre-submission checks: the same public checks, packet verification and isolated crash experiment on both Python versions. This fork-only workflow pins the final candidate SHA.
  • Local tokenization checks reproduce all eight recorded prompts, input IDs and decoded output IDs; all 12 local weight-shard hashes match cached pinned-snapshot LFS metadata. The packet verifier checks 64 file hashes and all 88 quoted source references.
  • Windows public-test collection cannot run because existing code imports the Unix-only resource module; this is documented. No tests were disabled, and the Linux checks above provide the required repository validation.

Research-only change under research/ai_generated_agi_architectures/; no runtime APIs, adapter dependencies or existing workflows change. No model weights, credentials or account screenshots are included.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant