feat: add visual Agent Eval system - #68
Merged
Merged
Conversation
lyyzka
force-pushed
the
codex/agent-eval-dashboard
branch
from
August 26, 2026 10:06
720bd32 to
1bfcbd4
Compare
Contributor
Author
|
@leecyang 第二轮 review 的 3 个阻塞项已在 |
Contributor
Author
|
已按本轮 review 将 Eval CI selector 改为 fail-closed,并推送 commit
本地验证已完成:lint、前后端 typecheck、guards、version check、Eval unit/runtime gates、全量 283 unit、build、Open Notebook 7/7、Skill validation。正在等待完整 PR CI。 |
Contributor
Author
|
@leecyang 本轮唯一阻塞项已修复并完成全矩阵验证,麻烦进行最终复审。 最终 commits:
最终 CI run
当前 PR 状态 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
knowledge.searchcitationsReview follow-up
$lingxiloop-verify-changeEval classification and$lingxiloop-eval-changeguidance for suites, baselines, deterministic/model Eval, Trace sanitization, comparison, and verificationeval:harness, and addeval:runtime: three deterministic, network-free Cases run the currentAgentOSRuntimethroughMemoryHostAdapter,ScriptedModelDriver, IPython/Host Bridge, dynamic RAG, and Approval before evaluationnpm run test:integration:evalevalFocusedfail closed: it is true only when every changed path is Eval-owned; shared Agent OS, migration, API/Admin, and integration infrastructure paths restore their owning testsfullMatrix=truefor package manifests, workflows, and the classifier itself, so CI/dependency selector changes are exercised before their focused result is trustedruntime.ts,migrate.ts,admin-router.ts, package manifests, workflow/classifier paths, and integration infrastructureVerification
npm run lintnpm run typechecknpm run server:typechecknpm run guard:brandnpm run guard:agent-osnpm run guard:llm-trackednpm run version:checknpm run test:eval— 14 passednpm run eval:check— frozen harness and real Agent OS runtime gates both passed at 100%npm test— 283 passednpm run buildevalFocused=false,fullMatrix=true, andintegration=fulllingxiloop-eval-changeSkill validation passedeval.test.tsand reject unknown files; database-backed execution is owned by PR CI because this machine has no configured PostgreSQL/RedisCloses #67