Skip to content

fix(ci): use Sol for triage and Sonnet 5 for PR evaluations - #1154

Merged
AbhitejJohn merged 6 commits into
mainfrom
abhitejjohn-agentic-workflow-repair
Sep 12, 2026
Merged

AbhitejJohn merged 6 commits into
mainfrom
abhitejjohn-agentic-workflow-repair

Conversation

@AbhitejJohn

@AbhitejJohn AbhitejJohn commented Sep 10, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Use separate defaults for operational workflows and skill evaluations:

Path Default
devops-health-check, devops-health-groom, devops-health-investigate, issue-triage gpt-5.6-sol
Default PR skill evaluation profile claude-sonnet-5 and unchanged gpt-5.6-luna

Regenerate the four agentic lock files with gh-aw v0.86.2. Preserve model-variable precedence, the copilot-pat-pool environment, secret mappings, permissions, prompts, safe outputs, and action/container pins. Replace Sonnet 4.6 in the default and full evaluation presets and use the existing Sonnet 5 -> Terra judge route. Keep Luna, explicit profile selection, and the newer preset unchanged. Correct the stale Mon/Wed/Fri schedule comment identified by Copilot.

Also repair concurrent SDK startup in the trusted evaluation launcher. This adds harness startup code, regression tests, launcher wiring in evaluation-run.yml, secretless test wiring in evaluation-workflow-tests.yml, and investigation guidance. Dependency manifests and lockfiles are unchanged by this repair.

Judge replacement is deferred. GPT executors retain Opus 4.8 as the primary judge and Haiku 4.5 as the scheduled secondary judge; PR and manual runs retain no secondary judge. No judge switch is included.

Actual PR evaluation path

evaluation.yml selects the default profile for ordinary PR events and unqualified /evaluate requests. Discovery expands that profile into executor/judge entries. The evaluate job passes those entries to evaluation-run.yml, which uses matrix.entry.model and matrix.entry.judge in evaluation commands. This changes the PR evaluator, not only the issue-triage model.

SDK startup repair

Proven local defects: the pinned SDK can start multiple transports for concurrent callers, and createSession/resumeSession can use a connected transport before filesystem-provider setup finishes. The no-model regression command node --test eng/evaluation-tools/sdk-startup.test.mjs originally failed four cases: three runtime starts instead of one, early session.create, early session.resume, and duplicate starts before a failed-start retry. After the guard, all six cases pass on SDK 1.0.11 and 1.0.13, including independent-client startup and parallel session creation after readiness. The tests use the real SDK methods with controlled transport setup, not live model calls. Upstream startup work is tracked in github/copilot-sdk#2585.

Run evidence: run 34627963458 reached Sonnet 5 execution, but both Sonnet 5 and unchanged Luna reported Cannot set session filesystem provider while sessions are active. The advanced-plugin legs each recorded expectedEvalCount=4, writtenResultCount=4, and measurementInvalidEvalCount=4. This is consistent with the reproduced startup defects; a complete repaired evaluation is still needed to confirm the end-to-end result.

eng/evaluation-tools/sdk-startup.mjs shares one startup promise per SDK client and waits for readiness before creating or resuming sessions. Startup errors propagate to every waiter; a later attempt can retry. It accepts only the two tested SDK versions and fails explicitly on an unknown version. vally.mjs loads this guard before the Vally CLI. The workflow copies both files from the trusted tooling checkout and places the launcher first on PATH, so evaluation and comparison use the same guard. No global SDK installation or dependency source is modified. Credential scope, trusted-checkout boundaries, model/judge routing, trial counts, concurrency, and result gates remain unchanged.

Core repair commit: a226ead6f. Current correction: b8083dd46, which also covers SDK 1.0.11 selected by the PR merge checks, while this branch uses SDK 1.0.13. Evaluation 34652125571 was started on a226ead6f; evaluation 34652384279 is bound to current head b8083dd46. Both were in progress at this update. Neither is claimed to be a passing evaluation.

The separate publish-session-data failure reports Authentication failed for https://github.com/dotnet/skills-data.git/. This repair does not change or repair that credential.

Judge experiment evidence — decision deferred

Public comparison records pinned to commit 27b203fd cover publication timestamps August 26–September 8, 2026 UTC. The experiment originated in #1002; scheduled run 34094126994 contains primary and cross-judge artifacts.

Filter actual row identifiers primaryJudge=claude-opus-4.8 and secondJudge=claude-haiku-4.5; the top-level label still names Sonnet 4.6 and is stale. The data contains 1,101 paired observations across 12 publication snapshots, 8 commits, 92 distinct plugin/skill keys, and 3 GPT executors. These are repeated observations, not independent trials.

Cohort Pass/fail agreement State + preference-regression agreement
All retained pairs 956/1,101 = 86.83% 945/1,101 = 85.83%
Both valid, conclusive, not underpowered 452/595 = 75.97% 443/595 = 74.45%
Luna, both valid/conclusive/powered 216/298 = 72.48% 212/298 = 71.14%
Codex 5.3, both valid/conclusive/powered 168/199 = 84.42% 164/199 = 82.41%
Sol, both valid/conclusive/powered 68/98 = 69.39% 67/98 = 68.37%

In the valid cohort, 109 observations pass with Opus but return VALID_NO_CHANGE with Haiku; 34 go the other way. Haiku produces fewer passes in this sample. These disagreements are not proven Haiku errors because there is no human reference verdict.

This does not prove equivalence. The overall count includes 502 pairs where both results are INVALID_INCONCLUSIVE. Repeated skills, changing commits, and selection of the valid cohort limit generalization. The data does not meter judge tokens, dollar cost, or duration; no measured savings are claimed. Independent cross-family review reproduced these counts and limitations. These results are retained for discussion, not as justification for a judge change in this PR.

Reproduce counts from the pinned JSON with:

const rows = data.entries.filter(r =>
  r.primaryJudge === 'claude-opus-4.8' &&
  r.secondJudge === 'claude-haiku-4.5');
const valid = r =>
  ['VALID_PASS', 'VALID_NO_CHANGE'].includes(r.primaryState) &&
  ['VALID_PASS', 'VALID_NO_CHANGE'].includes(r.secondState) &&
  r.primaryConclusive === true && r.secondConclusive === true &&
  r.primaryUnderpowered === false && r.secondUnderpowered === false;
const summarize = rs => ({
  pairs: rs.length,
  passAgree: rs.filter(r => r.primaryPassed === r.secondPassed).length,
  stateAgree: rs.filter(r => r.primaryState === r.secondState &&
    r.primaryPreferenceRegressed === r.secondPreferenceRegressed).length,
  opusPassHaikuNot: rs.filter(r => r.primaryPassed && !r.secondPassed).length,
  haikuPassOpusNot: rs.filter(r => !r.primaryPassed && r.secondPassed).length
});
console.log(summarize(rows));
console.log(summarize(rows.filter(valid)));

Expected all-pair output: 1101, 956, 945, 110, 35. Expected valid-cohort output: 595, 452, 443, 109, 34, in the field order above.

Operational evidence and limits

The seven model-failure reports #1120, #1125, #1126, #1129, #1130, #1132, and #1135 selected Sonnet 4.6 and returned HTTP 400: unavailable for integrator agentic-workflows. Run 34091012534 lists Sol and Sonnet 5, but not Sonnet 4.6. This proves advertised integration availability, not a successful current invocation.

Historical health-check, groom, and triage agents completed with auto-selected Sonnet 5. Those do not prove current Vally evaluation access.

This does not resolve HTTP 401 reports #1147, #1141, #1131, and #1128. No issues are automatically closed.

Validation and review

The Sol/Sonnet executor revision passed strict compilation of all four agentic workflows, 22 workflow tests, actionlint on evaluation.yml, and independent Claude/GPT/Gemini review. The generated locks retain 22 pre-existing actionlint diagnostics; baseline and changed diagnostics match. This is not a clean full-lock lint result. Judge restoration in commit 2c6c0ff76 passed both targeted tests (2/2) and actionlint on evaluation.yml. At that commit, the restored evaluation configuration and both investigation guides matched pre-switch commit cf4f4ad08 exactly, except for the corrected Sonnet 5 schedule comment. The later SDK repair adds the separate configuration and documentation described above.

SDK repair: six startup regressions passed on each supported SDK version; 22 workflow tests passed; the launcher returned Vally version 0.14.0; actionlint passed for both changed workflow files. Independent Claude/GPT/Gemini review found no significant issues in the core repair. The first remote smoke test rejected SDK 1.0.11; the current correction adds that explicitly tested version rather than removing the version guard. Current-head remote checks and evaluation are not yet claimed green.

Live workflow evaluations have been dispatched from this PR branch. No merge or credential changes were made.

Replace the concrete Sonnet 4.6 fallback in the four scoped workflows while preserving repository model overrides and PAT environment boundaries. Regenerate locks with gh-aw v0.86.2 without changing action or container pins.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings September 10, 2026 21:40

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

All reviewed changes are consistent and no blocking issues were identified.

Review tier: Lite
Findings: None

What changed in this PR

Updates four health and triage workflows to use claude-sonnet-5 and regenerates their lock files.

Changes:

  • Updated model defaults while preserving overrides and workflow configuration.
  • Regenerated compiled lock files with gh-aw v0.86.2.
File Description
.github/​workflows/​issue-triage.md Updated model default
.github/​workflows/​issue-triage.lock.yml Regenerated compiled workflow
.github/​workflows/​devops-health-investigate.md Updated model default
.github/​workflows/​devops-health-investigate.lock.yml Regenerated compiled workflow
.github/​workflows/​devops-health-groom.md Updated model default
.github/​workflows/​devops-health-groom.lock.yml Regenerated compiled workflow
.github/​workflows/​devops-health-check.md Updated model default
.github/​workflows/​devops-health-check.lock.yml Regenerated compiled workflow

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Use gpt-5.6-sol for four health and triage workflows. Use claude-sonnet-5 alongside unchanged gpt-5.6-luna in default and full evaluation profiles, retain model overrides and judge routing, and cover actual profile selection in workflow tests.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI review requested due to automatic review settings September 10, 2026 21:57
@AbhitejJohn AbhitejJohn changed the title fix(ci): use Sonnet 5 for health and triage workflows fix(ci): use Sol for triage and Sonnet 5 for PR evaluations Sep 10, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

evaluation.yml has an unresolved comparison-label issue and a stale schedule-documentation nit.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Review tier: Lite
Findings: 1 Low severity

New issues introduced by this change (1)
Severity Finding
Low severity .github/​workflows/​evaluation.yml — Update the remaining schedule comment to Sonnet 5

Comment thread .github/workflows/evaluation.yml
@github-actions github-actions Bot added the waiting-on-author PR state label label Sep 10, 2026
@github-actions

Copy link
Copy Markdown
Contributor

👋 @AbhitejJohn — this PR has 1 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)

Replace Opus 4.8 judging with Haiku 4.5 while retaining Opus executor profiles and cross-family Terra judging for Claude and MAI. Disable duplicate secondary judging, update routing regression tests and guidance, and correct the stale Sonnet schedule comment.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI review requested due to automatic review settings September 10, 2026 22:57

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Scheduled judge routing does not match the stated unchanged-behavior requirement.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Review tier: Lite
Findings: 1 Medium severity

New issues introduced by this change (1)
Severity Finding
Medium severity .github/​workflows/​evaluation.yml — Preserve the scheduled judge contract
Issues resolved since last review (1)
Severity Finding
Low severity .github/​workflows/​evaluation.yml — Update the remaining schedule comment to Sonnet 5 View resolved comment

Comment thread .github/workflows/evaluation.yml Outdated
@AbhitejJohn AbhitejJohn changed the title fix(ci): use Sol for triage and Sonnet 5 for PR evaluations fix(ci): use Sol for triage, Sonnet 5 for PRs, and Haiku judges Sep 10, 2026
@AbhitejJohn AbhitejJohn changed the title fix(ci): use Sol for triage, Sonnet 5 for PRs, and Haiku judges fix(ci): use Sol for triage and Sonnet 5 for PR evaluations Sep 10, 2026
Restore Opus 4.8 primary judges and scheduled Haiku 4.5 secondary judges at the user's request. Keep the corrected Sonnet schedule comment, Sol triage defaults, Sonnet 5 evaluation executors, and expanded schedule regression coverage.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI review requested due to automatic review settings September 10, 2026 23:06

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Corrective judge-restoration validation is still pending.

Review tier: Lite
Findings: None

Issues resolved since last review (2)
Severity Finding
Medium severity .github/​workflows/​evaluation.yml — Preserve the scheduled judge contract View resolved comment
Low severity .github/​workflows/​evaluation.yml — Update the remaining schedule comment to Sonnet 5 View resolved comment

@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress and removed waiting-on-author PR state label labels Sep 10, 2026
github-actions Bot added a commit that referenced this pull request Sep 10, 2026
@github-actions

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate 2c6c0ff76c85cd034e791eea1a4f8e8636bd9827 to retry this exact commit.

2 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.

@github-actions github-actions Bot removed the pr-state/evals-in-progress PR evaluations are in progress label Sep 11, 2026
github-actions Bot added a commit that referenced this pull request Sep 11, 2026
@github-actions

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate 2c6c0ff76c85cd034e791eea1a4f8e8636bd9827 to retry this exact commit.

22 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.

Guard the pinned SDK's concurrent lazy startup and wait for filesystem-provider readiness before create/resume. Use the trusted launcher for evaluation and comparison without changing model routing, trial concurrency, or result gates.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI review requested due to automatic review settings September 11, 2026 22:01
Validate the same startup guard against SDK 1.0.11 used by the PR merge checks and SDK 1.0.13 on this branch. Keep unknown SDK versions fail-closed. Mock the runtime path explicitly in transport-only tests.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

No unresolved blocking issues were identified in the final review.

Review tier: Lite
Findings: None

Copilot AI review requested due to automatic review settings September 11, 2026 22:06
@AbhitejJohn
AbhitejJohn deployed to copilot-pat-pool September 11, 2026 22:09 — with GitHub Actions Active
@AbhitejJohn
AbhitejJohn deployed to copilot-pat-pool September 11, 2026 22:09 — with GitHub Actions Active
@AbhitejJohn
AbhitejJohn deployed to copilot-pat-pool September 11, 2026 22:09 — with GitHub Actions Active
@AbhitejJohn
AbhitejJohn deployed to copilot-pat-pool September 11, 2026 22:09 — with GitHub Actions Active

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Current-head remote checks and complete repaired evaluations are still pending.

Review tier: Lite
Findings: None

@github-actions github-actions Bot added pr-state/evals-in-progress PR evaluations are in progress waiting-on-review PR state label and removed pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress labels Sep 11, 2026
@github-actions

Copy link
Copy Markdown
Contributor

✅ Evaluation passed for b8083dd. cc @AbhitejJohn @JanKrivanek — please review.

github-actions Bot added a commit that referenced this pull request Sep 11, 2026
@github-actions

Copy link
Copy Markdown
Contributor

📊 Skill Evaluation Results

16 model/skill results across 8 skills and 2 models — ✅ 3 improved, ➖ 6 not proven improved, ⚠️ 6 invalid or underpowered, ⛔ 0 activation contract failures, 📉 1 preference losses (report only).

Measurement identity: evaluated commit b8083dd46b40bf81675c3d1dac06a33d8e67d0c6; 2 judge models.

Measurement health: 16 expected / 16 observed / 16 written; 0 missing, 0 unexpected, 6 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Skill Model Verdict Gate evidence Overfit Warnings Next action
dotnet-aot-compat claude-sonnet-5 ⚠️ Underpowered n=1; 0W/1T/0L; d=0; p=1.000; net +0.0% 🔴 0.58 — Predeclare more independent, discriminating stimuli; repeated runs do not add power.
dotnet-aot-compat gpt-5.6-luna ⚠️ Underpowered n=1; 1W/0T/0L; d=1; p=0.500; net +100.0% 🟡 0.28 — Predeclare more independent, discriminating stimuli; repeated runs do not add power.
exp-mock-usage-analysis claude-sonnet-5 ✅ Improved n=6; 5W/1T/0L; d=5; p=0.031; net +83.3% 🟡 0.23 — Review overfit evidence.
exp-mock-usage-analysis gpt-5.6-luna ✅ Improved n=6; 5W/1T/0L; d=5; p=0.031; net +83.3% ✅ 0.13 — None.
exp-test-maintainability claude-sonnet-5 ⚠️ Underpowered n=4; 3W/1T/0L; d=3; p=0.125; net +75.0% 🟡 0.39 — Predeclare more independent, discriminating stimuli; repeated runs do not add power.
exp-test-maintainability gpt-5.6-luna ⚠️ Underpowered n=4; 3W/0T/1L; d=4; p=0.312; net +50.0% ✅ 0.10 — Predeclare more independent, discriminating stimuli; repeated runs do not add power.
migrate-dotnet10-to-dotnet11 claude-sonnet-5 ➖ Not proven improved n=11; 6W/0T/5L; d=11; p=0.500; net +9.1% ✅ 0.09 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
migrate-dotnet10-to-dotnet11 gpt-5.6-luna ➖ Not proven improved n=11; 3W/3T/5L; d=8; p=0.363; net -18.2% ✅ 0.07 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
migrate-dotnet8-to-dotnet9 claude-sonnet-5 ➖ Not proven improved n=12; 6W/2T/4L; d=10; p=0.377; net +16.7% ✅ 0.07 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
migrate-dotnet8-to-dotnet9 gpt-5.6-luna ➖ Not proven improved n=12; 7W/3T/2L; d=9; p=0.090; net +41.7% ✅ 0.07 Activation: isolated 12/12; plugin 11/12 Inspect tied or lost stimuli and fix inconsistent skill behavior.
migrate-dotnet9-to-dotnet10 claude-sonnet-5 ✅ Improved n=17; 10W/4T/3L; d=13; p=0.046; net +41.2% ✅ 0.18 — None.
migrate-dotnet9-to-dotnet10 gpt-5.6-luna 📉 Preference loss (report only) n=17; 1W/1T/15L; d=16; p=0.000; net -82.4% ✅ 0.07 — Inspect losing stimuli and fix skill behavior; this is not objective completion proof.
migrate-nullable-references claude-sonnet-5 ⚠️ Underpowered n=3; 0W/2T/1L; d=1; p=0.500; net -33.3% 🟡 0.31 Activation: isolated 3/3; plugin 2/3 Predeclare more independent, discriminating stimuli; repeated runs do not add power.
migrate-nullable-references gpt-5.6-luna ⚠️ Underpowered n=3; 0W/2T/1L; d=1; p=0.500; net -33.3% ✅ 0.07 — Predeclare more independent, discriminating stimuli; repeated runs do not add power.
thread-abort-migration claude-sonnet-5 ➖ Not proven improved n=5; 4W/0T/1L; d=5; p=0.188; net +60.0% 🟡 0.33 Activation: isolated 4/5; plugin 4/5 Inspect tied or lost stimuli and fix inconsistent skill behavior.
thread-abort-migration gpt-5.6-luna ➖ Not proven improved n=5; 4W/0T/1L; d=5; p=0.188; net +60.0% 🟡 0.33 Activation: isolated 4/5; plugin 5/5 Inspect tied or lost stimuli and fix inconsistent skill behavior.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the skill.
  • ⛔ Activation contract failed — the isolated target skill activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/skill result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
⚠️ Underpowered — dotnet-aot-compat (claude-sonnet-5)

Why: Net win +0.0% (0W/1T/0L over 1 preference-eligible stimulus vote(s), sign test p=1.000), mean preference +0.0% across 1 paired run(s) — underpowered (1 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=1; 0W/1T/0L; d=0; p=1.000; net +0.0%

Overfit: High (score 0.58)

Repeated-run reliability (not used by the gate): 1 paired run (0W/1T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Make Azure.ResourceManager AOT-compatible Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Make Azure.ResourceManager AOT-compatible: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⚠️ Underpowered — dotnet-aot-compat (gpt-5.6-luna)

Why: Net win +100.0% (1W/0T/0L over 1 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +100.0% across 1 paired run(s) — underpowered (1 preference-eligible stimulus vote(s); a credible verdict needs at least 5, and this eval won every one of them) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=1; 1W/0T/0L; d=1; p=0.500; net +100.0%

Overfit: Moderate (score 0.28)

Repeated-run reliability (not used by the gate): 1 paired run (1W/0T/0L).

⚠️ Underpowered — exp-test-maintainability (claude-sonnet-5)

Why: Net win +75.0% (3W/1T/0L over 4 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +45.0% across 4 paired run(s) — underpowered (4 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=4; 3W/1T/0L; d=3; p=0.125; net +75.0%

Overfit: Moderate (score 0.39)

Repeated-run reliability (not used by the gate): 4 paired runs (3W/1T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Recognize tests with minimal boilerplate that need no refactoring Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Recognize tests with minimal boilerplate that need no refactoring: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⚠️ Underpowered — exp-test-maintainability (gpt-5.6-luna)

Why: Net win +50.0% (3W/0T/1L over 4 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +35.0% across 4 paired run(s) — underpowered (4 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=4; 3W/0T/1L; d=4; p=0.312; net +50.0%

Overfit: Low (score 0.10)

Repeated-run reliability (not used by the gate): 4 paired runs (3W/0T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Recognize tests with minimal boilerplate that need no refactoring Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Recognize tests with minimal boilerplate that need no refactoring: Both responses are accurate, avoid over-engineering, and correctly frame the code as already well-structured. B is more polished with priorities and trade-offs and better addresses the shared-field point. However, the core of the task—identifying the repeated individual test m...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⚠️ Underpowered — migrate-nullable-references (claude-sonnet-5)

Why: Net win -33.3% (0W/2T/1L over 3 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -33.3% across 3 paired run(s) — underpowered (3 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=3; 0W/2T/1L; d=1; p=0.500; net -33.3%

Warnings: Activation: isolated 3/3; plugin 2/3

Overfit: Moderate (score 0.31)

Repeated-run reliability (not used by the gate): 3 paired runs (0W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Enable NRT in ASP.NET Core Web API with EF Core Eligible -100.0% -100.0% 0/0/1
= Enable NRT in a small library with mixed nullability Eligible +0.0% +0.0% 0/1/0
= File-by-file migration: only modify the targeted file Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Enable NRT in ASP.NET Core Web API with EF Core: A's annotations consistently reflect the entity and API optionality requirements, especially the optional summary/category response fields. B compiles, and gets EF navigations and DbSets right, but its blanket required DTO strings misrepresents optional Summary and CategoryNam...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⚠️ Underpowered — migrate-nullable-references (gpt-5.6-luna)

Why: Net win -33.3% (0W/2T/1L over 3 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -13.3% across 3 paired run(s) — underpowered (3 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=3; 0W/2T/1L; d=1; p=0.500; net -33.3%

Overfit: Low (score 0.07)

Repeated-run reliability (not used by the gate): 3 paired runs (0W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Enable NRT in ASP.NET Core Web API with EF Core Eligible -100.0% -40.0% 0/0/1
= Enable NRT in a small library with mixed nullability Eligible +0.0% +0.0% 0/1/0
= File-by-file migration: only modify the targeted file Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Enable NRT in ASP.NET Core Web API with EF Core: Both responses enabled nullable, kept DbSets non-nullable, correctly marked Book.Category as nullable, and achieved zero-warning builds. The differentiator is semantic annotation quality on optional/descriptive fields. Response B's visible timeline reveals a blanket regex appl...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

📉 Preference loss (report only) — migrate-dotnet9-to-dotnet10 (gpt-5.6-luna)

Why: Net win -82.4% (1W/1T/15L over 17 preference-eligible stimulus vote(s), sign test p=0.000), mean preference -36.5% across 17 paired run(s) — credibly worse

Next action: Inspect losing stimuli and fix skill behavior; this is not objective completion proof.

State: VALID_NO_CHANGE (preference_regression_report_only)

Gate evidence: n=17; 1W/1T/15L; d=16; p=0.000; net -82.4%

Overfit: Low (score 0.07)

Repeated-run reliability (not used by the gate): 17 paired runs (1W/1T/15L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ ASP.NET Core app with OpenAPI transformers using Microsoft.OpenApi v1 APIs Eligible -100.0% -40.0% 0/0/1
▼ ASP.NET Core app with WebHostBuilder, OpenAPI, and forwarded headers Eligible -100.0% -40.0% 0/0/1
▼ App using SslStream properties and SystemEvents Eligible -100.0% -40.0% 0/0/1
▼ Blazor WASM app with generic math shift masking and tar operations Eligible -100.0% -100.0% 0/0/1
▼ C# 14 compiler breaking changes — field keyword, extension keyword, disposal Eligible -100.0% -40.0% 0/0/1
▼ Console app with System.Linq.Async, SIGTERM, and BufferedStream Eligible -100.0% -40.0% 0/0/1
▼ Containerized single-file app with P/Invoke and IDispatchEx Eligible -100.0% -40.0% 0/0/1
▼ Cryptography app with OpenSSL, X.509, and Rfc2898DeriveBytes Eligible -100.0% -40.0% 0/0/1
▼ EF Core app with Azure SQL JSON columns and parameterized collections Eligible -100.0% -40.0% 0/0/1
▼ EF Core app with dynamic ExecuteUpdate and complex types Eligible -100.0% -40.0% 0/0/1
= Expression tree code broken by C# 14 span overload resolution Eligible +0.0% +0.0% 0/1/0
▼ JSON polymorphism with conflicting property names and XmlSerializer Eligible -100.0% -40.0% 0/0/1
▼ Library with NuGet auditing, transitive deps, and InlineArray Eligible -100.0% -40.0% 0/0/1
▼ SDK and NuGet obscure tooling changes Eligible -100.0% -40.0% 0/0/1
▼ SQLite app with DateTimeOffset timezone handling Eligible -100.0% -40.0% 0/0/1
▼ WinForms and WPF desktop app with System.Drawing and DynamicResource Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • ASP.NET Core app with OpenAPI transformers using Microsoft.OpenApi v1 APIs: Both responses are high quality, correct, and cover all five rubric points with essentially the same solution. A is marginally more accurate: it grounds its answer in fetched source files, uses the more complete OpenApiSecuritySchemeReference constructor with the host document...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — migrate-dotnet10-to-dotnet11 (claude-sonnet-5)

Why: Net win +9.1% (6W/0T/5L over 11 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +14.5% across 11 paired run(s) — not credible (sign test p=0.500 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=11; 6W/0T/5L; d=11; p=0.500; net +9.1%

Overfit: Low (score 0.09)

Repeated-run reliability (not used by the gate): 11 paired runs (6W/0T/5L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ BackgroundService exceptions and ZipArchive CRC32 validation Eligible -100.0% -40.0% 0/0/1
▼ Basic TFM update with Docker and global.json Eligible -100.0% -40.0% 0/0/1
▼ Cryptography app using DSA on macOS Eligible -100.0% -40.0% 0/0/1
▼ EF Core app with Cosmos DB provider using sync APIs Eligible -100.0% -40.0% 0/0/1
▼ mTLS and AIA certificate chain validation Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • BackgroundService exceptions and ZipArchive CRC32 validation: Both answers address the two requested breaking changes well, but A is more technically nuanced and actionable on BackgroundService host-task faulting and is marginally more complete on ZIP validation. B is concise but oversimplifies the host behavior.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — migrate-dotnet10-to-dotnet11 (gpt-5.6-luna)

Why: Net win -18.2% (3W/3T/5L over 11 preference-eligible stimulus vote(s), sign test p=0.363), mean preference -7.3% across 11 paired run(s) — no improvement

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=11; 3W/3T/5L; d=8; p=0.363; net -18.2%

Overfit: Low (score 0.07)

Repeated-run reliability (not used by the gate): 11 paired runs (3W/3T/5L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ ASP.NET Core app with OpenAPI customizations and Blazor Virtualize Eligible -100.0% -40.0% 0/0/1
▼ BackgroundService exceptions and ZipArchive CRC32 validation Eligible -100.0% -40.0% 0/0/1
= C# 15 compiler breaking changes — Span safe-context, nameof, with() Eligible +0.0% +0.0% 0/1/0
▼ C# 15 dynamic operator and ref readonly delegate issues Eligible -100.0% -40.0% 0/0/1
= Cryptography app using DSA on macOS Eligible +0.0% +0.0% 0/1/0
= EF Core SQL Server with Entra ID auth and Design package dependency Eligible +0.0% +0.0% 0/1/0
▼ EF Core app with Cosmos DB provider using sync APIs Eligible -100.0% -40.0% 0/0/1
▼ mTLS and AIA certificate chain validation Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • ASP.NET Core app with OpenAPI customizations and Blazor Virtualize: Both responses correctly and completely address both parts of the question: the Microsoft.OpenApi v2→v3 breaking changes and the Virtualize OverscanCount default change from 3 to 15, each with appropriate remediation. Response A is slightly stronger because it provides a concr...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — migrate-dotnet8-to-dotnet9 (claude-sonnet-5)

Why: Net win +16.7% (6W/2T/4L over 12 preference-eligible stimulus vote(s), sign test p=0.377), mean preference +16.7% across 12 paired run(s) — not credible (sign test p=0.377 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=12; 6W/2T/4L; d=10; p=0.377; net +16.7%

Overfit: Low (score 0.07)

Repeated-run reliability (not used by the gate): 12 paired runs (6W/2T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ App with JsonDocument null deserialization and BinaryFormatter fallback Eligible -100.0% -40.0% 0/0/1
▼ App with empty environment variables, ZIP encoding, and keyed DI services Eligible -100.0% -40.0% 0/0/1
= C# 13 compiler breaking changes — InlineArray on record, iterator safe context, collection expressions Eligible +0.0% +0.0% 0/1/0
▼ EF Core Cosmos DB app with existing documents and composite id format Eligible -100.0% -40.0% 0/0/1
= EF Core app with migration patterns and Cosmos DB discriminator Eligible +0.0% +0.0% 0/1/0
▼ Library with String.Trim span overload, keyed services, and InlineArray Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • App with JsonDocument null deserialization and BinaryFormatter fallback: A more nearly fulfills the requested migration by targeting net9.0 and resolving all three identified .NET 9 breaking behaviors, including replacing BinaryFormatter. B correctly diagnoses the issue and responsibly notes serialization-format tradeoffs, but does not finish the r...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — migrate-dotnet8-to-dotnet9 (gpt-5.6-luna)

Why: Net win +41.7% (7W/3T/2L over 12 preference-eligible stimulus vote(s), sign test p=0.090), mean preference +31.7% across 12 paired run(s) — not credible (sign test p=0.090 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=12; 7W/3T/2L; d=9; p=0.090; net +41.7%

Warnings: Activation: isolated 12/12; plugin 11/12

Overfit: Low (score 0.07)

Repeated-run reliability (not used by the gate): 12 paired runs (7W/3T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ C# 13 compiler breaking changes — InlineArray on record, iterator safe context, collection expressions Eligible -100.0% -40.0% 0/0/1
= Containerized app with env var precedence reversal and zlib removal Eligible +0.0% +0.0% 0/1/0
▼ Containerized app with zlib dependency and runtime configuration Eligible -100.0% -40.0% 0/0/1
= EF Core Cosmos DB app with existing documents and composite id format Eligible +0.0% +0.0% 0/1/0
= EF Core app with migration patterns and Cosmos DB discriminator Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • C# 13 compiler breaking changes — InlineArray on record, iterator safe context, collection expressions: Both responses are technically accurate and cover all three breaking changes correctly with the right fixes. A did genuine research (fetched Microsoft docs and the language proposal) and provides slightly richer, still-accurate detail (AllowUnsafeBlocks, cross-yield pointer ca...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — thread-abort-migration (claude-sonnet-5)

Why: Net win +60.0% (4W/0T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +24.0% across 5 paired run(s) — not credible (sign test p=0.188 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%

Warnings: Activation: isolated 4/5; plugin 4/5

Overfit: Moderate (score 0.33)

Repeated-run reliability (not used by the gate): 5 paired runs (4W/0T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Thread.Join and Thread.Sleep only — should not migrate Eligible +100.0% +40.0% 1/0/0
▼ Worker thread with abort-based cancellation Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Worker thread with abort-based cancellation: A supplies the more complete and robust primary migration: it uses an appropriate dedicated background thread for potentially blocking work, does bounded cooperative Join without unsafe disposal after a timeout, includes generic exception handling, and offers concrete cancella...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — thread-abort-migration (gpt-5.6-luna)

Why: Net win +60.0% (4W/0T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +24.0% across 5 paired run(s) — not credible (sign test p=0.188 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%

Warnings: Activation: isolated 4/5; plugin 5/5

Overfit: Moderate (score 0.33)

Repeated-run reliability (not used by the gate): 5 paired runs (4W/0T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Blocking WaitHandle with Thread.Interrupt Eligible -100.0% -40.0% 0/0/1
▲ Thread.Join and Thread.Sleep only — should not migrate Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Blocking WaitHandle with Thread.Interrupt: Both responses are high quality, correct, and take the same fundamentally sound Channel<T> + cooperative cancellation approach. B edges out on two specific rubric points (explicit PlatformNotSupportedException note and the WaitSleepJoin/wait-state warning). However, A is more ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — exp-mock-usage-analysis (claude-sonnet-5)

Why: Net win +83.3% (5W/1T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +63.3% across 6 paired run(s) — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=6; 5W/1T/0L; d=5; p=0.031; net +83.3%

Overfit: Moderate (score 0.23)

Repeated-run reliability (not used by the gate): 6 paired runs (5W/1T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Analyze mock usage in FakeItEasy tests Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Analyze mock usage in FakeItEasy tests: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

Routine passing details for 2 results are in Full Results.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1154 in dotnet/skills, download eval artifacts with gh run download 34652384279 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/b8083dd46b40bf81675c3d1dac06a33d8e67d0c6/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

@github-actions

Copy link
Copy Markdown
Contributor

📊 Skill Evaluation Results

26 model/skill results across 13 skills and 2 models — ✅ 15 improved, ➖ 7 not proven improved, ⚠️ 4 invalid or underpowered, ⛔ 0 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit a226ead6ff24b56338979737651609f9aef8e736; 2 judge models.

Measurement health: 26 expected / 26 observed / 26 written; 0 missing, 0 unexpected, 4 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Skill Model Verdict Gate evidence Overfit Warnings Next action
analyzing-dotnet-performance claude-sonnet-5 ➖ Not proven improved n=11; 7W/2T/2L; d=9; p=0.090; net +45.5% 🔴 0.50 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
analyzing-dotnet-performance gpt-5.6-luna ✅ Improved n=11; 8W/3T/0L; d=8; p=0.004; net +72.7% ✅ 0.12 — None.
android-tombstone-symbolication claude-sonnet-5 ➖ Not proven improved n=8; 1W/1T/6L; d=7; p=0.063; net -62.5% 🟡 0.48 Activation: isolated 7/8; plugin 8/8 Inspect tied or lost stimuli and fix inconsistent skill behavior.
android-tombstone-symbolication gpt-5.6-luna ➖ Not proven improved n=8; 4W/3T/1L; d=5; p=0.188; net +37.5% 🟡 0.21 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
apple-crash-symbolication claude-sonnet-5 ⚠️ Underpowered n=3; 0W/1T/2L; d=2; p=0.250; net -66.7% 🟡 0.46 Activation: isolated 2/3; plugin 3/3 Predeclare more independent, discriminating stimuli; repeated runs do not add power.
apple-crash-symbolication gpt-5.6-luna ⚠️ Underpowered n=3; 2W/1T/0L; d=2; p=0.250; net +66.7% ✅ 0.16 Activation: isolated 2/3; plugin 3/3 Predeclare more independent, discriminating stimuli; repeated runs do not add power.
clr-activation-debugging claude-sonnet-5 ➖ Not proven improved n=7; 3W/0T/4L; d=7; p=0.500; net -14.3% 🟡 0.31 Activation: isolated 7/7; plugin 6/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
clr-activation-debugging gpt-5.6-luna ➖ Not proven improved n=7; 5W/1T/1L; d=6; p=0.109; net +57.1% ✅ 0.18 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
dotnet-trace-collect claude-sonnet-5 ✅ Improved n=17; 14W/3T/0L; d=14; p=0.000; net +82.4% 🔴 0.53 — Review overfit evidence.
dotnet-trace-collect gpt-5.6-luna ✅ Improved n=17; 14W/1T/2L; d=16; p=0.002; net +70.6% 🟡 0.21 — Review overfit evidence.
dump-collect claude-sonnet-5 ✅ Improved n=9; 6W/3T/0L; d=6; p=0.016; net +66.7% 🔴 0.59 Activation: isolated 8/9; plugin 8/9 Fix activation gaps; Review overfit evidence.
dump-collect gpt-5.6-luna ➖ Not proven improved n=9; 6W/1T/2L; d=8; p=0.145; net +44.4% ✅ 0.16 Activation: isolated 9/9; plugin 7/9 Inspect tied or lost stimuli and fix inconsistent skill behavior.
microbenchmarking claude-sonnet-5 ⚠️ Underpowered n=1; 1W/0T/0L; d=1; p=0.500; net +100.0% 🟡 0.41 — Predeclare more independent, discriminating stimuli; repeated runs do not add power.
microbenchmarking gpt-5.6-luna ⚠️ Underpowered n=1; 1W/0T/0L; d=1; p=0.500; net +100.0% ✅ 0.17 — Predeclare more independent, discriminating stimuli; repeated runs do not add power.
migrate-mstest-v1v2-to-v3 claude-sonnet-5 ✅ Improved n=11; 8W/2T/1L; d=9; p=0.020; net +63.6% 🟡 0.38 — Review overfit evidence.
migrate-mstest-v1v2-to-v3 gpt-5.6-luna ✅ Improved n=11; 11W/0T/0L; d=11; p=0.000; net +100.0% ✅ 0.11 — None.
migrate-mstest-v3-to-v4 claude-sonnet-5 ✅ Improved n=15; 8W/6T/1L; d=9; p=0.020; net +46.7% 🟡 0.35 — Review overfit evidence.
migrate-mstest-v3-to-v4 gpt-5.6-luna ✅ Improved n=15; 9W/4T/2L; d=11; p=0.033; net +46.7% 🟡 0.22 — Review overfit evidence.
migrate-nunit-to-mstest claude-sonnet-5 ➖ Not proven improved n=14; 4W/10T/0L; d=4; p=0.063; net +28.6% 🟡 0.29 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
migrate-nunit-to-mstest gpt-5.6-luna ✅ Improved n=14; 8W/5T/1L; d=9; p=0.020; net +50.0% ✅ 0.16 — None.
migrate-vstest-to-mtp claude-sonnet-5 ✅ Improved n=11; 9W/1T/1L; d=10; p=0.011; net +72.7% 🟡 0.33 — Review overfit evidence.
migrate-vstest-to-mtp gpt-5.6-luna ✅ Improved n=11; 9W/2T/0L; d=9; p=0.002; net +81.8% ✅ 0.14 — None.
migrate-xunit-to-mstest claude-sonnet-5 ✅ Improved n=13; 7W/5T/1L; d=8; p=0.035; net +46.2% 🟡 0.43 — Review overfit evidence.
migrate-xunit-to-mstest gpt-5.6-luna ✅ Improved n=13; 9W/4T/0L; d=9; p=0.002; net +69.2% ✅ 0.15 — None.
migrate-xunit-to-xunit-v3 claude-sonnet-5 ✅ Improved n=12; 6W/6T/0L; d=6; p=0.016; net +50.0% 🟡 0.36 — Review overfit evidence.
migrate-xunit-to-xunit-v3 gpt-5.6-luna ✅ Improved n=12; 9W/2T/1L; d=10; p=0.011; net +66.7% ✅ 0.11 — None.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the skill.
  • ⛔ Activation contract failed — the isolated target skill activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/skill result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
⚠️ Underpowered — apple-crash-symbolication (claude-sonnet-5)

Why: Net win -66.7% (0W/1T/2L over 3 preference-eligible stimulus vote(s), sign test p=0.250), mean preference -46.7% across 3 paired run(s) — underpowered (3 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=3; 0W/1T/2L; d=2; p=0.250; net -66.7%

Warnings: Activation: isolated 2/3; plugin 3/3

Overfit: Moderate (score 0.46)

Repeated-run reliability (not used by the gate): 3 paired runs (0W/1T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Investigate root cause of a .NET MAUI iOS crash Eligible -100.0% -40.0% 0/0/1
= Parse .NET frames and locate dSYMs from an iOS crash log Eligible +0.0% +0.0% 0/1/0
▼ Reject Android tombstone passed as iOS crash log Eligible -100.0% -100.0% 0/0/1

Illustrative judge evidence:

  • Investigate root cause of a .NET MAUI iOS crash: A is the more evidence-based and technically careful result: it validates and interprets the CoreCLR symbols, distinguishes the long-lived runtime/main-loop frames from the unknown application failure, and gives actionable app-dSYM symbolication commands. B is readable and off...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⚠️ Underpowered — apple-crash-symbolication (gpt-5.6-luna)

Why: Net win +66.7% (2W/1T/0L over 3 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +46.7% across 3 paired run(s) — underpowered (3 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=3; 2W/1T/0L; d=2; p=0.250; net +66.7%

Warnings: Activation: isolated 2/3; plugin 3/3

Overfit: Low (score 0.16)

Repeated-run reliability (not used by the gate): 3 paired runs (2W/1T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Reject Android tombstone passed as iOS crash log Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Reject Android tombstone passed as iOS crash log: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⚠️ Underpowered — microbenchmarking (claude-sonnet-5)

Why: Net win +100.0% (1W/0T/0L over 1 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +40.0% across 1 paired run(s) — underpowered (1 preference-eligible stimulus vote(s); a credible verdict needs at least 5, and this eval won every one of them) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=1; 1W/0T/0L; d=1; p=0.500; net +100.0%

Overfit: Moderate (score 0.41)

Repeated-run reliability (not used by the gate): 1 paired run (1W/0T/0L).

⚠️ Underpowered — microbenchmarking (gpt-5.6-luna)

Why: Net win +100.0% (1W/0T/0L over 1 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +100.0% across 1 paired run(s) — underpowered (1 preference-eligible stimulus vote(s); a credible verdict needs at least 5, and this eval won every one of them) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=1; 1W/0T/0L; d=1; p=0.500; net +100.0%

Overfit: Low (score 0.17)

Repeated-run reliability (not used by the gate): 1 paired run (1W/0T/0L).

➖ Not proven improved — analyzing-dotnet-performance (claude-sonnet-5)

Why: Net win +45.5% (7W/2T/2L over 11 preference-eligible stimulus vote(s), sign test p=0.090), mean preference +29.1% across 11 paired run(s) — not credible (sign test p=0.090 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=11; 7W/2T/2L; d=9; p=0.090; net +45.5%

Overfit: High (score 0.50)

Repeated-run reliability (not used by the gate): 11 paired runs (7W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Detects CurrentCulture comparer and compiled regex budget in inflection rules Eligible +0.0% +0.0% 0/1/0
▼ Finds per-call Dictionary allocation not hoisted to static Eligible -100.0% -40.0% 0/0/1
▼ Finds repeated enumeration in time and collection formatting Eligible -100.0% -40.0% 0/0/1
= Flags Span inconsistencies and compound method chains in truncation library Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Detects CurrentCulture comparer and compiled regex budget in inflection rules: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — android-tombstone-symbolication (claude-sonnet-5)

Why: Net win -62.5% (1W/1T/6L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference -25.0% across 8 paired run(s) — no improvement

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 1W/1T/6L; d=7; p=0.063; net -62.5%

Warnings: Activation: isolated 7/8; plugin 8/8

Overfit: Moderate (score 0.48)

Repeated-run reliability (not used by the gate): 8 paired runs (1W/1T/6L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Handle .NET frames with no BuildId metadata Eligible -100.0% -40.0% 0/0/1
▼ Recognize NativeAOT tombstone with app binary and libSystem.Native.so Eligible -100.0% -40.0% 0/0/1
▼ Recognize tombstone with no .NET frames Eligible -100.0% -40.0% 0/0/1
= Reject iOS crash log as wrong format Eligible +0.0% +0.0% 0/1/0
▼ Symbolicate CoreCLR frames in an Android tombstone Eligible -100.0% -40.0% 0/0/1
▼ Symbolicate multi-thread tombstone Eligible -100.0% -40.0% 0/0/1
▼ Symbolicate tombstone with multiple .NET libraries and different BuildIds Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Handle .NET frames with no BuildId metadata: Both responses correctly determine that symbolication cannot proceed from this truncated tombstone and give sound next steps. A is marginally stronger because it reports the parsed frame count and individual library/offset evidence, and its local-runtime-pack/build metadata gu...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — android-tombstone-symbolication (gpt-5.6-luna)

Why: Net win +37.5% (4W/3T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +30.0% across 8 paired run(s) — not credible (sign test p=0.188 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 4W/3T/1L; d=5; p=0.188; net +37.5%

Overfit: Moderate (score 0.21)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Handle .NET frames with no BuildId metadata Eligible +0.0% +0.0% 0/1/0
= Recognize tombstone with no .NET frames Eligible +0.0% +0.0% 0/1/0
= Reject iOS crash log as wrong format Eligible +0.0% +0.0% 0/1/0
▼ Symbolicate CoreCLR frames in an Android tombstone Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Handle .NET frames with no BuildId metadata: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — clr-activation-debugging (claude-sonnet-5)

Why: Net win -14.3% (3W/0T/4L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -22.9% across 7 paired run(s) — no improvement

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/0T/4L; d=7; p=0.500; net -14.3%

Warnings: Activation: isolated 7/7; plugin 6/7

Overfit: Moderate (score 0.31)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/0T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Decline non-CLR-activation issue Eligible +100.0% +40.0% 1/0/0
▼ Diagnose FOD suppressed but activation still failing Eligible -100.0% -100.0% 0/0/1
▼ Diagnose unexpected FOD dialog from native build tool Eligible -100.0% -100.0% 0/0/1
▼ Explain why same binary behaves differently under different launch methods Eligible -100.0% -40.0% 0/0/1
▼ Identify multiple activation sequences in a single log Eligible -100.0% -100.0% 0/0/1

Illustrative judge evidence:

  • Diagnose FOD suppressed but activation still failing: A finds and interprets the available activation log, gives a coherent causal chain for both the activation failure and silent behavior, and proposes a concrete fix. B merely asks the user to provide logs that were available in the environment, so it does not answer the task at...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — clr-activation-debugging (gpt-5.6-luna)

Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +22.9% across 7 paired run(s) — not credible (sign test p=0.109 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%

Overfit: Low (score 0.18)

Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Analyze healthy managed EXE activation Eligible -100.0% -100.0% 0/0/1
= Identify multiple activation sequences in a single log Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Analyze healthy managed EXE activation: Response A located the log file in the filesystem (after its view tool failed, it recovered via bash/find) and delivered a correct, thorough assessment confirming normal activation, runtime v4.0.30319, and absence of errors. Response B activated a relevant skill but its file r...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — dump-collect (gpt-5.6-luna)

Why: Net win +44.4% (6W/1T/2L over 9 preference-eligible stimulus vote(s), sign test p=0.145), mean preference +24.4% across 9 paired run(s) — not credible (sign test p=0.145 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 6W/1T/2L; d=8; p=0.145; net +44.4%

Warnings: Activation: isolated 9/9; plugin 7/9

Overfit: Low (score 0.16)

Repeated-run reliability (not used by the gate): 9 paired runs (6W/1T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Configure CoreCLR dump collection in Alpine Docker as non-root Eligible +0.0% +0.0% 0/1/0
▼ Configure automatic crash dumps for CoreCLR app on Linux Eligible -100.0% -40.0% 0/0/1
▲ Decline dump analysis request Eligible +100.0% +100.0% 1/0/0
▼ Recover crash dump from macOS NativeAOT without createdump Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Configure CoreCLR dump collection in Alpine Docker as non-root: Both responses fully satisfy every rubric criterion: enabling DOTNET_DbgEnableMiniDump, adding SYS_PTRACE, handling non-root dump directory permissions, avoiding false Alpine claims, and not analyzing dumps. Response A adds a named-volume workflow with dump extraction steps an...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — migrate-nunit-to-mstest (claude-sonnet-5)

Why: Net win +28.6% (4W/10T/0L over 14 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +11.4% across 14 paired run(s) — not credible — 10 of 14 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=14; 4W/10T/0L; d=4; p=0.063; net +28.6%

Overfit: Moderate (score 0.29)

Repeated-run reliability (not used by the gate): 14 paired runs (4W/10T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Convert NUnit combinatorial parameter data Eligible +0.0% +0.0% 0/1/0
= Convert NUnit theory datapoint discovery Eligible +0.0% +0.0% 0/1/0
= Convert cooperative cancellation attribute Eligible +0.0% +0.0% 0/1/0
= Preserve NUnit pairwise case set Eligible +0.0% +0.0% 0/1/0
= Preserve NUnit single-instance fixture state Eligible +0.0% +0.0% 0/1/0
= Preserve STA across async continuations Eligible +0.0% +0.0% 0/1/0
= Preserve TestCaseSource result and row metadata Eligible +0.0% +0.0% 0/1/0
= Preserve all-inconclusive NUnit theory result Eligible +0.0% +0.0% 0/1/0
= Preserve assembly-wide SetUpFixture scope Eligible +0.0% +0.0% 0/1/0
= Translate explicit NUnit parallelization Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Convert NUnit combinatorial parameter data: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — dotnet-trace-collect (claude-sonnet-5)

Why: Net win +82.4% (14W/3T/0L over 17 preference-eligible stimulus vote(s), sign test p=0.000), mean preference +57.6% across 17 paired run(s) — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=17; 14W/3T/0L; d=14; p=0.000; net +82.4%

Overfit: High (score 0.53)

Repeated-run reliability (not used by the gate): 17 paired runs (14W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Excessive GC on Linux (.NET 8) Eligible +0.0% +0.0% 0/1/0
= Memory leak on .NET Framework Windows Eligible +0.0% +0.0% 0/1/0
= Memory leak on Linux (.NET 8) Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Excessive GC on Linux (.NET 8): Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — dotnet-trace-collect (gpt-5.6-luna)

Why: Net win +70.6% (14W/1T/2L over 17 preference-eligible stimulus vote(s), sign test p=0.002), mean preference +45.9% across 17 paired run(s) — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=17; 14W/1T/2L; d=16; p=0.002; net +70.6%

Overfit: Moderate (score 0.21)

Repeated-run reliability (not used by the gate): 17 paired runs (14W/1T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Container installation without .NET SDK Eligible -100.0% -100.0% 0/0/1
= Excessive GC on Linux (.NET 8) Eligible +0.0% +0.0% 0/1/0
▼ Linux pre-.NET 10 needing native call stacks Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Container installation without .NET SDK: Response A directly answers the intent of the question—getting dotnet-trace without the SDK—via the aka.ms self-contained single-file download and a curl command, satisfying the two most specific rubric criteria. Response B, while technically valid and thorough, uses the .NET ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — dump-collect (claude-sonnet-5)

Why: Net win +66.7% (6W/3T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +33.3% across 9 paired run(s) — credibly better

Next action: Fix activation gaps; Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=9; 6W/3T/0L; d=6; p=0.016; net +66.7%

Warnings: Activation: isolated 8/9; plugin 8/9

Overfit: High (score 0.59)

Repeated-run reliability (not used by the gate): 9 paired runs (6W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Advisory: NativeAOT Kubernetes dump collection setup Eligible +0.0% +0.0% 0/1/0
= Configure automatic crash dumps for CoreCLR app on Linux Eligible +0.0% +0.0% 0/1/0
▲ Decline dump analysis request Eligible +100.0% +40.0% 1/0/0
= Detect runtime and configure crash dumps for unknown .NET app on Linux Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Advisory: NativeAOT Kubernetes dump collection setup: Position-swap inconsistent (forward: tie, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — migrate-mstest-v1v2-to-v3 (claude-sonnet-5)

Why: Net win +63.6% (8W/2T/1L over 11 preference-eligible stimulus vote(s), sign test p=0.020), mean preference +30.9% across 11 paired run(s) — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=11; 8W/2T/1L; d=9; p=0.020; net +63.6%

Overfit: Moderate (score 0.38)

Repeated-run reliability (not used by the gate): 11 paired runs (8W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Correctly identify MSTest v1 vs v2 and recommend different migration paths Eligible +0.0% +0.0% 0/1/0
= Fix Assert.AreEqual object overload errors after v3 upgrade Eligible +0.0% +0.0% 0/1/0
▼ Migrate from .testsettings to .runsettings Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Correctly identify MSTest v1 vs v2 and recommend different migration paths: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — migrate-mstest-v3-to-v4 (claude-sonnet-5)

Why: Net win +46.7% (8W/6T/1L over 15 preference-eligible stimulus vote(s), sign test p=0.020), mean preference +26.7% across 15 paired run(s) — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=15; 8W/6T/1L; d=9; p=0.020; net +46.7%

Overfit: Moderate (score 0.35)

Repeated-run reliability (not used by the gate): 15 paired runs (8W/6T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Fix TestMethodAttribute CallerInfo constructor breaking change Eligible +0.0% +0.0% 0/1/0
= Fix TestMethodAttribute and TestMethod display name constructor Eligible +0.0% +0.0% 0/1/0
= Fix multiple v4 breaking changes: Assert, ClassCleanup, TestContext, Timeout Eligible +0.0% +0.0% 0/1/0
▼ Full MSTest v3 to v4 migration with multiple breaking changes Eligible -100.0% -40.0% 0/0/1
= Handle net6.0 target framework dropped in MSTest v4 Eligible +0.0% +0.0% 0/1/0
= Understand behavioral changes after MSTest v4 upgrade Eligible +0.0% +0.0% 0/1/0
= Verified MSTest v3 to v4 package update with build and test Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Fix TestMethodAttribute CallerInfo constructor breaking change: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — migrate-mstest-v3-to-v4 (gpt-5.6-luna)

Why: Net win +46.7% (9W/4T/2L over 15 preference-eligible stimulus vote(s), sign test p=0.033), mean preference +26.7% across 15 paired run(s) — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=15; 9W/4T/2L; d=11; p=0.033; net +46.7%

Overfit: Moderate (score 0.22)

Repeated-run reliability (not used by the gate): 15 paired runs (9W/4T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Fix Assert.IsInstanceOfType out parameter removal Eligible +0.0% +0.0% 0/1/0
= Fix multiple v4 breaking changes: Assert, ClassCleanup, TestContext, Timeout Eligible +0.0% +0.0% 0/1/0
= Full MSTest v3 to v4 migration with multiple breaking changes Eligible +0.0% +0.0% 0/1/0
▼ Handle net6.0 target framework dropped in MSTest v4 Eligible -100.0% -40.0% 0/0/1
▼ Migrate custom TestMethodAttribute from Execute to ExecuteAsync Eligible -100.0% -40.0% 0/0/1
= Understand behavioral changes after MSTest v4 upgrade Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Fix Assert.IsInstanceOfType out parameter removal: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — migrate-vstest-to-mtp (claude-sonnet-5)

Why: Net win +72.7% (9W/1T/1L over 11 preference-eligible stimulus vote(s), sign test p=0.011), mean preference +50.9% across 11 paired run(s) — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=11; 9W/1T/1L; d=10; p=0.011; net +72.7%

Overfit: Moderate (score 0.33)

Repeated-run reliability (not used by the gate): 11 paired runs (9W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Set OutputType=Exe only for test projects in Directory.Build.props Eligible -100.0% -40.0% 0/0/1
= Update Azure DevOps pipeline from VSTest task to MTP Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Set OutputType=Exe only for test projects in Directory.Build.props: The answers are substantively equivalent and correct. A is marginally more complete: it explicitly says every MTP-related property should remain in the conditioned group and offers a sensible per-project fallback if project naming is not reliable. B's example is nevertheless f...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — migrate-xunit-to-mstest (claude-sonnet-5)

Why: Net win +46.2% (7W/5T/1L over 13 preference-eligible stimulus vote(s), sign test p=0.035), mean preference +27.7% across 13 paired run(s) — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=13; 7W/5T/1L; d=8; p=0.035; net +46.2%

Overfit: Moderate (score 0.43)

Repeated-run reliability (not used by the gate): 13 paired runs (7W/5T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Convert ITestOutputHelper to TestContext Eligible +0.0% +0.0% 0/1/0
= Convert Skip, Trait, and Timeout Eligible +0.0% +0.0% 0/1/0
▼ Handle ICollectionFixture explicitly (do not silently widen scope) Eligible -100.0% -40.0% 0/0/1
= Map exception assertions correctly (xUnit Throws -> MSTest ThrowsExactly) Eligible +0.0% +0.0% 0/1/0
= Preserve type and sequence assertion semantics Eligible +0.0% +0.0% 0/1/0
= Preserve xUnit parallelization default with [assembly: Parallelize] Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Convert ITestOutputHelper to TestContext: Both runs complete the migration correctly: xUnit packages and APIs are removed, MSTest attributes/assertions/TestContext are introduced, the net8.0 target and VSTest configuration are retained, and each reports a successful one-test run. A uses newer verified package versions...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — migrate-xunit-to-xunit-v3 (claude-sonnet-5)

Why: Net win +50.0% (6W/6T/0L over 12 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +25.0% across 12 paired run(s) — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=12; 6W/6T/0L; d=6; p=0.016; net +50.0%

Overfit: Moderate (score 0.36)

Repeated-run reliability (not used by the gate): 12 paired runs (6W/6T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Consolidate xunit.extensibility packages and remove xunit.abstractions Eligible +0.0% +0.0% 0/1/0
= Convert async void test methods to async Task Eligible +0.0% +0.0% 0/1/0
= Migrate project with YTest.MTP.XUnit2 to xUnit.net v3 preserving MTP Eligible +0.0% +0.0% 0/1/0
= Recognize project already on xUnit.net v3 — no migration needed Eligible +0.0% +0.0% 0/1/0
= Update BeforeAfterTestAttribute overrides with IXunitTest parameter Eligible +0.0% +0.0% 0/1/0
= Update custom FactAttribute to include source information parameters Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Consolidate xunit.extensibility packages and remove xunit.abstractions: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

Routine passing details for 6 results are in Full Results.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1154 in dotnet/skills, download eval artifacts with gh run download 34652125571 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/a226ead6ff24b56338979737651609f9aef8e736/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

@AbhitejJohn
AbhitejJohn merged commit 4c72b17 into main Sep 12, 2026
52 of 53 checks passed
@AbhitejJohn
AbhitejJohn deleted the abhitejjohn-agentic-workflow-repair branch September 12, 2026 00:01

This branch was successfully deployed

1 active deployment
copilot-pat-pool — b8083dd4 Deployed Sep 11, 2026 by AbhitejJohn via evaluate / vally (dotnet-experimental--claude-sonnet-5) #9092
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

waiting-on-review PR state label

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants