fix(ci): use Sol for triage and Sonnet 5 for PR evaluations - #1154
Conversation
Replace the concrete Sonnet 4.6 fallback in the four scoped workflows while preserving repository model overrides and PAT environment boundaries. Regenerate locks with gh-aw v0.86.2 without changing action or container pins. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
All reviewed changes are consistent and no blocking issues were identified.
Review tier: Lite
Findings: None
What changed in this PR
Updates four health and triage workflows to use claude-sonnet-5 and regenerates their lock files.
Changes:
- Updated model defaults while preserving overrides and workflow configuration.
- Regenerated compiled lock files with gh-aw v0.86.2.
| File | Description |
|---|---|
.github/workflows/issue-triage.md |
Updated model default |
.github/workflows/issue-triage.lock.yml |
Regenerated compiled workflow |
.github/workflows/devops-health-investigate.md |
Updated model default |
.github/workflows/devops-health-investigate.lock.yml |
Regenerated compiled workflow |
.github/workflows/devops-health-groom.md |
Updated model default |
.github/workflows/devops-health-groom.lock.yml |
Regenerated compiled workflow |
.github/workflows/devops-health-check.md |
Updated model default |
.github/workflows/devops-health-check.lock.yml |
Regenerated compiled workflow |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Use gpt-5.6-sol for four health and triage workflows. Use claude-sonnet-5 alongside unchanged gpt-5.6-luna in default and full evaluation profiles, retain model overrides and judge routing, and cover actual profile selection in workflow tests. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
evaluation.yml has an unresolved comparison-label issue and a stale schedule-documentation nit.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review tier: Lite
Findings: 1
New issues introduced by this change (1)
| Severity | Finding |
|---|---|
.github/workflows/evaluation.yml — Update the remaining schedule comment to Sonnet 5 |
|
👋 @AbhitejJohn — this PR has 1 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the |
Replace Opus 4.8 judging with Haiku 4.5 while retaining Opus executor profiles and cross-family Terra judging for Claude and MAI. Disable duplicate secondary judging, update routing regression tests and guidance, and correct the stale Sonnet schedule comment. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Scheduled judge routing does not match the stated unchanged-behavior requirement.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review tier: Lite
Findings: 1
New issues introduced by this change (1)
| Severity | Finding |
|---|---|
.github/workflows/evaluation.yml — Preserve the scheduled judge contract |
Issues resolved since last review (1)
| Severity | Finding |
|---|---|
.github/workflows/evaluation.yml — Update the remaining schedule comment to Sonnet 5 View resolved comment |
Restore Opus 4.8 primary judges and scheduled Haiku 4.5 secondary judges at the user's request. Keep the corrected Sonnet schedule comment, Sol triage defaults, Sonnet 5 evaluation executors, and expanded schedule regression coverage. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
Corrective judge-restoration validation is still pending.
Review tier: Lite
Findings: None
Issues resolved since last review (2)
| Severity | Finding |
|---|---|
.github/workflows/evaluation.yml — Preserve the scheduled judge contract View resolved comment |
|
.github/workflows/evaluation.yml — Update the remaining schedule comment to Sonnet 5 View resolved comment |
|
❌ Evaluation did not complete successfully (the evaluate job reported 2 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete. |
|
❌ Evaluation did not complete successfully (the evaluate job reported 22 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete. |
Guard the pinned SDK's concurrent lazy startup and wait for filesystem-provider readiness before create/resume. Use the trusted launcher for evaluation and comparison without changing model routing, trial concurrency, or result gates. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Validate the same startup guard against SDK 1.0.11 used by the PR merge checks and SDK 1.0.13 on this branch. Keep unknown SDK versions fail-closed. Mock the runtime path explicitly in transport-only tests. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
✅ Evaluation passed for |
📊 Skill Evaluation Results16 model/skill results across 8 skills and 2 models — ✅ 3 improved, ➖ 6 not proven improved, Measurement identity: evaluated commit Measurement health: 16 expected / 16 observed / 16 written; 0 missing, 0 unexpected, 6 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
|
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Make Azure.ResourceManager AOT-compatible | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Make Azure.ResourceManager AOT-compatible:Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
⚠️ Underpowered — dotnet-aot-compat (gpt-5.6-luna)
Why: Net win +100.0% (1W/0T/0L over 1 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +100.0% across 1 paired run(s) — underpowered (1 preference-eligible stimulus vote(s); a credible verdict needs at least 5, and this eval won every one of them) — add distinct, discriminating stimuli; repeated runs do not increase task breadth
Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.
State: INVALID_INCONCLUSIVE (underpowered)
Gate evidence: n=1; 1W/0T/0L; d=1; p=0.500; net +100.0%
Overfit: Moderate (score 0.28)
Repeated-run reliability (not used by the gate): 1 paired run (1W/0T/0L).
⚠️ Underpowered — exp-test-maintainability (claude-sonnet-5)
Why: Net win +75.0% (3W/1T/0L over 4 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +45.0% across 4 paired run(s) — underpowered (4 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth
Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.
State: INVALID_INCONCLUSIVE (underpowered)
Gate evidence: n=4; 3W/1T/0L; d=3; p=0.125; net +75.0%
Overfit: Moderate (score 0.39)
Repeated-run reliability (not used by the gate): 4 paired runs (3W/1T/0L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Recognize tests with minimal boilerplate that need no refactoring | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Recognize tests with minimal boilerplate that need no refactoring:Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
⚠️ Underpowered — exp-test-maintainability (gpt-5.6-luna)
Why: Net win +50.0% (3W/0T/1L over 4 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +35.0% across 4 paired run(s) — underpowered (4 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth
Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.
State: INVALID_INCONCLUSIVE (underpowered)
Gate evidence: n=4; 3W/0T/1L; d=4; p=0.312; net +50.0%
Overfit: Low (score 0.10)
Repeated-run reliability (not used by the gate): 4 paired runs (3W/0T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Recognize tests with minimal boilerplate that need no refactoring | Eligible | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
Recognize tests with minimal boilerplate that need no refactoring:Both responses are accurate, avoid over-engineering, and correctly frame the code as already well-structured. B is more polished with priorities and trade-offs and better addresses the shared-field point. However, the core of the task—identifying the repeated individual test m...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
⚠️ Underpowered — migrate-nullable-references (claude-sonnet-5)
Why: Net win -33.3% (0W/2T/1L over 3 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -33.3% across 3 paired run(s) — underpowered (3 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth
Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.
State: INVALID_INCONCLUSIVE (underpowered)
Gate evidence: n=3; 0W/2T/1L; d=1; p=0.500; net -33.3%
Warnings: Activation: isolated 3/3; plugin 2/3
Overfit: Moderate (score 0.31)
Repeated-run reliability (not used by the gate): 3 paired runs (0W/2T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Enable NRT in ASP.NET Core Web API with EF Core | Eligible | -100.0% | -100.0% | 0/0/1 |
| = Enable NRT in a small library with mixed nullability | Eligible | +0.0% | +0.0% | 0/1/0 |
| = File-by-file migration: only modify the targeted file | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Enable NRT in ASP.NET Core Web API with EF Core:A's annotations consistently reflect the entity and API optionality requirements, especially the optional summary/category response fields. B compiles, and gets EF navigations and DbSets right, but its blanket required DTO strings misrepresents optional Summary and CategoryNam...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
⚠️ Underpowered — migrate-nullable-references (gpt-5.6-luna)
Why: Net win -33.3% (0W/2T/1L over 3 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -13.3% across 3 paired run(s) — underpowered (3 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth
Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.
State: INVALID_INCONCLUSIVE (underpowered)
Gate evidence: n=3; 0W/2T/1L; d=1; p=0.500; net -33.3%
Overfit: Low (score 0.07)
Repeated-run reliability (not used by the gate): 3 paired runs (0W/2T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Enable NRT in ASP.NET Core Web API with EF Core | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Enable NRT in a small library with mixed nullability | Eligible | +0.0% | +0.0% | 0/1/0 |
| = File-by-file migration: only modify the targeted file | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Enable NRT in ASP.NET Core Web API with EF Core:Both responses enabled nullable, kept DbSets non-nullable, correctly marked Book.Category as nullable, and achieved zero-warning builds. The differentiator is semantic annotation quality on optional/descriptive fields. Response B's visible timeline reveals a blanket regex appl...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
📉 Preference loss (report only) — migrate-dotnet9-to-dotnet10 (gpt-5.6-luna)
Why: Net win -82.4% (1W/1T/15L over 17 preference-eligible stimulus vote(s), sign test p=0.000), mean preference -36.5% across 17 paired run(s) — credibly worse
Next action: Inspect losing stimuli and fix skill behavior; this is not objective completion proof.
State: VALID_NO_CHANGE (preference_regression_report_only)
Gate evidence: n=17; 1W/1T/15L; d=16; p=0.000; net -82.4%
Overfit: Low (score 0.07)
Repeated-run reliability (not used by the gate): 17 paired runs (1W/1T/15L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ ASP.NET Core app with OpenAPI transformers using Microsoft.OpenApi v1 APIs | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ ASP.NET Core app with WebHostBuilder, OpenAPI, and forwarded headers | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ App using SslStream properties and SystemEvents | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Blazor WASM app with generic math shift masking and tar operations | Eligible | -100.0% | -100.0% | 0/0/1 |
| ▼ C# 14 compiler breaking changes — field keyword, extension keyword, disposal | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Console app with System.Linq.Async, SIGTERM, and BufferedStream | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Containerized single-file app with P/Invoke and IDispatchEx | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Cryptography app with OpenSSL, X.509, and Rfc2898DeriveBytes | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ EF Core app with Azure SQL JSON columns and parameterized collections | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ EF Core app with dynamic ExecuteUpdate and complex types | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Expression tree code broken by C# 14 span overload resolution | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ JSON polymorphism with conflicting property names and XmlSerializer | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Library with NuGet auditing, transitive deps, and InlineArray | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ SDK and NuGet obscure tooling changes | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ SQLite app with DateTimeOffset timezone handling | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ WinForms and WPF desktop app with System.Drawing and DynamicResource | Eligible | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
ASP.NET Core app with OpenAPI transformers using Microsoft.OpenApi v1 APIs:Both responses are high quality, correct, and cover all five rubric points with essentially the same solution. A is marginally more accurate: it grounds its answer in fetched source files, uses the more complete OpenApiSecuritySchemeReference constructor with the host document...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — migrate-dotnet10-to-dotnet11 (claude-sonnet-5)
Why: Net win +9.1% (6W/0T/5L over 11 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +14.5% across 11 paired run(s) — not credible (sign test p=0.500 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=11; 6W/0T/5L; d=11; p=0.500; net +9.1%
Overfit: Low (score 0.09)
Repeated-run reliability (not used by the gate): 11 paired runs (6W/0T/5L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ BackgroundService exceptions and ZipArchive CRC32 validation | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Basic TFM update with Docker and global.json | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Cryptography app using DSA on macOS | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ EF Core app with Cosmos DB provider using sync APIs | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ mTLS and AIA certificate chain validation | Eligible | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
BackgroundService exceptions and ZipArchive CRC32 validation:Both answers address the two requested breaking changes well, but A is more technically nuanced and actionable on BackgroundService host-task faulting and is marginally more complete on ZIP validation. B is concise but oversimplifies the host behavior.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — migrate-dotnet10-to-dotnet11 (gpt-5.6-luna)
Why: Net win -18.2% (3W/3T/5L over 11 preference-eligible stimulus vote(s), sign test p=0.363), mean preference -7.3% across 11 paired run(s) — no improvement
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=11; 3W/3T/5L; d=8; p=0.363; net -18.2%
Overfit: Low (score 0.07)
Repeated-run reliability (not used by the gate): 11 paired runs (3W/3T/5L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ ASP.NET Core app with OpenAPI customizations and Blazor Virtualize | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ BackgroundService exceptions and ZipArchive CRC32 validation | Eligible | -100.0% | -40.0% | 0/0/1 |
| = C# 15 compiler breaking changes — Span safe-context, nameof, with() | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ C# 15 dynamic operator and ref readonly delegate issues | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Cryptography app using DSA on macOS | Eligible | +0.0% | +0.0% | 0/1/0 |
| = EF Core SQL Server with Entra ID auth and Design package dependency | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ EF Core app with Cosmos DB provider using sync APIs | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ mTLS and AIA certificate chain validation | Eligible | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
ASP.NET Core app with OpenAPI customizations and Blazor Virtualize:Both responses correctly and completely address both parts of the question: the Microsoft.OpenApi v2→v3 breaking changes and the Virtualize OverscanCount default change from 3 to 15, each with appropriate remediation. Response A is slightly stronger because it provides a concr...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — migrate-dotnet8-to-dotnet9 (claude-sonnet-5)
Why: Net win +16.7% (6W/2T/4L over 12 preference-eligible stimulus vote(s), sign test p=0.377), mean preference +16.7% across 12 paired run(s) — not credible (sign test p=0.377 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=12; 6W/2T/4L; d=10; p=0.377; net +16.7%
Overfit: Low (score 0.07)
Repeated-run reliability (not used by the gate): 12 paired runs (6W/2T/4L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ App with JsonDocument null deserialization and BinaryFormatter fallback | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ App with empty environment variables, ZIP encoding, and keyed DI services | Eligible | -100.0% | -40.0% | 0/0/1 |
| = C# 13 compiler breaking changes — InlineArray on record, iterator safe context, collection expressions | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ EF Core Cosmos DB app with existing documents and composite id format | Eligible | -100.0% | -40.0% | 0/0/1 |
| = EF Core app with migration patterns and Cosmos DB discriminator | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Library with String.Trim span overload, keyed services, and InlineArray | Eligible | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
App with JsonDocument null deserialization and BinaryFormatter fallback:A more nearly fulfills the requested migration by targeting net9.0 and resolving all three identified .NET 9 breaking behaviors, including replacing BinaryFormatter. B correctly diagnoses the issue and responsibly notes serialization-format tradeoffs, but does not finish the r...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — migrate-dotnet8-to-dotnet9 (gpt-5.6-luna)
Why: Net win +41.7% (7W/3T/2L over 12 preference-eligible stimulus vote(s), sign test p=0.090), mean preference +31.7% across 12 paired run(s) — not credible (sign test p=0.090 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=12; 7W/3T/2L; d=9; p=0.090; net +41.7%
Warnings: Activation: isolated 12/12; plugin 11/12
Overfit: Low (score 0.07)
Repeated-run reliability (not used by the gate): 12 paired runs (7W/3T/2L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ C# 13 compiler breaking changes — InlineArray on record, iterator safe context, collection expressions | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Containerized app with env var precedence reversal and zlib removal | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Containerized app with zlib dependency and runtime configuration | Eligible | -100.0% | -40.0% | 0/0/1 |
| = EF Core Cosmos DB app with existing documents and composite id format | Eligible | +0.0% | +0.0% | 0/1/0 |
| = EF Core app with migration patterns and Cosmos DB discriminator | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
C# 13 compiler breaking changes — InlineArray on record, iterator safe context, collection expressions:Both responses are technically accurate and cover all three breaking changes correctly with the right fixes. A did genuine research (fetched Microsoft docs and the language proposal) and provides slightly richer, still-accurate detail (AllowUnsafeBlocks, cross-yield pointer ca...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — thread-abort-migration (claude-sonnet-5)
Why: Net win +60.0% (4W/0T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +24.0% across 5 paired run(s) — not credible (sign test p=0.188 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%
Warnings: Activation: isolated 4/5; plugin 4/5
Overfit: Moderate (score 0.33)
Repeated-run reliability (not used by the gate): 5 paired runs (4W/0T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▲ Thread.Join and Thread.Sleep only — should not migrate | Eligible | +100.0% | +40.0% | 1/0/0 |
| ▼ Worker thread with abort-based cancellation | Eligible | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
Worker thread with abort-based cancellation:A supplies the more complete and robust primary migration: it uses an appropriate dedicated background thread for potentially blocking work, does bounded cooperative Join without unsafe disposal after a timeout, includes generic exception handling, and offers concrete cancella...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — thread-abort-migration (gpt-5.6-luna)
Why: Net win +60.0% (4W/0T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +24.0% across 5 paired run(s) — not credible (sign test p=0.188 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%
Warnings: Activation: isolated 4/5; plugin 5/5
Overfit: Moderate (score 0.33)
Repeated-run reliability (not used by the gate): 5 paired runs (4W/0T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Blocking WaitHandle with Thread.Interrupt | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▲ Thread.Join and Thread.Sleep only — should not migrate | Eligible | +100.0% | +40.0% | 1/0/0 |
Illustrative judge evidence:
Blocking WaitHandle with Thread.Interrupt:Both responses are high quality, correct, and take the same fundamentally sound Channel<T> + cooperative cancellation approach. B edges out on two specific rubric points (explicit PlatformNotSupportedException note and the WaitSleepJoin/wait-state warning). However, A is more ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — exp-mock-usage-analysis (claude-sonnet-5)
Why: Net win +83.3% (5W/1T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +63.3% across 6 paired run(s) — credibly better
Next action: Review overfit evidence.
State: VALID_PASS (credible_preference_improvement)
Gate evidence: n=6; 5W/1T/0L; d=5; p=0.031; net +83.3%
Overfit: Moderate (score 0.23)
Repeated-run reliability (not used by the gate): 6 paired runs (5W/1T/0L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Analyze mock usage in FakeItEasy tests | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Analyze mock usage in FakeItEasy tests:Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Routine passing details for 2 results are in Full Results.
🔍 Full Results - all metrics and investigation details
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1154 in dotnet/skills, download eval artifacts with
gh run download 34652384279 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/b8083dd46b40bf81675c3d1dac06a33d8e67d0c6/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
📊 Skill Evaluation Results26 model/skill results across 13 skills and 2 models — ✅ 15 improved, ➖ 7 not proven improved, Measurement identity: evaluated commit Measurement health: 26 expected / 26 observed / 26 written; 0 missing, 0 unexpected, 4 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
|
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Investigate root cause of a .NET MAUI iOS crash | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Parse .NET frames and locate dSYMs from an iOS crash log | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Reject Android tombstone passed as iOS crash log | Eligible | -100.0% | -100.0% | 0/0/1 |
Illustrative judge evidence:
Investigate root cause of a .NET MAUI iOS crash:A is the more evidence-based and technically careful result: it validates and interprets the CoreCLR symbols, distinguishes the long-lived runtime/main-loop frames from the unknown application failure, and gives actionable app-dSYM symbolication commands. B is readable and off...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
⚠️ Underpowered — apple-crash-symbolication (gpt-5.6-luna)
Why: Net win +66.7% (2W/1T/0L over 3 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +46.7% across 3 paired run(s) — underpowered (3 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth
Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.
State: INVALID_INCONCLUSIVE (underpowered)
Gate evidence: n=3; 2W/1T/0L; d=2; p=0.250; net +66.7%
Warnings: Activation: isolated 2/3; plugin 3/3
Overfit: Low (score 0.16)
Repeated-run reliability (not used by the gate): 3 paired runs (2W/1T/0L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Reject Android tombstone passed as iOS crash log | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Reject Android tombstone passed as iOS crash log:Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
⚠️ Underpowered — microbenchmarking (claude-sonnet-5)
Why: Net win +100.0% (1W/0T/0L over 1 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +40.0% across 1 paired run(s) — underpowered (1 preference-eligible stimulus vote(s); a credible verdict needs at least 5, and this eval won every one of them) — add distinct, discriminating stimuli; repeated runs do not increase task breadth
Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.
State: INVALID_INCONCLUSIVE (underpowered)
Gate evidence: n=1; 1W/0T/0L; d=1; p=0.500; net +100.0%
Overfit: Moderate (score 0.41)
Repeated-run reliability (not used by the gate): 1 paired run (1W/0T/0L).
⚠️ Underpowered — microbenchmarking (gpt-5.6-luna)
Why: Net win +100.0% (1W/0T/0L over 1 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +100.0% across 1 paired run(s) — underpowered (1 preference-eligible stimulus vote(s); a credible verdict needs at least 5, and this eval won every one of them) — add distinct, discriminating stimuli; repeated runs do not increase task breadth
Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.
State: INVALID_INCONCLUSIVE (underpowered)
Gate evidence: n=1; 1W/0T/0L; d=1; p=0.500; net +100.0%
Overfit: Low (score 0.17)
Repeated-run reliability (not used by the gate): 1 paired run (1W/0T/0L).
➖ Not proven improved — analyzing-dotnet-performance (claude-sonnet-5)
Why: Net win +45.5% (7W/2T/2L over 11 preference-eligible stimulus vote(s), sign test p=0.090), mean preference +29.1% across 11 paired run(s) — not credible (sign test p=0.090 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=11; 7W/2T/2L; d=9; p=0.090; net +45.5%
Overfit: High (score 0.50)
Repeated-run reliability (not used by the gate): 11 paired runs (7W/2T/2L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Detects CurrentCulture comparer and compiled regex budget in inflection rules | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Finds per-call Dictionary allocation not hoisted to static | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Finds repeated enumeration in time and collection formatting | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Flags Span inconsistencies and compound method chains in truncation library | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Detects CurrentCulture comparer and compiled regex budget in inflection rules:Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — android-tombstone-symbolication (claude-sonnet-5)
Why: Net win -62.5% (1W/1T/6L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference -25.0% across 8 paired run(s) — no improvement
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=8; 1W/1T/6L; d=7; p=0.063; net -62.5%
Warnings: Activation: isolated 7/8; plugin 8/8
Overfit: Moderate (score 0.48)
Repeated-run reliability (not used by the gate): 8 paired runs (1W/1T/6L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Handle .NET frames with no BuildId metadata | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Recognize NativeAOT tombstone with app binary and libSystem.Native.so | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Recognize tombstone with no .NET frames | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Reject iOS crash log as wrong format | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Symbolicate CoreCLR frames in an Android tombstone | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Symbolicate multi-thread tombstone | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Symbolicate tombstone with multiple .NET libraries and different BuildIds | Eligible | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
Handle .NET frames with no BuildId metadata:Both responses correctly determine that symbolication cannot proceed from this truncated tombstone and give sound next steps. A is marginally stronger because it reports the parsed frame count and individual library/offset evidence, and its local-runtime-pack/build metadata gu...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — android-tombstone-symbolication (gpt-5.6-luna)
Why: Net win +37.5% (4W/3T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +30.0% across 8 paired run(s) — not credible (sign test p=0.188 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=8; 4W/3T/1L; d=5; p=0.188; net +37.5%
Overfit: Moderate (score 0.21)
Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Handle .NET frames with no BuildId metadata | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Recognize tombstone with no .NET frames | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Reject iOS crash log as wrong format | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Symbolicate CoreCLR frames in an Android tombstone | Eligible | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
Handle .NET frames with no BuildId metadata:Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — clr-activation-debugging (claude-sonnet-5)
Why: Net win -14.3% (3W/0T/4L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -22.9% across 7 paired run(s) — no improvement
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=7; 3W/0T/4L; d=7; p=0.500; net -14.3%
Warnings: Activation: isolated 7/7; plugin 6/7
Overfit: Moderate (score 0.31)
Repeated-run reliability (not used by the gate): 7 paired runs (3W/0T/4L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▲ Decline non-CLR-activation issue | Eligible | +100.0% | +40.0% | 1/0/0 |
| ▼ Diagnose FOD suppressed but activation still failing | Eligible | -100.0% | -100.0% | 0/0/1 |
| ▼ Diagnose unexpected FOD dialog from native build tool | Eligible | -100.0% | -100.0% | 0/0/1 |
| ▼ Explain why same binary behaves differently under different launch methods | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Identify multiple activation sequences in a single log | Eligible | -100.0% | -100.0% | 0/0/1 |
Illustrative judge evidence:
Diagnose FOD suppressed but activation still failing:A finds and interprets the available activation log, gives a coherent causal chain for both the activation failure and silent behavior, and proposes a concrete fix. B merely asks the user to provide logs that were available in the environment, so it does not answer the task at...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — clr-activation-debugging (gpt-5.6-luna)
Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +22.9% across 7 paired run(s) — not credible (sign test p=0.109 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%
Overfit: Low (score 0.18)
Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Analyze healthy managed EXE activation | Eligible | -100.0% | -100.0% | 0/0/1 |
| = Identify multiple activation sequences in a single log | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Analyze healthy managed EXE activation:Response A located the log file in the filesystem (after its view tool failed, it recovered via bash/find) and delivered a correct, thorough assessment confirming normal activation, runtime v4.0.30319, and absence of errors. Response B activated a relevant skill but its file r...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — dump-collect (gpt-5.6-luna)
Why: Net win +44.4% (6W/1T/2L over 9 preference-eligible stimulus vote(s), sign test p=0.145), mean preference +24.4% across 9 paired run(s) — not credible (sign test p=0.145 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=9; 6W/1T/2L; d=8; p=0.145; net +44.4%
Warnings: Activation: isolated 9/9; plugin 7/9
Overfit: Low (score 0.16)
Repeated-run reliability (not used by the gate): 9 paired runs (6W/1T/2L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Configure CoreCLR dump collection in Alpine Docker as non-root | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Configure automatic crash dumps for CoreCLR app on Linux | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▲ Decline dump analysis request | Eligible | +100.0% | +100.0% | 1/0/0 |
| ▼ Recover crash dump from macOS NativeAOT without createdump | Eligible | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
Configure CoreCLR dump collection in Alpine Docker as non-root:Both responses fully satisfy every rubric criterion: enabling DOTNET_DbgEnableMiniDump, adding SYS_PTRACE, handling non-root dump directory permissions, avoiding false Alpine claims, and not analyzing dumps. Response A adds a named-volume workflow with dump extraction steps an...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — migrate-nunit-to-mstest (claude-sonnet-5)
Why: Net win +28.6% (4W/10T/0L over 14 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +11.4% across 14 paired run(s) — not credible — 10 of 14 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=14; 4W/10T/0L; d=4; p=0.063; net +28.6%
Overfit: Moderate (score 0.29)
Repeated-run reliability (not used by the gate): 14 paired runs (4W/10T/0L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Convert NUnit combinatorial parameter data | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Convert NUnit theory datapoint discovery | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Convert cooperative cancellation attribute | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Preserve NUnit pairwise case set | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Preserve NUnit single-instance fixture state | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Preserve STA across async continuations | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Preserve TestCaseSource result and row metadata | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Preserve all-inconclusive NUnit theory result | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Preserve assembly-wide SetUpFixture scope | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Translate explicit NUnit parallelization | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Convert NUnit combinatorial parameter data:Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — dotnet-trace-collect (claude-sonnet-5)
Why: Net win +82.4% (14W/3T/0L over 17 preference-eligible stimulus vote(s), sign test p=0.000), mean preference +57.6% across 17 paired run(s) — credibly better
Next action: Review overfit evidence.
State: VALID_PASS (credible_preference_improvement)
Gate evidence: n=17; 14W/3T/0L; d=14; p=0.000; net +82.4%
Overfit: High (score 0.53)
Repeated-run reliability (not used by the gate): 17 paired runs (14W/3T/0L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Excessive GC on Linux (.NET 8) | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Memory leak on .NET Framework Windows | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Memory leak on Linux (.NET 8) | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Excessive GC on Linux (.NET 8):Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — dotnet-trace-collect (gpt-5.6-luna)
Why: Net win +70.6% (14W/1T/2L over 17 preference-eligible stimulus vote(s), sign test p=0.002), mean preference +45.9% across 17 paired run(s) — credibly better
Next action: Review overfit evidence.
State: VALID_PASS (credible_preference_improvement)
Gate evidence: n=17; 14W/1T/2L; d=16; p=0.002; net +70.6%
Overfit: Moderate (score 0.21)
Repeated-run reliability (not used by the gate): 17 paired runs (14W/1T/2L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Container installation without .NET SDK | Eligible | -100.0% | -100.0% | 0/0/1 |
| = Excessive GC on Linux (.NET 8) | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Linux pre-.NET 10 needing native call stacks | Eligible | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
Container installation without .NET SDK:Response A directly answers the intent of the question—getting dotnet-trace without the SDK—via the aka.ms self-contained single-file download and a curl command, satisfying the two most specific rubric criteria. Response B, while technically valid and thorough, uses the .NET ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — dump-collect (claude-sonnet-5)
Why: Net win +66.7% (6W/3T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +33.3% across 9 paired run(s) — credibly better
Next action: Fix activation gaps; Review overfit evidence.
State: VALID_PASS (credible_preference_improvement)
Gate evidence: n=9; 6W/3T/0L; d=6; p=0.016; net +66.7%
Warnings: Activation: isolated 8/9; plugin 8/9
Overfit: High (score 0.59)
Repeated-run reliability (not used by the gate): 9 paired runs (6W/3T/0L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Advisory: NativeAOT Kubernetes dump collection setup | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Configure automatic crash dumps for CoreCLR app on Linux | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▲ Decline dump analysis request | Eligible | +100.0% | +40.0% | 1/0/0 |
| = Detect runtime and configure crash dumps for unknown .NET app on Linux | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Advisory: NativeAOT Kubernetes dump collection setup:Position-swap inconsistent (forward: tie, reverse: B). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — migrate-mstest-v1v2-to-v3 (claude-sonnet-5)
Why: Net win +63.6% (8W/2T/1L over 11 preference-eligible stimulus vote(s), sign test p=0.020), mean preference +30.9% across 11 paired run(s) — credibly better
Next action: Review overfit evidence.
State: VALID_PASS (credible_preference_improvement)
Gate evidence: n=11; 8W/2T/1L; d=9; p=0.020; net +63.6%
Overfit: Moderate (score 0.38)
Repeated-run reliability (not used by the gate): 11 paired runs (8W/2T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Correctly identify MSTest v1 vs v2 and recommend different migration paths | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Fix Assert.AreEqual object overload errors after v3 upgrade | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Migrate from .testsettings to .runsettings | Eligible | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
Correctly identify MSTest v1 vs v2 and recommend different migration paths:Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — migrate-mstest-v3-to-v4 (claude-sonnet-5)
Why: Net win +46.7% (8W/6T/1L over 15 preference-eligible stimulus vote(s), sign test p=0.020), mean preference +26.7% across 15 paired run(s) — credibly better
Next action: Review overfit evidence.
State: VALID_PASS (credible_preference_improvement)
Gate evidence: n=15; 8W/6T/1L; d=9; p=0.020; net +46.7%
Overfit: Moderate (score 0.35)
Repeated-run reliability (not used by the gate): 15 paired runs (8W/6T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Fix TestMethodAttribute CallerInfo constructor breaking change | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Fix TestMethodAttribute and TestMethod display name constructor | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Fix multiple v4 breaking changes: Assert, ClassCleanup, TestContext, Timeout | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Full MSTest v3 to v4 migration with multiple breaking changes | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Handle net6.0 target framework dropped in MSTest v4 | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Understand behavioral changes after MSTest v4 upgrade | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Verified MSTest v3 to v4 package update with build and test | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Fix TestMethodAttribute CallerInfo constructor breaking change:Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — migrate-mstest-v3-to-v4 (gpt-5.6-luna)
Why: Net win +46.7% (9W/4T/2L over 15 preference-eligible stimulus vote(s), sign test p=0.033), mean preference +26.7% across 15 paired run(s) — credibly better
Next action: Review overfit evidence.
State: VALID_PASS (credible_preference_improvement)
Gate evidence: n=15; 9W/4T/2L; d=11; p=0.033; net +46.7%
Overfit: Moderate (score 0.22)
Repeated-run reliability (not used by the gate): 15 paired runs (9W/4T/2L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Fix Assert.IsInstanceOfType out parameter removal | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Fix multiple v4 breaking changes: Assert, ClassCleanup, TestContext, Timeout | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Full MSTest v3 to v4 migration with multiple breaking changes | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Handle net6.0 target framework dropped in MSTest v4 | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Migrate custom TestMethodAttribute from Execute to ExecuteAsync | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Understand behavioral changes after MSTest v4 upgrade | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Fix Assert.IsInstanceOfType out parameter removal:Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — migrate-vstest-to-mtp (claude-sonnet-5)
Why: Net win +72.7% (9W/1T/1L over 11 preference-eligible stimulus vote(s), sign test p=0.011), mean preference +50.9% across 11 paired run(s) — credibly better
Next action: Review overfit evidence.
State: VALID_PASS (credible_preference_improvement)
Gate evidence: n=11; 9W/1T/1L; d=10; p=0.011; net +72.7%
Overfit: Moderate (score 0.33)
Repeated-run reliability (not used by the gate): 11 paired runs (9W/1T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Set OutputType=Exe only for test projects in Directory.Build.props | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Update Azure DevOps pipeline from VSTest task to MTP | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Set OutputType=Exe only for test projects in Directory.Build.props:The answers are substantively equivalent and correct. A is marginally more complete: it explicitly says every MTP-related property should remain in the conditioned group and offers a sensible per-project fallback if project naming is not reliable. B's example is nevertheless f...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — migrate-xunit-to-mstest (claude-sonnet-5)
Why: Net win +46.2% (7W/5T/1L over 13 preference-eligible stimulus vote(s), sign test p=0.035), mean preference +27.7% across 13 paired run(s) — credibly better
Next action: Review overfit evidence.
State: VALID_PASS (credible_preference_improvement)
Gate evidence: n=13; 7W/5T/1L; d=8; p=0.035; net +46.2%
Overfit: Moderate (score 0.43)
Repeated-run reliability (not used by the gate): 13 paired runs (7W/5T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Convert ITestOutputHelper to TestContext | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Convert Skip, Trait, and Timeout | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Handle ICollectionFixture explicitly (do not silently widen scope) | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Map exception assertions correctly (xUnit Throws -> MSTest ThrowsExactly) | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Preserve type and sequence assertion semantics | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Preserve xUnit parallelization default with [assembly: Parallelize] | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Convert ITestOutputHelper to TestContext:Both runs complete the migration correctly: xUnit packages and APIs are removed, MSTest attributes/assertions/TestContext are introduced, the net8.0 target and VSTest configuration are retained, and each reports a successful one-test run. A uses newer verified package versions...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — migrate-xunit-to-xunit-v3 (claude-sonnet-5)
Why: Net win +50.0% (6W/6T/0L over 12 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +25.0% across 12 paired run(s) — credibly better
Next action: Review overfit evidence.
State: VALID_PASS (credible_preference_improvement)
Gate evidence: n=12; 6W/6T/0L; d=6; p=0.016; net +50.0%
Overfit: Moderate (score 0.36)
Repeated-run reliability (not used by the gate): 12 paired runs (6W/6T/0L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Consolidate xunit.extensibility packages and remove xunit.abstractions | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Convert async void test methods to async Task | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Migrate project with YTest.MTP.XUnit2 to xUnit.net v3 preserving MTP | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Recognize project already on xUnit.net v3 — no migration needed | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Update BeforeAfterTestAttribute overrides with IXunitTest parameter | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Update custom FactAttribute to include source information parameters | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Consolidate xunit.extensibility packages and remove xunit.abstractions:Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Routine passing details for 6 results are in Full Results.
🔍 Full Results - all metrics and investigation details
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1154 in dotnet/skills, download eval artifacts with
gh run download 34652125571 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/a226ead6ff24b56338979737651609f9aef8e736/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.


Summary
Use separate defaults for operational workflows and skill evaluations:
devops-health-check,devops-health-groom,devops-health-investigate,issue-triagegpt-5.6-solclaude-sonnet-5and unchangedgpt-5.6-lunaRegenerate the four agentic lock files with gh-aw v0.86.2. Preserve model-variable precedence, the
copilot-pat-poolenvironment, secret mappings, permissions, prompts, safe outputs, and action/container pins. Replace Sonnet 4.6 in thedefaultandfullevaluation presets and use the existing Sonnet 5 -> Terra judge route. Keep Luna, explicit profile selection, and thenewerpreset unchanged. Correct the stale Mon/Wed/Fri schedule comment identified by Copilot.Also repair concurrent SDK startup in the trusted evaluation launcher. This adds harness startup code, regression tests, launcher wiring in
evaluation-run.yml, secretless test wiring inevaluation-workflow-tests.yml, and investigation guidance. Dependency manifests and lockfiles are unchanged by this repair.Judge replacement is deferred. GPT executors retain Opus 4.8 as the primary judge and Haiku 4.5 as the scheduled secondary judge; PR and manual runs retain no secondary judge. No judge switch is included.
Actual PR evaluation path
evaluation.ymlselects thedefaultprofile for ordinary PR events and unqualified/evaluaterequests. Discovery expands that profile into executor/judge entries. Theevaluatejob passes those entries toevaluation-run.yml, which usesmatrix.entry.modelandmatrix.entry.judgein evaluation commands. This changes the PR evaluator, not only the issue-triage model.SDK startup repair
Proven local defects: the pinned SDK can start multiple transports for concurrent callers, and
createSession/resumeSessioncan use a connected transport before filesystem-provider setup finishes. The no-model regression commandnode --test eng/evaluation-tools/sdk-startup.test.mjsoriginally failed four cases: three runtime starts instead of one, earlysession.create, earlysession.resume, and duplicate starts before a failed-start retry. After the guard, all six cases pass on SDK 1.0.11 and 1.0.13, including independent-client startup and parallel session creation after readiness. The tests use the real SDK methods with controlled transport setup, not live model calls. Upstream startup work is tracked in github/copilot-sdk#2585.Run evidence: run 34627963458 reached Sonnet 5 execution, but both Sonnet 5 and unchanged Luna reported
Cannot set session filesystem provider while sessions are active. The advanced-plugin legs each recordedexpectedEvalCount=4,writtenResultCount=4, andmeasurementInvalidEvalCount=4. This is consistent with the reproduced startup defects; a complete repaired evaluation is still needed to confirm the end-to-end result.eng/evaluation-tools/sdk-startup.mjsshares one startup promise per SDK client and waits for readiness before creating or resuming sessions. Startup errors propagate to every waiter; a later attempt can retry. It accepts only the two tested SDK versions and fails explicitly on an unknown version.vally.mjsloads this guard before the Vally CLI. The workflow copies both files from the trusted tooling checkout and places the launcher first onPATH, so evaluation and comparison use the same guard. No global SDK installation or dependency source is modified. Credential scope, trusted-checkout boundaries, model/judge routing, trial counts, concurrency, and result gates remain unchanged.Core repair commit:
a226ead6f. Current correction:b8083dd46, which also covers SDK 1.0.11 selected by the PR merge checks, while this branch uses SDK 1.0.13. Evaluation 34652125571 was started ona226ead6f; evaluation 34652384279 is bound to current headb8083dd46. Both were in progress at this update. Neither is claimed to be a passing evaluation.The separate
publish-session-datafailure reportsAuthentication failedforhttps://github.com/dotnet/skills-data.git/. This repair does not change or repair that credential.Judge experiment evidence — decision deferred
Public comparison records pinned to commit 27b203fd cover publication timestamps August 26–September 8, 2026 UTC. The experiment originated in #1002; scheduled run 34094126994 contains primary and cross-judge artifacts.
Filter actual row identifiers
primaryJudge=claude-opus-4.8andsecondJudge=claude-haiku-4.5; the top-level label still names Sonnet 4.6 and is stale. The data contains 1,101 paired observations across 12 publication snapshots, 8 commits, 92 distinct plugin/skill keys, and 3 GPT executors. These are repeated observations, not independent trials.In the valid cohort, 109 observations pass with Opus but return
VALID_NO_CHANGEwith Haiku; 34 go the other way. Haiku produces fewer passes in this sample. These disagreements are not proven Haiku errors because there is no human reference verdict.This does not prove equivalence. The overall count includes 502 pairs where both results are
INVALID_INCONCLUSIVE. Repeated skills, changing commits, and selection of the valid cohort limit generalization. The data does not meter judge tokens, dollar cost, or duration; no measured savings are claimed. Independent cross-family review reproduced these counts and limitations. These results are retained for discussion, not as justification for a judge change in this PR.Reproduce counts from the pinned JSON with:
Expected all-pair output:
1101, 956, 945, 110, 35. Expected valid-cohort output:595, 452, 443, 109, 34, in the field order above.Operational evidence and limits
The seven model-failure reports #1120, #1125, #1126, #1129, #1130, #1132, and #1135 selected Sonnet 4.6 and returned HTTP 400: unavailable for integrator
agentic-workflows. Run 34091012534 lists Sol and Sonnet 5, but not Sonnet 4.6. This proves advertised integration availability, not a successful current invocation.Historical health-check, groom, and triage agents completed with auto-selected Sonnet 5. Those do not prove current Vally evaluation access.
This does not resolve HTTP 401 reports #1147, #1141, #1131, and #1128. No issues are automatically closed.
Validation and review
The Sol/Sonnet executor revision passed strict compilation of all four agentic workflows, 22 workflow tests, actionlint on
evaluation.yml, and independent Claude/GPT/Gemini review. The generated locks retain 22 pre-existing actionlint diagnostics; baseline and changed diagnostics match. This is not a clean full-lock lint result. Judge restoration in commit2c6c0ff76passed both targeted tests (2/2) and actionlint onevaluation.yml. At that commit, the restored evaluation configuration and both investigation guides matched pre-switch commitcf4f4ad08exactly, except for the corrected Sonnet 5 schedule comment. The later SDK repair adds the separate configuration and documentation described above.SDK repair: six startup regressions passed on each supported SDK version; 22 workflow tests passed; the launcher returned Vally version
0.14.0; actionlint passed for both changed workflow files. Independent Claude/GPT/Gemini review found no significant issues in the core repair. The first remote smoke test rejected SDK 1.0.11; the current correction adds that explicitly tested version rather than removing the version guard. Current-head remote checks and evaluation are not yet claimed green.Live workflow evaluations have been dispatched from this PR branch. No merge or credential changes were made.