Skip to content

Fix strict-mode wire schema to actually be strict-tool-use compatible - #101

Merged
exactml merged 3 commits into
masterfrom
fix/strict-schema-compat-v2
Sep 20, 2026
Merged

exactml merged 3 commits into
masterfrom
fix/strict-schema-compat-v2

Conversation

@exactml

@exactml exactml commented Sep 20, 2026

Copy link
Copy Markdown
Owner

Summary

This fixes a production regression from v0.2.3, which made every Anthropic-backed review fail 100% of the time.

  • AnthropicProvider.generate_structured's strict: true tool definition was rejected outright by the API (400 invalid_request_error: "For 'number' type, properties maximum, minimum are not supported") on every single call, because Finding.confidence's ge=0.0, le=1.0 and _FindingsResponse.findings' max_length render as minimum/maximum/maxItems -- keywords outside strict tool use's supported JSON Schema subset. Unlike the pre-v0.2.3 failure mode (intermittent, sometimes worked), this failed deterministically every time, making v0.2.3 strictly worse than v0.2.2 for anyone with an Anthropic models.reviewer configured.
  • _strict_input_schema now adapts the wire schema before sending it: strips confirmed-unsupported keywords (minimum, maximum, exclusiveMinimum, exclusiveMaximum, multipleOf, minLength, maxLength, maxItems, uniqueItems), rewrites a nullable field's anyOf: [{type: X}, {type: "null"}] (pydantic's rendering of X | None) into the {"type": [X, "null"]} form Anthropic's docs describe, and validates additionalProperties: false wherever it appears. Pydantic's own validation of the response still enforces the stripped bounds locally -- only the server-side guarantee for that bound is lost, not the check itself.
  • Design note, since the first draft of this fix used a denylist and it already had a gap (missing uniqueItems, caught in review): the final version checks every schema keyword against two explicit lists -- confirmed-unsupported (stripped) and confirmed-supported (kept) -- and raises ValueError immediately on anything in neither. A keyword nobody's vetted yet now fails a test the moment a new constrained field is added, instead of silently shipping to production the way this regression did.
  • Includes a regression test built directly against the real _FindingsResponse schema, plus tests for the allowlist's fail-loud behavior on an unvetted keyword (e.g. pattern) and the uniqueItems gap specifically.

Relates to #96 -- not a new issue, this is the correction to the fix from #99.

Test plan

  • make lint -- clean
  • make test -- 165 passed, including:
    • a regression test against the real _FindingsResponse schema asserting no forbidden keyword leaks into the sent request
    • the stripped bounds (confidence's ge/le) still reject an out-of-range value locally, proving the server-side guarantee loss doesn't become a client-side gap too
    • an unvetted keyword (e.g. pattern) raises ValueError before any request is sent
    • uniqueItems (the exact gap found in review of the first draft) is confirmed stripped
  • Once released, re-verify against a live PR with a non-trivial diff (the failure only reproduces with an actual Anthropic API call)

🤖 Generated with Claude Code

exactml and others added 3 commits September 20, 2026 15:10
v0.2.3's strict: true broke every Anthropic-backed review outright --
Finding.confidence's ge/le bounds and _FindingsResponse.findings'
max_length render as minimum/maximum/maxItems, none of which strict
tool use's supported JSON Schema subset allows, so the API rejected
the request with a 400 on every single call. Strips those keywords
from the wire schema (pydantic still enforces them locally on the
response) and rewrites a nullable anyOf into the single-type-array
form Anthropic's docs describe. Includes a regression test built
directly against _FindingsResponse.

Relates to ISSUE-96

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The denylist in the previous commit was missing uniqueItems -- a
gap found by review. Switched to two explicit lists: keywords
confirmed unsupported (stripped, pydantic still enforces them
locally) and keywords confirmed supported (kept as-is). Anything
in neither list raises ValueError immediately instead of silently
passing an unvetted keyword through, which is exactly how v0.2.3
shipped broken in the first place.

Relates to ISSUE-96

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@exactml exactml added the bug Something isn't working label Sep 20, 2026
@exactml
exactml merged commit a2d2467 into master Sep 20, 2026
1 of 2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant