feat: Validate Actor input and apply input-schema defaults - #72
Merged
Merged
Conversation
A build now records the input schema from the Actor's pushed source, and every run started against that build has its input validated against it with the schema's defaults applied - the last piece of run-start behaviour that was listed as unsupported. Resolution at build time follows what apify-cli looks for locally: the `input` field of `.actor/actor.json` (inline object or a path relative to `.actor/`), then `.actor/INPUT_SCHEMA.json`, then `INPUT_SCHEMA.json` at the Actor root, matched case-insensitively with every outcome stated in the build log. A schema that is itself invalid fails the build rather than producing an image whose every later run would silently skip validation. At run start, a build with a schema gets the platform's own behaviour: defaults fill every omitted field (nested objects and array items included), the filled-in input is what lands in the run's INPUT record - so a call with no input at all still runs on the defaults - and an input the schema rejects answers 400 without creating a run or a container. Validation itself is the platform's, through @apify/input_schema and the same AJV configuration the API uses, so the message a developer reads locally is the one the real API would give. Two deliberate local differences, both documented: proxy group availability is not checked (no proxy groups are emulated here), and encrypted secret input fields stay unsupported. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011U9yUTUn78bqGDn5obYSXs
…sd93fv # Conflicts: # requirements/unsupported.md
Master's contribution conventions landed after this branch's feature commit: requirements describe what, not how, and comments carry the missing context rather than restating the code. Applied to the input-schema spec sections and the modules they describe; no behaviour change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011U9yUTUn78bqGDn5obYSXs
…sd93fv # Conflicts: # requirements/api.md # requirements/unsupported.md # src/api/routes/actors.ts
The e2e case asserted only the dataset item count for a call with no input, which the sample Actor's own internal fallback for a missing maxPages produces just as well. It now reads back the run's INPUT record through the CLI, which only the runtime writes, so it really distinguishes the schema's defaults from the Actor defaulting on its own. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011U9yUTUn78bqGDn5obYSXs
The input-schema resolver was a near copy of the Dockerfile one, down to the Actor-root containment check. Both now read the pushed source through one module, so a traversal fix cannot land in only one of them, and an unparseable .actor/actor.json is reported by the resolver itself instead of relying on the Dockerfile resolver having run first. Also: one fail-the-build path in place of three copies, the two platform error types moved to the shared error factories, the schema decision moved out of the HTTP route into the service, defaults no longer clone an input nobody else holds, and the comments that restated their own code are gone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011U9yUTUn78bqGDn5obYSXs
…ntions The sections described how the feature is implemented rather than what it does for a user: the schema lookup was explained in terms of what apify-cli does internally, and the API section talked about stored schemas being compiled. Both now state the behaviour and nothing else. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011U9yUTUn78bqGDn5obYSXs
Every sample repeated its input schema's defaults as an in-code fallback, so a missing field was silently papered over and the default lived in two places. The runtime (like the platform) fills those fields in before the run starts, so the fallbacks are gone and the schema is the only place a default is written. The two Playwright samples' start URLs had only a prefill, so they now carry a default too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011U9yUTUn78bqGDn5obYSXs
A run's log now says how many `apify-actor-start` events it was pre-charged and what they cost, and warns when the requested memory is one the Apify platform would refuse. The memory is still applied as asked. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011U9yUTUn78bqGDn5obYSXs
Both simple samples now ship pricing for `apify-default-dataset-item` alongside `apify-actor-start`, so one run of a priced sample shows every way an event can be charged - two by the Actor, two by the runtime. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011U9yUTUn78bqGDn5obYSXs
The CPU and memory figures are accumulated in memory, and were written to the run record only after the terminal status. A client that waited for the run to finish and then read the record could find them missing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011U9yUTUn78bqGDn5obYSXs
The bundled `pricing.json` prices an Actor that has none yet; applying it to one that is already priced is refused, since the array is append-only. Both the README and the skill now give the append command. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011U9yUTUn78bqGDn5obYSXs
A rejected `pricingInfos` update now says which field of which entry differs from the stored one, and what was sent and stored there, instead of leaving the caller to diff two arrays by eye. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011U9yUTUn78bqGDn5obYSXs
`usageUsd.ACTOR_COMPUTE_UNITS` is rounded to six decimals, so a product whose seventh decimal is a 5 moves by the full tolerance of `toBeCloseTo(..., 6)` and failed on the boundary. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011U9yUTUn78bqGDn5obYSXs
The block's own lines already carry the runtime marker, so the two `=` rules added width without adding information. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011U9yUTUn78bqGDn5obYSXs
…ferences Seven bullets restating platform behaviour become one that names it, one listing what this runtime does differently, and the local dev-folder caveat. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011U9yUTUn78bqGDn5obYSXs
Pijukatel
marked this pull request as ready for review
September 23, 2026 12:47
Pijukatel
pushed a commit
that referenced
this pull request
Sep 23, 2026
Two conflicts in requirements/api.md: - The run-start input validation #72 documents sits outside the section this branch deletes, so it is kept as #72 wrote it, above this branch's "Five endpoints are exceptions to the {data} envelope" line. - The Actor-runtime section: kept this branch's, as before. The OpenAPI document is normative for the upstream-fallback contract, and #72 extends the never-relayed error types, so setApiFallbackState's exhaustive list now names invalid-input and invalid-input-schema too. No conflict flagged that - #72's own edit to the same list lands inside the prose this branch deletes - and the list claims to be exhaustive, so leaving them out would have made the specification wrong. Nothing else in #72 reaches the namespace: input validation is faithful to the platform, message for message, so it warrants no platform note. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016VvUV6661cyYbxoHZv2Pby
Pijukatel
pushed a commit
that referenced
this pull request
Sep 24, 2026
Clean merge of #74 (run memory from .actor/actor.json). No conflicts: its edits to runs.ts and SKILL.md land away from this branch's. No change to the OpenAPI document. Memory sizing is platform-faithful behaviour on the emulated /v2 surface, not runtime behaviour grafted onto an endpoint, so it warrants no x-actor-runtime-platform-notes entry - the same call made for input validation (#72), pay-per-event charging (#69) and the runs/last shortcuts (#68). Its local differences (log notes, no plan cap, a bad memory field failing the build) are emulation-fidelity notes, which live in requirements/actor-driver.md, where #74 put them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016VvUV6661cyYbxoHZv2Pby
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
An Actor's input schema is now read at build time - from the
inputfield of.actor/actor.json,.actor/INPUT_SCHEMA.json, orINPUT_SCHEMA.jsonat the Actor root - and every run of that build has its input validated against it with the schema's defaults applied.A run started with no input at all therefore runs on the schema's defaults, and a run whose input the schema rejects never starts, answering with the same message the Apify API gives. A build whose input schema cannot be read, or is not a valid input schema, fails with the reason in its log instead of producing an image whose runs skip validation.
A run's log also states how many
apify-actor-startevents it was pre-charged and what they cost, and warns when the requested memory is one the platform would refuse - the memory is still applied as asked.Two local differences: Apify Proxy group availability is not checked, and encrypted secret input fields stay unsupported.