Skip to content

Out-of-sample queries on fitted and loaded models (newdata across predict/do/ate, compile caching) #511

Description

@daimon-pymclabs

In a production setting a fitted (or loaded, see #507) model is queried repeatedly on new data. Today that path is uneven:

  • predictions() / comparisons() / slopes() accept newdata=, but predict(), do(), ate() and friends do not. Out-of-sample predict() relies on the user calling pm.set_data() themselves.
  • Panel models with lag() / adstock() bake n_units / n_times into the scan graph, so scoring a new panel shape means building a new model.
  • Each query compiles PyTensor functions from scratch, so cold-start latency is high for a service answering many small requests.

Proposal

  • A consistent newdata= argument (business units, validated per Data contracts: validate new data against the training schema and add an explicit missing-data policy #512) across predict, do, ate/att/atu/cate, and prob.
  • A documented route for scoring new panel units / horizons (e.g. forecasting forward from the last observed state).
  • Cache compiled posterior-predictive / intervention functions on the model so repeated queries with the same shape skip recompilation, plus an explicit warmup() hook for services.
  • Optional: a thin serializable query/result layer (dict or JSON in, dict or JSON out) that a service wrapper can call directly.

Part of #509 (production readiness).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions