Skip to content

Bootstrap prod Azure inference + deploy the 3 models at a high rate limit #34

Description

@Kenny-Heitritter

Question

In qbraid-infrastructure terraform/environments/prod/azure/ (today foundation-only: RG qbraid-prod + Log Analytics, no cognitive accounts), bootstrap a prod AI Foundry cognitive account (qbraid-prod-ai-services) mirroring staging's main.tf / monitoring.tf / outputs.tf, and deploy Sol/Terra/Luna with a much higher sku_capacity than staging.

Decide the prod rate-limit target per model, bounded by the quota ceiling from the discovery ticket. Apply through the (gated) prod CI pipeline.

Not user-facing: the gateway will not route prod traffic to these deployments in this effort — this ticket provisions prod, it does not expose it.

Metadata

Metadata

Assignees

No one assigned

    Labels

    wayfinder:taskWayfinder ticket: manual unblocking task

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions