Cog containers are Docker containers that serve an HTTP server for running your model. You can deploy them anywhere that Docker containers run.
The server inside Cog containers is coglet, a Rust-based inference server that handles HTTP requests, worker process management, and run execution.
This guide assumes you have a model packaged with Cog. If you don't, follow our getting started guide, or start from one of the examples in the Cog repository.
Build your model into a Docker image:
cog build -t my-modelThe image contains your model code, dependencies, the Cog runtime, and everything in between. It serves an HTTP server on port 5000 when run.
cog push builds the project and publishes it to an OCI-compliant registry:
cog push registry.example.com/acme/my-model:v1The project configuration decides what Cog publishes. A project with image in
cog.yaml publishes a container image. A project with model publishes an OCI
bundle containing the image and any managed weights. Passing a target doesn't
switch between these formats.
The positional target is optional when image or model already supplies a
destination. When present, it overrides that configured destination and every
COG_MODEL* environment variable. The target must be tag-addressable -- Cog
rejects digest-pinned targets. An untagged image target uses Docker's latest
tag, while an untagged bundle target gets a timestamp tag generated by Cog.
For a bundle with managed weights, use the same repository where cog weights import published the weight manifests. The model tag may differ. If the
repository doesn't contain every required weight manifest, Cog fails before
pushing the image and tells you to run cog weights import for that repository.
Without --json, a successful push prints the published references as a
human-readable tree on stderr.
Use --json when another program needs the immutable references produced by a
push. For example, this pushes a bundle and prints its model, image, and managed
weight references:
cog push registry.example.com/acme/my-model:v1 --jsonVersion 1 output for a bundle with managed weights has this shape:
{
"version": 1,
"model": "registry.example.com/acme/my-model@sha256:1111111111111111111111111111111111111111111111111111111111111111",
"image": "registry.example.com/acme/my-model@sha256:2222222222222222222222222222222222222222222222222222222222222222",
"weights": [
{
"name": "transformer",
"reference": "registry.example.com/acme/my-model@sha256:3333333333333333333333333333333333333333333333333333333333333333"
}
]
}The version 1 fields are:
version: Always1.image: Always present. The value is the published image's complete, digest-pinned reference.model: Present only for bundle pushes. The value is the complete, digest-pinned bundle reference.weights: Present only when the bundle has managed weights. Each entry has anameand a complete, digest-pinnedreference. Entries keep theircog.yamlorder.
JSON whitespace and object key order aren't part of the contract.
Stdout stays empty while the command runs. Build and upload progress, warnings, diagnostics, and registry-provider messages go to stderr. After the push and provider post-processing both succeed, stdout receives exactly one JSON document. JSON mode suppresses the human-readable reference tree.
Any build, push, digest-resolution, provider, or serialization failure exits nonzero without writing a JSON document to stdout. JSON mode requires Cog to resolve every published reference to an immutable digest, so it can report failure even when an image was pushed successfully. Non-JSON image pushes keep their existing fallback when a registry can't resolve the pushed image's digest.
You have several options for running a built image.
Run the image directly with Docker. This is the approach you'd use for production deployment.
# If your model uses a CPU:
docker run -d -p 5001:5000 my-model
# If your model uses a GPU:
docker run -d -p 5001:5000 --gpus all my-modelThe server listens on port 5000 inside the container (mapped to 5001 above).
For local development, cog serve builds the image and starts the server
with your project directory mounted in:
cog serveBy default the server runs on port 8393 and the container port is published on
127.0.0.1 (localhost), so it is only reachable from your local machine. The
server process inside the container binds to 0.0.0.0; use --host to control
which host interface the Docker port mapping is published on.
Use -p to choose a different port:
cog serve -p 5000Once the server is running, make predictions by sending a POST request
to the /predictions endpoint.
Inputs go inside an "input" object in the JSON body.
Note
The examples below use localhost:5001, matching the Docker command above
(-p 5001:5000). If you used cog serve, use localhost:8393 by default,
or the port you passed with -p.
curl http://localhost:5001/predictions -X POST \
-H "Content-Type: application/json" \
-d '{"input": {"prompt": "a photo of a cat", "steps": 50}}'{
"status": "succeeded",
"output": "data:image/png;base64,...",
"metrics": {
"predict_time": 4.52
}
}Important
Inputs must be wrapped in an "input" object.
{"input": {"scale": 2.0}} is correct; {"scale": 2.0} is not.
To discover what inputs your model accepts, view the OpenAPI schema:
curl http://localhost:5001/openapi.jsonFile inputs (cog.Path or cog.File types) are passed as strings
inside the "input" object.
There are two ways to do this:
1. HTTP/HTTPS URLs
Pass a URL to a publicly accessible file. The server downloads it inside the container:
curl http://localhost:5001/predictions -X POST \
-H "Content-Type: application/json" \
-d '{"input": {"image": "https://example.com/photo.jpg"}}'2. Data URLs (base64)
To pass a local file, encode it as a data URL:
# Construct a data URL from a local file
DATA_URL=$(python3 -c "
import base64, mimetypes
with open('input.jpg', 'rb') as f:
data = base64.b64encode(f.read()).decode()
mime = mimetypes.guess_type('input.jpg')[0] or 'application/octet-stream'
print(f'data:{mime};base64,{data}')
")
curl http://localhost:5001/predictions -X POST \
-H "Content-Type: application/json" \
-d "{\"input\": {\"image\": \"$DATA_URL\"}}"Note
The HTTP API only accepts JSON (application/json).
Multipart form uploads are not supported.
When you use cog run -i image=@photo.jpg,
the CLI handles the base64 encoding for you automatically.
When a model returns a file output (cog.Path or cog.File),
the response contains a base64-encoded data URL by default:
{
"status": "succeeded",
"output": "data:image/png;base64,iVBORw0KGgo..."
}To have the server upload output files to external storage instead,
start the server with the --upload-url flag. The server then uploads each
file output to that URL prefix and returns the resulting URL in the response.
With cog serve:
cog serve --upload-url https://example.com/upload/When running the image directly with Docker, override the command to start the
server with --upload-url:
docker run -d -p 5001:5000 my-model \
python -m cog.server.http --upload-url https://example.com/upload/With an upload URL configured, file outputs are uploaded and the response contains the uploaded URL instead of a data URL:
{
"status": "succeeded",
"output": "https://example.com/upload/image.png"
}The Docker and cog serve commands above leave an HTTP server running.
If you instead want to run a single prediction and exit — without starting a
server — use cog run:
cog run my-model -i image=@input.jpgThis starts the container, runs one prediction, prints the result, and exits.
File inputs are passed with the @ prefix (e.g. -i image=@photo.jpg),
and the CLI handles base64 encoding for you.
The server exposes a GET /health-check endpoint that returns the current status of the model container. Use this for readiness probes in orchestration systems like Kubernetes.
curl http://localhost:5001/health-checkThe response includes a status field with values like STARTING, READY, BUSY, SETUP_FAILED, or DEFUNCT. See the HTTP API reference for full details.
If you started the container with docker run -d, stop it with:
docker kill <container-id>If you used cog serve, press Ctrl+C in the terminal.
(cog run exits on its own once the prediction finishes, so there's nothing
to stop.)
By default, the server processes one run at a time. To enable concurrent runs, make your run() method async and decorate it with @cog.concurrent(max=N):
import cog
class Runner(cog.BaseRunner):
@cog.concurrent(max=4)
async def run(self) -> str:
return "hello world"The deprecated concurrency.max field in cog.yaml is still supported and takes precedence over the decorator by baking COG_MAX_CONCURRENCY into the image.
You can configure runtime behavior with environment variables:
COG_SETUP_TIMEOUT: Maximum time in seconds for thesetup()method (default: no timeout).COG_MAX_CONCURRENCY: Number of concurrent prediction slots (default: 1). Overrides both@cog.concurrentand deprecatedcog.yamlconcurrency.
See the environment variables reference for the full list.
- HTTP API reference for full endpoint documentation
- Private registries for using private Python package registries
cog.yamlreference for configuration options