Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 21 additions & 6 deletions RECOVERY.md
Original file line number Diff line number Diff line change
Expand Up @@ -71,14 +71,14 @@ run to `succeeded` over a GPU that was never touched.
| `blocked` | Approved, but another recovery is already running on this node. The approval is kept, so it resumes on its own |
| `draining` | Node is being drained before a reset |
| `in-progress` | Recovery Job is running |
| `succeeded` / `failed` | Job finished. A failure inside `spec.maxRetries` is re-queued for approval |
| `succeeded` / `failed` | Job finished. `failed` is terminal: only an approval naming the event starts another attempt |

`status.events[].stateMessage` explains any state that is not self-explanatory (which recovery holds
the node, which image cannot be pulled, what a stalled drain was waiting on). An empty value means
there is nothing to add.

The plan-level `status.state` is `idle`, `active`, or `error`. `error` means an admin is needed: an
event out of retries, or one blocked on missing firmware.
The plan-level `status.state` is `idle`, `active`, or `error`. `error` means an admin is needed: a
failed event, or one blocked on missing firmware.

## Safety mechanisms

Expand All @@ -104,8 +104,13 @@ event out of retries, or one blocked on missing firmware.
against their registries before a Job is created, so a mistyped reference does not produce a Job
that reports `in-progress` from `ImagePullBackOff`. Checked once per spec generation; disable with
`spec.skipImageVerification` where nodes hold pull credentials the operator cannot see.
* **Retry budget.** `spec.maxRetries` bounds automatic retries; an exhausted event needs an explicit
per-event re-approval, so a standing group approval cannot loop a dying card forever.
* **No automatic retries.** A failed event is terminal. A reset that did not bring the card back is
unlikely to on an identical second run.
* **One pod per approval.** Recovery Jobs are created with `backoffLimit: 0`, so a pod that exits
non-zero — in `xpu-smi` or in the reflash's firmware-copy init container — fails the Job.
* **A verdict is always reached.** An event stays `in-progress` only while its Job can still
conclude. Past the Job's deadline plus a minute, a Job that has reported nothing — or that has been
deleted from under the operator — fails the event.
* **Finalizer.** Deleting a plan blocks until every recovery Job is terminal, and releases all drain
taints first.
* **The operator tolerates its own drain taint**, since it is the only thing that removes it.
Expand All @@ -120,7 +125,6 @@ metadata:
spec:
deviceId: "0xe20b" # mandatory; one plan per GPU model
defaultResetType: "slot" # mandatory: "slot" or "amc"
maxRetries: 3

drain:
enable: true
Expand Down Expand Up @@ -177,6 +181,17 @@ spec:
The operator fills in a missing `id`, and marks non-persistent approvals `consumed: true` once acted
upon; consumed entries stay as an audit trail and can be removed by hand.

A `consumed` approval authorises nothing further, including another attempt at the event it started.
Restarting a failed event therefore means a new entry naming it, which is also the record of who
decided to try again:

```yaml
spec:
approvals:
- id: second-look-evt-node03-reflash-0000-4b-00-0
eventId: evt-node03-reflash-0000-4b-00-0
```

### kubectl plugin

`kubectl-gpurecovery` wraps the patching, with tab completion for plan names, event IDs and approval
Expand Down
17 changes: 2 additions & 15 deletions api/v1alpha1/gpurecoveryplan_types.go
Original file line number Diff line number Diff line change
Expand Up @@ -124,15 +124,6 @@ type GPURecoveryPlanSpec struct {
// needs (spec.xpuSmi.image, and the firmware image for a reflash).
// +optional
SkipImageVerification bool `json:"skipImageVerification,omitempty"`

// MaxRetries is the maximum number of times a failed recovery event is automatically
// re-queued for approval and retried while its device taint persists. Once this limit
// is reached the event stays in the failed state and requires manual intervention
// (e.g. delete the event entry or increase MaxRetries). Setting 0 disables automatic
// retries entirely.
// +kubebuilder:default=3
// +kubebuilder:validation:Minimum=0
MaxRetries int32 `json:"maxRetries"`
}

// DrainSpec configures the node drain a reset-type recovery performs before it touches the
Expand Down Expand Up @@ -392,15 +383,11 @@ type RecoveryEvent struct {
// +optional
JobName string `json:"jobName,omitempty"`

// PastJobs is the names of all Jobs created for this event across all attempts.
// Jobs are retained alive until the event is removed (i.e. the device taint clears),
// so their Pods remain available for diagnostics throughout the event lifecycle.
// PastJobs is the names of all Jobs created for this event across all attempts, so its length
// is the number of attempts made so far.
// +optional
PastJobs []string `json:"pastJobs,omitempty"`

// RetryCount is the number of times this recovery has been retried after failure.
RetryCount int32 `json:"retryCount"`

// LastUpdated is the timestamp of the most recent state change for this event.
LastUpdated *metav1.Time `json:"lastUpdated"`
}
Expand Down
23 changes: 2 additions & 21 deletions charts/gpu-base-operator/crds/gpurecoveryplans.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -242,17 +242,6 @@ spec:
- file
- source
type: object
maxRetries:
default: 3
description: |-
MaxRetries is the maximum number of times a failed recovery event is automatically
re-queued for approval and retried while its device taint persists. Once this limit
is reached the event stays in the failed state and requires manual intervention
(e.g. delete the event entry or increase MaxRetries). Setting 0 disables automatic
retries entirely.
format: int32
minimum: 0
type: integer
skipImageVerification:
description: |-
SkipImageVerification disables the pre-flight registry check on the images a recovery Job
Expand Down Expand Up @@ -361,7 +350,6 @@ spec:
required:
- defaultResetType
- deviceId
- maxRetries
type: object
status:
description: GPURecoveryPlanStatus defines the observed state of GPURecoveryPlan.
Expand Down Expand Up @@ -429,9 +417,8 @@ spec:
type: string
pastJobs:
description: |-
PastJobs is the names of all Jobs created for this event across all attempts.
Jobs are retained alive until the event is removed (i.e. the device taint clears),
so their Pods remain available for diagnostics throughout the event lifecycle.
PastJobs is the names of all Jobs created for this event across all attempts, so its length
is the number of attempts made so far.
items:
type: string
type: array
Expand Down Expand Up @@ -486,11 +473,6 @@ spec:
required:
- type
type: object
retryCount:
description: RetryCount is the number of times this recovery
has been retried after failure.
format: int32
type: integer
state:
allOf:
- enum:
Expand Down Expand Up @@ -525,7 +507,6 @@ spec:
- lastUpdated
- nodeName
- recoveryType
- retryCount
- state
type: object
type: array
Expand Down
23 changes: 2 additions & 21 deletions config/crd/bases/intel.com_gpurecoveryplans.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -242,17 +242,6 @@ spec:
- file
- source
type: object
maxRetries:
default: 3
description: |-
MaxRetries is the maximum number of times a failed recovery event is automatically
re-queued for approval and retried while its device taint persists. Once this limit
is reached the event stays in the failed state and requires manual intervention
(e.g. delete the event entry or increase MaxRetries). Setting 0 disables automatic
retries entirely.
format: int32
minimum: 0
type: integer
skipImageVerification:
description: |-
SkipImageVerification disables the pre-flight registry check on the images a recovery Job
Expand Down Expand Up @@ -361,7 +350,6 @@ spec:
required:
- defaultResetType
- deviceId
- maxRetries
type: object
status:
description: GPURecoveryPlanStatus defines the observed state of GPURecoveryPlan.
Expand Down Expand Up @@ -429,9 +417,8 @@ spec:
type: string
pastJobs:
description: |-
PastJobs is the names of all Jobs created for this event across all attempts.
Jobs are retained alive until the event is removed (i.e. the device taint clears),
so their Pods remain available for diagnostics throughout the event lifecycle.
PastJobs is the names of all Jobs created for this event across all attempts, so its length
is the number of attempts made so far.
items:
type: string
type: array
Expand Down Expand Up @@ -486,11 +473,6 @@ spec:
required:
- type
type: object
retryCount:
description: RetryCount is the number of times this recovery
has been retried after failure.
format: int32
type: integer
state:
allOf:
- enum:
Expand Down Expand Up @@ -525,7 +507,6 @@ spec:
- lastUpdated
- nodeName
- recoveryType
- retryCount
- state
type: object
type: array
Expand Down
4 changes: 4 additions & 0 deletions internal/controller/gpurecoveryplan_const.go
Original file line number Diff line number Diff line change
Expand Up @@ -96,6 +96,10 @@ const (
defaultResetJobTimeout int64 = 300
defaultReflashJobTimeout int64 = 600

// recoveryJobVerdictGrace is how long past a recovery Job's own activeDeadlineSeconds the
// operator waits for the Job controller's verdict before failing the event itself.
recoveryJobVerdictGrace = 60 * time.Second

// maxRecoveryNameLen is the hard ceiling on a recovery Job name, and therefore on the
// event ID it is built from.
//
Expand Down
Loading