Understand retry schedules, failure thresholds, cooldown and half-open probes, and how circuit gates affect failed workflow recovery.
Circuit breaker versus retry policy
A retry policy answers: “When may this job try again, and how many executions are allowed?” A circuit breaker answers: “Should this endpoint admit more work while it appears unhealthy?” They act at different levels and are most useful together.
Changing a retry delay does not repair a provider outage. Opening a circuit does not erase jobs or automatically make their operations safe to repeat.
Per-job timing and per-endpoint admission
| Control | Scope | Controls | Does not guarantee |
|---|---|---|---|
| Retry policy | Job snapshot | Attempt ceiling, exponential delay, cap and jitter | Provider recovery or unique side effects |
| Retry-after hint | Individual failure | Earliest acceptable next attempt | Exact dispatch time |
| Circuit breaker | Endpoint | Admission after repeated failures and recovery probes | That every failure was transient |
| Endpoint pause | Endpoint | Operator-controlled suspension of claims | Cancellation of remote effects |
| Concurrency and rate limits | Endpoint | Available execution capacity | That a due job starts immediately |
GetRatchet snapshots the resolved retry policy when enqueueing. A selected tool contract overrides the endpoint policy, and an explicit enqueue maximum overrides its default attempt count. Later policy edits do not rewrite that job's scheduled due time. Circuit state remains an admission gate at claim time; a custom policy cannot bypass it.
Closed, open and half-open
The current implementation opens the endpoint circuit after five consecutive failures and uses a 60-second cooldown. These are implementation constants, not customer-configurable retry-policy fields.
While closed, eligible jobs can be claimed subject to the other gates. When open, claims wait for the cooldown. A due eligible job can then act as the half-open probe. A successful normal attempt closes the circuit and resets the failure count. A failed probe reopens it and starts another cooldown. Probe control prevents treating cooldown expiry as permission for every waiting job to hammer the provider simultaneously.
The cooldown is not a timer that guarantees a request at exactly 60 seconds. A compatible worker must poll, a job must be due, and endpoint pause, rate and concurrency gates still apply. A lost probe lease is handled by the durable claim/recovery logic, rather than accepting stale reports as proof of health.
A concrete outage timeline
Suppose five invoice jobs fail against one unavailable payment endpoint. Those failures open its circuit. Each retryable job retains its own next due time according to its snapshotted policy. A job due in ten seconds is still blocked by the circuit; short custom backoff does not override endpoint health.
After cooldown, an eligible worker may claim a probe. If that succeeds, ordinary claims can resume when their other conditions are met. Jobs whose attempts already exhausted do not spontaneously become pending again. An operator must inspect and deliberately recover them, taking account of potential duplicate effects.
If a provider sends Retry-After: 120, that job's minimum delay remains two minutes even if another job recovers the circuit sooner. Conversely, an open circuit can hold it longer than two minutes. The selected due time is a lower bound, not a reservation of execution capacity.
Retryability, cancellation and exhaustion
NonRetryableError stops automatic attempts for that job; it is not a bypass for the endpoint's failure accounting. A normal failed attempt can contribute to the circuit even when the handler decides another attempt would be pointless. Input and output schema failures remain non-retryable. Unknown handler errors remain retryable by default, so classify documented provider errors locally when that default is inappropriate.
Cancellation prevents further automatic attempts and does not increment the normal circuit failure counter. Handlers must cooperate with cancellation; remote side effects may already have happened. Reaching the maximum attempt count exhausts a job independently of whether the circuit later recovers. Synthetic tests remain one-attempt checks and do not repair an open production circuit.
Choose the right recovery action
If the destination is still unavailable, let the cooldown and bounded policy control traffic. If credentials or input are invalid, fix the underlying condition before replay. If the worker is absent or registered under another version, repair deployment compatibility instead of changing the retry policy. If the result is uncertain, reconcile with the destination before issuing another irreversible action.
Read HTTP-specific error decisions, the exact jitter calculation, and idempotency for external effects. The durable worker reference covers leases, report fencing and reviewed replay in more detail.