Worker leases, retries, versions, and safe replay.
Durable worker operations
Run registerTool({ name, version, handler }) in a separate long-lived Node process, then await ratchet.worker.start({ concurrency, signal }). Deploy at least one worker for each exact name/version still present in queued or retry-scheduled jobs. Keep the old version running until its jobs finish; registering a newer version does not silently migrate them.
The producer calls enqueue with a stable idempotency key, then can request run completion. The run remains RUNNING while required steps are pending and becomes SUCCEEDED after the last one succeeds. An exhausted or cancelled step makes the run FAILED. A failed run does not automatically cancel other in-flight steps; inspect them before retrying.
The queue is PostgreSQL. A claim atomically locks a due job and its endpoint, increments the attempt and lease token, and creates an execution row. Heartbeats extend the lease. If a worker vanishes, another worker with the same handler version can reclaim the job after lease expiry. The previous attempt is marked failed with a warning that an external side effect may have completed. A stale worker report receives HTTP 409. A worker should stop normal polling on shutdown and let active calls finish; worker.stop() does that.
By default, failures use STANDARD exponential backoff with ±20% jitter, starting at 30 seconds and capped at 15 minutes before jitter. Custom policies and per-error decisions are described below. After the configured maximum attempts, the job is EXHAUSTED. An input or output schema rejection is non-retryable. A tool timeout aborts the handler's AbortSignal; handlers must propagate it to their own HTTP calls. A handler that ignores it can continue running after GetRatchet records a timeout, so destination-side idempotency remains essential.
If queue age rises, check that a worker is online, its key works, and it registers the exact queued handler versions. If a circuit is open, inspect recent attempts and the external service before resuming traffic. A paused service holds queued jobs without consuming attempts; existing in-flight calls may finish. If database access fails, producers cannot enqueue and workers cannot claim. Do not run tools outside GetRatchet to work around that failure unless the caller deliberately accepts losing the execution record.
Do not log input, output, credentials or full error bodies from the worker. GetRatchet stores encrypted recovery input and result data; the dashboard shows redacted previews. Keep RESULT_ENCRYPTION_KEY available on every API and worker deployment. Follow the rotation procedure before changing it.
Replaying a failed step
An administrator can preview GET /api/v1/steps/:id/replay and then request POST /api/v1/steps/:id/replay with {"acknowledgeDuplicateRisk":true,"maxAttempts":1}. The server locks the durable job and accepts this only after it is exhausted or cancelled. By default, replay keeps its encrypted original input and exact handler version, increments the attempt number, and writes a replay event with the actor key ID. Confirm destination-side idempotency before replay: the previous worker may have completed an external side effect even when its result was not recorded. A paused service holds the replay until resumed.
To use a different handler, first publish a new contract for the same tool name and inspect the replay preview's upgrades list. The preview validates the encrypted saved input against each target input schema without returning the raw input. An incompatible or FORBIDDEN target cannot be selected. After reviewing schema and side-effect changes, submit targetContractVersion, expectedHandlerVersion from the preview, and acknowledgeVersionChange: true along with the normal replay fields. The server rechecks the target and saved input under the job lock, updates the step's contract and handler version and snapshots the target contract's timeout and retry schedule for the next attempt, and records both old and new versions in the replay event. Previous executions remain unchanged. The target worker must register the new exact handler version; otherwise the job stays queued. Schema compatibility only proves the saved input has the expected shape; the handler may still behave differently. Check destination-side idempotency and the target handler's effects before approving the replay.
Worker polling defaults to every 30 seconds. Shorter intervals increase database operations and Vercel function invocations; budget capacity before lowering this value. A worker needs a separate WORKER key authorized for every exact name@version it registers. A producer needs an INGEST key. The worker ID remains bound to its registering key.
Service policies are read at GET /api/v1/endpoints/:id/policy and changed by an administrator with PUT and an expectedVersion. Presets are STANDARD (five attempts, 30-second initial backoff), AGGRESSIVE (eight attempts, five-second initial backoff), and RELAXED (three attempts, two-minute initial backoff). A service also has a default timeout, maximum concurrent leased jobs, and optional attempts-per-minute limit. Durable steps store the policy version; jobs snapshot timeout, maximum attempts and retry preset. Policy updates do not change those values or previously scheduled due times. Concurrency and rate gates apply to new claims immediately. Paused services accept new jobs but do not claim them.
Custom durable retry policies
| Preset | Total attempts | Initial delay | Multiplier | Cap before jitter | Jitter |
|---|---|---|---|---|---|
| STANDARD | 5 | 30 seconds | 2 | 15 minutes | ±20% |
| AGGRESSIVE | 8 | 5 seconds | 2 | 5 minutes | ±20% |
| RELAXED | 3 | 2 minutes | 2 | 30 minutes | ±20% |
| CUSTOM | 1–10 | 1,000–1,800,000 ms | 1–4 | initial delay–86,400,000 ms | ±0–50% |
The attempt count includes the first execution. Attempts and millisecond fields are integers; multiplier and jitter may be fractional. Unknown fields, incomplete CUSTOM policies, and custom fields attached to built-in presets are rejected. maxDelayMs must be at least initialDelayMs.
Update an endpoint with the following body. Policy reads return policyPreset and retryPolicy, which is the complete CUSTOM object or null for built-in presets.
{
"expectedVersion": 1,
"preset": "CUSTOM",
"maxAttempts": 6,
"initialDelayMs": 10000,
"multiplier": 2,
"maxDelayMs": 600000,
"jitterPercent": 20,
"defaultTimeoutMs": 30000,
"maxConcurrent": 4,
"rateLimitPerMinute": null
}To publish a custom tool contract, supply retryPreset: "CUSTOM" and retryPolicy: { preset: "CUSTOM", maxAttempts: 6, initialDelayMs: 10000, multiplier: 2, maxDelayMs: 600000, jitterPercent: 20 } alongside its existing contract fields. Other presets must omit retryPolicy.
The selected tool contract overrides the endpoint policy, even when the contract uses STANDARD. Without a contract, the endpoint policy applies. An explicit enqueue maxAttempts overrides only the chosen policy's attempt count. Synthetic tests always have one attempt. Every new job snapshots all parameters and the effective attempt count; updating an endpoint or publishing another contract version never changes existing jobs or already scheduled dueAt. Legacy rows without a snapshot retain their old preset behavior. Explicit manual replay still has its existing separately approved additional-attempt budget and retains the job's delay policy.
After failure number n, compute min(maxDelayMs, initialDelayMs * multiplier ** (n - 1)), then apply a random factor from 1 - jitterPercent / 100 to 1 + jitterPercent / 100 and round to milliseconds. The cap is applied before jitter, preserving the built-in behavior. The final delay is never more than 24 hours. For six attempts with the example policy, nominal waits are 10, 20, 40, 80 and 160 seconds; execution time, pauses, open circuits, polling and concurrency/rate gates can make actual starts later.
Classify each handler failure locally
TypeScript exports RetryableError(message, { retryAfterMs? }) and NonRetryableError(message). Python exports RetryableError(message, retry_after_ms=...) and retains NonRetryableError(message). Ordinary unknown handler errors are retryable by default. Non-retryable failures exhaust the job immediately. Input/output schema failures and cancellation remain non-retryable, and successful handlers do not run the classifier.
For third-party error types, register classifyError(error, context) in TypeScript or classify_error(error, context) in Python. It returns { retryable: true, retryAfterMs?: number } / { retryable: false } in TypeScript, or {"retryable": True, "retry_after_ms": ...} / {"retryable": False} in Python. Explicit SDK errors and internal schema/cancellation failures take priority over the classifier. Classifier exceptions or invalid decisions fall back to retryable and reach the worker's local onError/on_error callback when provided. Do not log raw diagnostics containing provider credentials or payloads.
Only a bounded, redacted message and the whitelisted decision fields reach the report API. Error objects, stacks and classifier metadata are not serialized. The server never guesses retryability from error text or HTTP status codes.
HTTP provider example (TypeScript)
Inspect the provider's structured status locally; do not pass its response body into an error message.
import { RetryableError, NonRetryableError } from '@getratchet/sdk';
function retryAfterMs(value: string | null): number | undefined {
if (!value) return undefined;
const ms = /^\d+(\.\d+)?$/.test(value.trim())
? Number(value) * 1000 : Date.parse(value) - Date.now();
if (!Number.isFinite(ms)) return undefined;
if (ms > 86_400_000) throw new NonRetryableError('Provider requested a wait beyond 24 hours; review manually');
return Math.max(1000, Math.ceil(ms));
}
worker.registerTool({
name: 'send_welcome_email', version: '1',
async handler(input, context) {
const response = await fetch(providerUrl, {
method: 'POST', signal: context.signal,
headers: { 'Content-Type': 'application/json', 'Idempotency-Key': context.idempotencyKey },
body: JSON.stringify(input),
});
if (response.status === 429) throw new RetryableError('Provider rate limited', { retryAfterMs: retryAfterMs(response.headers.get('Retry-After')) });
if (response.status === 503) throw new RetryableError('Provider unavailable');
if (response.status === 400) throw new NonRetryableError('Provider rejected validation');
if (!response.ok) throw new Error('Provider request failed'); // unknown => retryable
return response.json();
},
});A third-party library can instead use a local classifier:
classifyError(error, context) {
// ProviderHttpError is your provider library's structured error type.
if (error instanceof ProviderHttpError) {
if (error.status === 400) return { retryable: false };
if (error.status === 429) {
try { return { retryable: true, retryAfterMs: retryAfterMs(error.headers.get('Retry-After')) }; }
catch (failure) {
if (failure instanceof NonRetryableError) return { retryable: false };
throw failure;
}
}
if (error.status === 503) return { retryable: true };
}
return { retryable: true };
}If a provider may request more than 24 hours, have the classifier explicitly return { retryable: false } for that case; throwing from a classifier intentionally falls back to retryable. The handler example above uses NonRetryableError directly.
HTTP provider example (Python)
import time
from email.utils import parsedate_to_datetime
from urllib.error import HTTPError
from getratchet import RetryableError, NonRetryableError
def provider_delay(value):
if not value:
return None
try:
seconds = float(value)
except ValueError:
try:
seconds = parsedate_to_datetime(value).timestamp() - time.time()
except (ValueError, TypeError, OverflowError):
return None
import math
if not math.isfinite(seconds):
return None
return max(1000, math.ceil(seconds * 1000))
def classify_provider_error(error, context):
if isinstance(error, HTTPError):
if error.code == 400:
return {"retryable": False}
if error.code == 429:
delay = provider_delay(error.headers.get('Retry-After'))
if delay is not None and delay > 86400000:
return {"retryable": False} # manual review, no early retry
return {"retryable": True, **({"retry_after_ms": delay} if delay is not None else {})}
if error.code == 503:
return {"retryable": True}
return {"retryable": True}
worker.register_tool('send_welcome_email', '1', send_welcome,
classify_error=classify_provider_error)
# Or raise RetryableError('Provider busy', retry_after_ms=60000)
# or NonRetryableError('Invalid recipient') directly from send_welcome.retryAfterMs on a FAILED report must be an integer from 1,000 through 86,400,000, and is forbidden with retryable: false or a successful result. Its absence preserves normal policy scheduling. The stored next due time uses now + max(jitteredPolicyDelay, retryAfterMs), limited to 24 hours. No jitter is applied to the explicit minimum. Run inspection exposes steps[].durableJob.dueAt and the policy snapshot. Gates can defer execution further; retry hints never bypass an open circuit, a pause, or an attempt limit.
Durable execution remains at least once. Even a non-retryable error cannot undo a side effect that happened before a crash or lost report. Keep a stable idempotency key across retries and pass it to the destination service; destination-side deduplication is required for irreversible writes, payments and emails.
Automatically generated idempotency keys
When omitted, the durable key hashes the run ID, tool name, handler version, selected contract version (or null), effective endpoint (defaults to the tool name), dependency step/status (or null), and input. Nested JSON object keys are sorted; array order is significant. Equivalent operations in one SDK therefore reuse a key regardless of object insertion order. TypeScript process-bound wrappers also include the effective endpoint. Explicit customer keys are never changed.
TypeScript and Python use the same operation fields and canonical JSON ordering, but native numeric serialization can differ between languages (for example Python 1.0 versus JavaScript 1). Use an explicit shared key when producers in different languages must deduplicate the same business operation. Generated keys changed in this release; carry forward explicit keys for operations queued by older SDKs. Selecting an implicit contract does not pin future publications: supply a contract version for a fixed operation identity.
A same-version replay keeps its original retry-policy snapshot. A version upgrade uses the target contract's preset and complete retry-policy snapshot. Replay maxAttempts remains the number of additional executions; the job's lifetime ceiling includes prior attempts. The upgraded policy's attempt-count metadata is that ceiling capped at the public policy limit of 10; the job's maxAttempts remains authoritative for replay exhaustion. Previous attempt history is retained and backoff uses the lifetime attempt number.