Answers to common execution and recovery questions.
GetRatchet FAQ
Can GetRatchet run an agent for me? No. It records and manages an agent's tool calls. Your application runs the agent, and your worker runs registered durable handlers.
Does wrap() survive a process crash? No. It executes a local function in the current process. Durable enqueue() persists a named handler call for a separate worker to claim after a restart.
Are durable side effects exactly once? No. A lease can expire after a destination completed an action but before GetRatchet recorded success. Pass a stable idempotency key to the destination for payments, writes, and emails.
Why is my run still running after finishRun()? The finish request closes registration of new steps. The run waits for required queued, running, or retry-scheduled durable steps to settle.
Why is a job queued but unclaimed? Check an online worker, exact handler name and version, worker-key scope, project and environment, paused service, open circuit, and queue/database health. A newer handler version never silently replaces the queued version.
Can I replay a successful step or change its handler version? The current replay and recovery APIs accept exhausted or cancelled durable work only and preserve the exact saved handler version and input. A successful step is skipped. Publish a new contract and start new work for a new handler version.
What if an endpoint has no destination-side idempotency? Treat replay as potentially duplicating the side effect. Inspect destination records before approving it. A FORBIDDEN tool contract blocks GetRatchet replay.
Can I test a tool without changing customer data? Publish a SAFE contract backed by a safe test handler and call the synthetic test API. The SAFE classification is your assertion; test handlers must still avoid real side effects.
Who can approve AI recovery? An owner or admin must review the current dry-run plan and explicitly acknowledge duplicate risk. AI text never schedules work. External AI processing is off by default, and deterministic analysis works without a provider key.
What data goes to the optional AI provider? Only bounded, abstract statuses, outcome categories, durations, circuit state, and replay-safety classifications. No names, IDs, payloads, prompts, raw errors, or credentials are sent by default. See AI advisory.
When does old history disappear? With billing enabled, completed run retention is 7 days on Free, 30 on Basic, and 30–365 on Pro (90 by default). Downgrades wait 14 days before reducing retention. The platform worker performs bounded cleanup independently of customer polling. Unfinished work is retained. Replay and idempotency lookup for a deleted run end with its removal.
Can I export my history? Pro exports (also available during downgrade grace) require administrator access and are paginated and exclude payloads, raw errors, and secrets. Export before the retention sweep if you need older completed runs.
Where are my workers hosted? In your own long-lived process or container. Vercel hosts the GetRatchet API and console; it does not execute your registered tool code.
What happens if telemetry export is down? Traces and metrics may be lost, but export failures must not fail or block tool execution. Check collector configuration and retry the end-to-end trace verification when it recovers.
How can I restore a backup? Follow deployment and restore. Restore into an isolated non-production database first and verify the migration history and encrypted-data key before planning any production recovery.
Can I customize durable retries? Yes. STANDARD (5 attempts, 30s initial, 15m cap), AGGRESSIVE (8, 5s, 5m), and RELAXED (3, 120s, 30m) retain doubling backoff and ±20% jitter. CUSTOM lets you set the total attempts, initial delay, multiplier, cap and jitter. Attempts include the first execution. The tool contract overrides the endpoint; explicit enqueue maxAttempts overrides the count. Jobs snapshot the complete policy, so later edits never change existing jobs or their scheduled due times.
Can different errors have different retry decisions? Yes. Throw RetryableError or NonRetryableError in either SDK, or register a local classifyError (TypeScript) / classify_error (Python) for provider errors. Unknown errors remain retryable. Schema failures and cancellation are not retried; a classifier that throws falls back to retryable and reports its failure to the local error callback. The server does not parse error text or provider status codes to make decisions.
How does Retry-After work? A worker can send retryAfterMs (Python: retry_after_ms), from 1,000 through 86,400,000 ms. The scheduler uses the greater of that minimum and the policy's jittered delay, capped at 24 hours. It never jitters the explicit minimum. Circuit, pause and capacity gates still apply. See provider examples for 429, 503, non-retryable 400 and unknown-error fallback. This remains at-least-once execution: destination-side idempotency is necessary because a side effect can complete before its report is lost.
Where does the documentation assistant get its answers?
The read-only assistant searches the same published Markdown guides as the documentation site, including the SDK, retry, billing, deployment and CLI references. Each deployment refreshes that source. Relevant excerpts link to their original guide; English source excerpts are labeled in French sessions. It does not run commands, change configuration or send questions to an external model. Never paste credentials or tool payloads. Run context remains optional and permission checked.
See the HTTP retry guide, backoff guide, idempotency guide and CLI reference.