Decide when to retry HTTP failures, respect Retry-After, and classify provider errors in TypeScript and Python workers.
Which HTTP errors should be retried?
A retry is another execution of an operation, not just another network packet. Before retrying a failed payment, email or write, establish whether the destination may already have applied it. Use a stable destination idempotency key, then decide whether another attempt can change the outcome.
GetRatchet does not interpret HTTP status codes on your behalf. The handler or its local classifier makes that decision. The server receives a bounded failure message, a retryable flag and, optionally, a minimum delay. It does not need the provider response body, exception stack or credentials.
Start with the provider's contract
| HTTP response | Typical decision | What to check first |
|---|---|---|
| 400 Bad Request | Stop and correct the input. | Some APIs use 400 for multiple error codes; inspect a documented machine-readable code locally. |
| 401 Unauthorized | Refresh credentials if supported, then retry within a bound. | Repeating the same expired credential is not recovery. Keep refresh logic in your worker. |
| 403 Forbidden | Usually stop. | Permissions, account restrictions and policy errors normally need intervention. |
| 408 Request Timeout | Often retryable. | Determine whether the destination processed any part of the operation. |
| 409 Conflict | Depends on the operation. | A version conflict may need a fresh read; a duplicate-operation response may represent prior success. Do not repeat unchanged input indefinitely. |
| 425 Too Early | Retry only according to the transport/provider contract. | This can concern replay risk with early data. A fresh request must avoid the unsafe transport condition. |
| 429 Too Many Requests | Usually retryable after waiting. | Respect Retry-After and consider shared account limits across all workers. |
| 500 Internal Server Error | Sometimes retryable. | It may be a temporary fault or a reproducible provider bug. Bound attempts and preserve idempotency. |
| 502 Bad Gateway | Often retryable. | The upstream may have completed the operation before its response was lost. |
| 503 Service Unavailable | Usually retryable. | Honor Retry-After when present and allow the circuit breaker to limit further traffic. |
| 504 Gateway Timeout | Often retryable. | A timeout does not prove that the upstream did nothing. |
This table is a starting point, not a universal policy. Provider error codes, operation semantics and authentication state can override the usual decision. A successful HTTP response can also contain an application-level failure; inspect the provider's documented response contract.
Respect Retry-After without retrying early
Retry-After can contain an integer number of seconds or an HTTP date. Convert it to a duration when handling the response. GetRatchet accepts a minimum delay from 1,000 to 86,400,000 milliseconds. If a provider asks for more than 24 hours, stop automatic attempts and arrange a later reviewed recovery; clamping that request down to 24 hours would retry too early.
These examples assume the HTTP request was made for an operation safe to retry. They do not forward raw response bodies. Unknown exceptions remain retryable, matching SDK behavior. Keep the producer and worker keys separate.
TypeScript worker classifier
import { createRatchet, RetryableError, NonRetryableError } from '@getratchet/sdk';
class ProviderFailure extends Error {
constructor(readonly status: number, readonly retryAfter: string | null) {
super('Provider request failed'); // no response body or credentials
}
}
function providerDelay(value: string | null): number | undefined {
if (!value) return undefined;
const raw = value.trim();
const ms = /^\d+$/.test(raw)
? Number(raw) * 1000 : Date.parse(raw) - Date.now();
return Number.isFinite(ms) ? Math.max(1000, Math.ceil(ms)) : undefined;
}
const worker = createRatchet({
baseUrl: 'https://getratchet.app',
apiKey: process.env.GETRATCHET_WORKER_KEY!,
});
worker.registerTool({
name: 'check_provider', version: '1',
async handler(_input, context) {
const response = await fetch(process.env.PROVIDER_HEALTH_URL!, {
signal: context.signal,
});
if (!response.ok) {
throw new ProviderFailure(response.status, response.headers.get('Retry-After'));
}
return { available: true };
},
classifyError(error) {
if (error instanceof ProviderFailure) {
if ([400, 401, 403].includes(error.status)) return { retryable: false };
if (error.status === 429 || error.status === 503) {
const retryAfterMs = providerDelay(error.retryAfter);
if (retryAfterMs !== undefined && retryAfterMs > 86_400_000) {
return { retryable: false }; // operator review, no premature retry
}
return { retryable: true, retryAfterMs };
}
}
return { retryable: true };
},
});
await worker.worker.start({ concurrency: 2 });
// In a handler you can instead throw either public SDK error directly:
// throw new RetryableError('Provider busy', { retryAfterMs: 60_000 });
// throw new NonRetryableError('Input needs correction');This example deliberately stops on 401 because it does not implement credential refresh. Add explicit provider-specific handling for 409 or 425 if your operation needs it; the fallback is not proof that every status is safe.
Python worker classifier
import math
import os
import time
from email.utils import parsedate_to_datetime
from urllib.error import HTTPError
from urllib.request import urlopen
from getratchet import Ratchet, RetryableError, NonRetryableError
def provider_delay(value):
if not value:
return None
raw = value.strip()
try:
seconds = (int(raw) if raw.isascii() and raw.isdigit()
else parsedate_to_datetime(raw).timestamp() - time.time())
return max(1000, math.ceil(seconds * 1000)) if math.isfinite(seconds) else None
except (ValueError, TypeError, OverflowError):
return None
def classify(error, context):
if isinstance(error, HTTPError):
if error.code in (400, 401, 403):
return {"retryable": False}
if error.code in (429, 503):
delay = provider_delay(error.headers.get("Retry-After"))
if delay is not None and delay > 86400000:
return {"retryable": False}
return {"retryable": True, **({"retry_after_ms": delay} if delay is not None else {})}
return {"retryable": True}
def check_provider(input, context):
context.check_cancelled()
with urlopen(os.environ["PROVIDER_HEALTH_URL"], timeout=10) as response:
return {"available": response.status == 200}
worker = Ratchet("https://getratchet.app", os.environ["GETRATCHET_WORKER_KEY"])
worker.register_tool("check_provider", "1", check_provider, classify_error=classify)
worker.worker.start(concurrency=2)
# A handler can also raise these directly:
# raise RetryableError("Provider busy", retry_after_ms=60000)
# raise NonRetryableError("Input needs correction")Python blocking HTTP calls need their own bounded timeout; cancellation is cooperative. Do not log the original HTTPError body or headers. In both SDKs, a classifier failure falls back to retryable behavior and can be reported to the worker's local on-error callback. Input/output schema failures and cancellation stay non-retryable regardless of the classifier.
Bound the damage
An error being retryable does not mean it will run immediately or forever. Maximum attempts include the initial execution. The next due time uses the greater of the jittered policy delay and the explicit minimum, within the 24-hour safety bound. Pauses, concurrency limits, rate limits and an open circuit can defer execution further. Synthetic tests remain limited to one attempt.
For the calculation, read exponential backoff with jitter. For side-effect safety, read destination-side idempotency. The worker operations reference covers error classes, schema validation and report fencing.