Run GetRatchet housekeeping independently on Railway.
Platform maintenance on Railway
The GetRatchet-owned process uses pnpm platform-worker, Dockerfile.platform-worker and railway.platform.json. It is separate from customer workers. It never registers handlers, claims DurableJob records or executes customer code. Recovery only re-queues eligible existing jobs; execution remains customer-owned.
Deploy the repository root as a Railway service. For new services, configure the Dockerfile builder with /Dockerfile.platform-worker, healthcheck /health, and restart-on-failure with ten retries in Railway settings. Railway no longer accepts Config as Code for new services; railway.platform.json remains a reference for legacy services. Start with one replica. The container installs pinned dependencies and generates Prisma; it does not run migrations or build the web app. Apply migrations once through the controlled application release process before starting it.
Configure:
DATABASE_URL: the platform runtime database connection for this environment, with the same organization-scoped maintenance access as the API. Never use a customer API key or provider/tool credential.NEXT_PUBLIC_APP_URL: canonical application URL;RESEND_API_KEYandEMAIL_FROMfor operational notifications.BILLING_ENFORCEMENT_ENABLED: identical to the API. Stripe API secrets are not needed for deadline reconciliation; webhooks belong to the web process.PORT: Railway-injected health port (default 8080).MAINTENANCE_INTERVAL_MS: pause between pages, clamped to 1–60 seconds; default 5 seconds.- Preview deployments must use the existing
VERCEL_ENV=preview,PREVIEW_DATABASE_URLandVERCEL_URLisolation configuration instead of inheriting production credentials.
No customer tool credentials or payload decryption key are required. Retention does not decrypt payloads. Keep database migration credentials out of the running container.
Health: HTTP /health, success 200; unavailable during shutdown or after five minutes without a fully successful page. Restart policy: ON_FAILURE, at most ten retries. Railway service health checks validate deployment; configure external monitoring for continuing health. Logs contain fixed event names, timestamps and counts only.
Each iteration selects at most 25 organizations using an ID cursor. Each organization performs one bounded retention sweep (25 completed runs, gated once/hour), one recovery scheduling step, up to 20 new offline notifications, up to three pending notifications and billing deadline reconciliation. Expired rate-limit deletion is capped at 100 rows/page. Retention/recovery/reconciliation use database locks or conditional updates, and notifications have delivery leases/deduplication. As with any external email send, a crash after sending but before recording delivery may duplicate a notification.
SIGTERM/SIGINT stops new work, waits for the current bounded task, closes HTTP and disconnects the database. Allow sufficient Railway termination grace for an active database or email request. Do not run multiple replicas initially; existing transactional guards are the protection for overlapping restarts, not a reason to scale blindly.
Rollout without a housekeeping gap
Customer claim polls keep the existing idempotent fallback unless PLATFORM_MAINTENANCE_POLL_FALLBACK=false. Start the platform worker, verify its health/logs and maintenance progress without customer polling, then set that variable to false on the web app. Customer SDKs need no changes. Roll back the flag if maintenance is unavailable. Billing deadlines are also computed on reads and do not depend on polling.