Skip to content
Get started
Get started
bunqueue in Production: What Survives a Crash
guide · production

What survives a crash.

Run the default engine as one process plus SQLite, or run multiple brokers against PostgreSQL 15–18. PostgreSQL 18.6 is the pinned and recommended release. This page explains crash behavior, monitoring, and the storage-specific checks before go-live.

quick answers

If it dies, what happens?

The short version first. These rows describe the default SQLite path; PostgreSQL differences follow.

FailureWhat happens to the job
Worker crashes mid-jobThe job is redelivered to another worker after its lease expires. It may run twice, never zero times.
Server killed hard (kill -9)Jobs written to disk survive and recover on restart. Jobs accepted in the last ~10 ms may be lost, unless they were added with durable: true.
Server restarted gracefully (SIGTERM, deploys)After shutdown completes successfully, buffered SQLite writes are flushed before exit. A flush error or forced termination follows the hard-kill row instead.
Job fails repeatedlyAfter its retry attempts are exhausted it moves to the DLQ (dead letter queue, a parking lot for failed jobs you can inspect and retry). See DLQ.
Disk fills upThe server stays up and /health reports degraded with HTTP 503. Durable pushes fail explicitly instead of pretending to succeed.
The whole machine is lostYou restore the last S3 snapshot. Your maximum data loss is the backup interval.
the crash timeline

The worker dies before the ACK.

This is the failure that defines a queue’s honesty. Here is the exact sequence.

  1. A worker pulls a job. The server marks it active and gives the worker a lease: a lock with a time-to-live (default 30 seconds, the worker’s lockDuration), renewed by heartbeats while the job runs.
  2. Your handler does its side effect. The email is sent, the charge is captured.
  3. The worker dies before its ACK reaches the server. An ACK is the worker’s “done” confirmation; without it the server has no idea the work happened.
  4. The lease expires and the server requeues the job with its attempt count incremented.
  5. Another worker pulls it. Your side effect runs again.

Once a job is durably admitted, the processing guarantee is at-least-once: it can run more than once, but the broker does not silently discard that generation. SQLite’s default write buffer precedes durable admission, so a hard crash can still lose jobs accepted inside its documented 10 ms window unless they use durable: true. No queue can atomically combine your external side effect with its own acknowledgment, so a crash can always land between them. What bunqueue guarantees around those edges:

  • A late ACK after ordinary lease expiry is accepted only while that exact generation still owns the processing slot and the job was not re-leased. If the job’s processing timeout already finalized that generation, the ACK/FAIL is instead an explicit idempotent no-op. Two generations can never both complete the same job.
  • Re-adding a job with the same custom jobId is a no-op, even under heavy concurrency. Producers that retry on timeout should always set one.
  • Duplicate ACKs are harmless.
durability

Two write modes, one decision.

By default bunqueue batches writes to SQLite every 10 ms for throughput. A durable write skips the application batch and commits before the add returns; host and physical-media durability remain operational concerns.

SQLite modePublished native workload medianBunqueue process-crash buffer window
Buffered (default)186,384 jobs/s, public on-disk addBulkup to 10 ms of accepted jobs
durable: true per job60,835 ops/s, sequential Embedded addsnone after add() resolves
await queue.add('send-newsletter', data); // buffered, fast
await queue.add('capture-payment', data, { durable: true }); // committed before this returns

Mixing is free: mark only the jobs that are money as durable and keep the rest batched. The 10 ms window applies to abrupt process termination; a successfully completed graceful shutdown flushes the buffer. Measured numbers per mode are in Benchmarks.

backup and restore

One file, snapshotted to S3.

In SQLite mode, the entire queue state is one database file, so disaster recovery is one object-storage snapshot. Works with AWS S3, Cloudflare R2, and MinIO; PostgreSQL uses database-native backup and PITR.

S3_BACKUP_ENABLED=1
S3_BUCKET=my-backups
S3_ACCESS_KEY_ID=...
S3_SECRET_ACCESS_KEY=...
S3_REGION=us-east-1 # or S3_ENDPOINT for R2/MinIO
S3_BACKUP_INTERVAL=21600000 # snapshot cadence, default 6 hours
S3_BACKUP_RETENTION=7 # snapshots kept, oldest pruned

This requires a persistent data path. Scheduled server backups flush the pending write buffer (and fail the cycle if storage backoff leaves anything pending), then use SQLite VACUUM INTO, so committed WAL data is captured even when a long-lived reader prevents checkpoint truncation.

bunqueue backup now # snapshot immediately
bunqueue backup list # list snapshots with size and age
bunqueue backup restore <key> --force # stop the server first
bunqueue backup status

Restore validates before it replaces: it verifies sizes and checksum, runs an integrity check on a temp copy, quarantines stale WAL/SHM sidecars, and only then swaps it in. A corrupt backup never touches your live database. Stop the server before restoring, size the interval to how much queue state you can afford to lose, and rehearse one restore before you need it. Provider setup lives in the S3 Backup guide.

monitoring

Three endpoints, no agent.

Everything is on the HTTP port (default 6790). Point Prometheus at it or just curl it.

EndpointWhat it gives you
GET /healthhealthy (200) or degraded (503), uptime, version, job counts, memory; only degraded responses include the storage object (with diskFull on SQLite)
GET /healthz, /liveBare liveness probes; remain 200 while the process responds
GET /readyReadiness; returns 503 when persistent storage is degraded
GET /prometheusMetrics in Prometheus text format: job counts, totals, latency, bounded per-queue gauges, process/connections, storage and backup freshness
GET /metricsThe same counters as JSON

Set METRICS_AUTH=true to require an AUTH_TOKENS bearer token on /prometheus; a missing token configuration fails closed with 503. Alert on these production symptoms:

  • DLQ growth: bunqueue_jobs_dlq climbing means handlers are failing terminally.
  • Queue depth: bunqueue_jobs_waiting + bunqueue_jobs_prioritized growing steadily means workers cannot keep up.
  • /health reporting degraded: inspect persistent-storage health. SQLite commonly reports a full disk; PostgreSQL may report connection, lifecycle, event-stream, projection-refresh, heartbeat, recovery, retention, or cron failures. /ready returns 503 until the affected subsystem recovers.
  • Backup freshness: an enabled scheduler that is stopped, failing, or has no success within two intervals puts recovery objectives at risk.
  • Telemetry truncation: bunqueue_queue_metrics_omitted > 0 means per-queue drill-down is incomplete; review cardinality before raising the cap.
  • Memory drift: RSS should plateau once the workload does; internal caches are hard-capped.

Dashboards and webhook alerting are covered in Monitoring.

postgresql failover

Know the recovery clock.

PostgreSQL uses the database clock for leases and broker ownership, so host clock skew cannot create two valid owners.

BUNQUEUE_POSTGRES_LEASE_DURATION_MS defaults to 30 seconds. It controls the background coordination cadence as follows:

OperationFormulaDefault
Broker heartbeatmax(1s, floor(leaseDurationMs / 3))10 s
Duplicate broker-ID stale takeovermax(leaseDurationMs, 3 × heartbeat interval)30 s
Expired processing-lease scanmin(15s, max(500ms, floor(leaseDurationMs / 2)))15 s

The processing lease itself is still determined by the worker’s lockDuration and renewed by heartbeats. After that lease expires, a surviving broker discovers it on the next recovery scan, so the normal recovery bound is the remaining job lease plus up to one scan interval, database/query scheduling not included. Do not lower these values merely to make failover look faster: measure database latency and pause behavior first, and alert when heartbeat or recovery health becomes degraded. See Storage backends.

postgresql upgrades

Treat the schema as a one-way boundary.

Schema initialization is automatic, but mixed bunqueue versions are not a supported steady state.

For the safest PostgreSQL upgrade:

  1. Verify database backups/PITR and rehearse the target version against a restored clone.
  2. Drain or stop every old broker before allowing the first new broker to initialize the schema.
  3. Start one new broker, wait for /ready, and verify the schema and authoritative queue counts.
  4. Start the remaining brokers at the same bunqueue version, then update clients.

Do not assume an application-only rollback is safe after a schema migration. A binary that supports an older schema version refuses to start against a newer recorded version. If the new broker has migrated the database, roll forward or restore the pre-upgrade database/PITR point together with the old binaries. Zero-downtime mixed-version rollout requires explicit compatibility evidence for the exact source and target versions; it is not implied by the additive shape of the current migration.

graceful shutdown

SIGTERM is a contract.

Deploys restart the server constantly. The shutdown path is what makes that boring.

On SIGTERM the server stops accepting connections and waits up to SHUTDOWN_TIMEOUT_MS (default 30,000 ms) for active jobs to finish. SQLite then flushes its pending write buffer and closes the database, which is why ordinary buffered jobs survive a graceful deploy. PostgreSQL instead closes operation admission, drains work already admitted to the manager, settles deferred writes and projection repairs, releases the broker’s remaining owned leases, and then closes its SQL pool.

Jobs still running when the deadline hits remain in the selected persistent backend as active and are recovered for retry by SQLite startup or PostgreSQL lease recovery. Size SHUTDOWN_TIMEOUT_MS a little above your longest handler, and give your orchestrator a termination grace period slightly above that (Kubernetes defaults to 30 s).

go-live checklist

Before you point traffic at it.

  • SQLite only: durable: true on every job you cannot re-derive; PostgreSQL writes are transactional without this option

  • idempotency keys in every handler side effect
  • custom jobId wherever producers retry

  • SQLite: S3_BACKUP_ENABLED=1; PostgreSQL: provider backup/PITR; rehearse one restore

  • alerts on DLQ growth and waiting-queue depth
  • /healthz and /ready wired into your orchestrator

  • AUTH_TOKENS set; METRICS_AUTH if metrics are exposed

  • native TLS or a private network between clients and server

  • SHUTDOWN_TIMEOUT_MS sized above your longest handler

  • SQLite disk alerts for .db/.db-wal, or PostgreSQL storage/connection alerts

  • worker lockDuration above your slowest job, or heartbeats on

Ready to deploy?

The deployment guide has the Docker, systemd, and PM2 configs to paste.