Operations
Operations for BlitzQ.
Processes to run
| Process | Command | How many |
|---|---|---|
| Web / producers | your application | any |
| Workers | blitzq worker module:app --queues ... --concurrency N | scale per queue |
| Scheduler | blitzq scheduler module:app | only if you use periodic tasks; run 2+ for availability (occurrences are deduplicated) |
Workers also promote delayed tasks and retries (--no-promote disables it), so a
scheduler process is optional for delayed tasks.
Signals and shutdown
SIGTERM or SIGINT (Ctrl-C) triggers a graceful shutdown. The worker stops
fetching, waits up to --shutdown-timeout (30 s) for running tasks, then cancels
the rest and returns their messages to the queue. A second signal skips the wait.
Give orchestrators a termination grace period longer than shutdown_timeout
(Kubernetes: terminationGracePeriodSeconds).
On Windows, Ctrl-C and Ctrl-Break are handled. SIGTERM from other processes is not deliverable to console apps, so use Ctrl-Break or a service manager there.
Inspecting a deployment
blitzq queue stats --app myapp.tasks:app # queue depth, in-progress, scheduled, DLQ, workers
blitzq task inspect <id> --app myapp.tasks:app # state, attempts, timings, error
blitzq task cancel <id> --app ... # certain for scheduled tasks, best effort otherwise
blitzq dead-letter list --app ... # newest first
blitzq task retry <id> --app ... # replay a dead letter with a fresh attempt budget
blitzq dead-letter purge --yes --app ...
blitzq queue purge <queue> --yes --app ...Instead of --app you can pass --redis-url, --mode and --namespace, or set
the environment variables BLITZQ_APP, BLITZQ_REDIS_URL, BLITZQ_MODE and
BLITZQ_NAMESPACE. Add --json for machine-readable output.
Metrics and logs
- Each worker keeps per-queue counters (
received,started,succeeded,failed,retried,dead_lettered,timeouts,cancelled,recovered,requeued,malformed) and histograms for execution time and enqueue-to-start latency. Labels are the queue name only, never task ids. - Workers publish a snapshot to Redis every 2 s;
blitzq queue stats --jsonincludes it. --metrics-port 9100serves Prometheus metrics (pip install "blitzq[monitoring]"):blitzq_tasks_total\{queue,event\},blitzq_task_duration_seconds,blitzq_queue_latency_seconds.--log-format jsonemits one JSON object per line with fields such astask_id,task,queueandattempt. Task arguments and results are never logged. Exception messages are logged, so don't put secrets in exception text.
Redis configuration
- Persistence decides durability. With
appendonly yesandappendfsync everysec, a Redis crash can lose about the last second of enqueues, acks and results. With RDB only, everything since the last snapshot can be lost. With no persistence, a restart loses all queued, scheduled and in-flight work. Replication to a replica is asynchronous, so a failover can lose recently acknowledged writes too. - Memory: set
maxmemorywithmaxmemory-policy noeviction. An eviction policy would silently delete queued messages and records. Withnoeviction, Redis rejects writes when full, so enqueues fail loudly instead. - Keys live under the namespace (
blitzq:by default). Records expire afterresult_ttl. Dead letters are capped atdead_letter_max(100 000; oldest dropped). Streams shrink as messages are acknowledged (XDEL). - Redis Cluster is not supported. Use a single primary (optionally with replicas and Sentinel).
Security
- Nothing received from Redis is unpickled or executed. Messages are MessagePack
or JSON decoded into plain data with msgspec, then validated against the
envelope schema. Malformed messages are dead-lettered. Messages larger than
Serializer(max_message_size=...)(default 8 MiB) are rejected on enqueue and receipt. - Anyone who can write to your Redis can enqueue any registered task with
arbitrary plain-data arguments. Protect Redis with authentication (ACL user and
password in the URL), TLS (
rediss://) and network isolation, and validate task arguments like any other untrusted input. - Task results may contain sensitive data. Don't expose
get_result/inspectover HTTP without application-level authorization. - Retries can amplify load on a struggling dependency. Use
max_delay, jitter, a bounded number of retries anddont_retry_onfor permanent errors. Per-queue concurrency caps how hard a queue can hit a downstream service.
Timeouts and hung tasks
timeout= cancels async tasks. Thread and process tasks cannot be killed: the
worker stops waiting and records a timeout, but the call keeps its concurrency
slot until it returns. A task that blocks forever therefore permanently consumes
one slot. Monitor running in queue stats and restart such workers.
Upgrading
The message format is versioned by position: new fields are appended with
defaults, so workers of version N+1 read messages of version N. Deploy new
workers before producers when adding tasks, so messages for new task names find
a worker that knows them. Messages for unknown tasks are dead-lettered (reason
unknown task) and can be replayed with blitzq task retry.