# DLQ Redrive Procedure This document captures the operational readiness artifact for plan item `DLQ redrive procedure`. ## Scope This procedure applies to the database-backed queue implementation used when: - `video-clipping.queue=db` It does not apply to: - `video-clipping.queue=memory` - external queue systems such as SQS unless an equivalent provider-specific procedure is written ## Current Queue Behavior Relevant queue message states: - `PENDING` - `PROCESSING` - `COMPLETED` - `DLQ` Relevant clip-job states: - `QUEUED` - `RUNNING` - `SUCCEEDED` - `FAILED` - `CANCEL_REQUESTED` - `CANCELLED` How a message reaches DLQ: - a worker claims a `PENDING` message - processing fails - the message is retried until `video-clipping.database-queue.max-attempts` - on the final failed attempt: - the queue message is marked `DLQ` - the clip job is marked `FAILED` - job events record `DLQ` and `FAILED` Important implementation facts: - There is no built-in one-click redrive command in this repository. - A restored or edited queue message can be picked up automatically by workers once it is returned to `PENDING`. - `PROCESSING` messages older than `video-clipping.database-queue.visibility-timeout-ms` are reclaimed automatically. ## Preconditions Before redriving any DLQ message: 1. Confirm the root cause is understood. 2. Confirm the root cause is fixed. 3. Confirm the failed job is safe to retry. 4. Confirm the source object referenced by the job still exists. 5. Confirm there is no active incident still affecting storage, workers, or the database. Never redrive blindly when: - DLQ is still growing - workers are still failing the same workload - a bad deployment is still live - object storage is unavailable - the source object has already been deleted by retention ## Triage Queries Inspect DLQ volume: ```sql select status, count(*) from queue_messages group by status order by status; ``` Inspect candidate DLQ rows: ```sql select id, job_id, status, attempt_count, available_at, claimed_at, last_error from queue_messages where status = 'DLQ' order by claimed_at desc nulls last, available_at desc limit 50; ``` Inspect the corresponding clip jobs: ```sql select id, asset_id, tenant_id, status, attempt_count, progress_percent, error_message, started_at, finished_at from clip_jobs where id in ( select job_id from queue_messages where status = 'DLQ' ) order by finished_at desc nulls last; ``` Inspect recent job events for one candidate: ```sql select event_type, message, created_at from job_events where aggregate_id = :job_id order by created_at asc; ``` ## Redrive Decision Rules A DLQ job is eligible for redrive only if all of the following are true: - job status is `FAILED` - queue message status is `DLQ` - the failure cause has been fixed - the source video object still exists - retrying will not violate tenant or operational constraints Do not redrive if: - the job was intentionally cancelled - the source object is gone - the failure was caused by permanently invalid media - duplicate replay would create unwanted business behavior ## Safe Redrive Strategy Use a staged approach: 1. Pause or scale down workers. 2. Select a small sample of DLQ messages. 3. Reset only those messages and jobs. 4. Start one worker replica. 5. Observe behavior. 6. Expand only if sample replay succeeds. Recommended initial batch size: - `1` to `5` jobs Do not bulk-redrive the full DLQ on the first pass. ## Manual Redrive Procedure ### Step 1: Quiesce Workers Before modifying queue rows: - scale worker replicas to `0`, or - otherwise stop worker polling This prevents workers from consuming partially edited rows mid-procedure. ### Step 2: Select Candidate Jobs Pick specific `job_id` and `queue_messages.id` values to redrive. Record: - queue message id - job id - tenant id - last error - source object key Verify source object key: ```sql select a.id, a.source_object_key, a.status, j.id as job_id, j.status as job_status from clip_jobs j join video_assets a on a.id = j.asset_id where j.id = :job_id; ``` ### Step 3: Snapshot Candidate Rows Before changing anything, capture the current database state for the candidate rows. Minimum snapshot queries: ```sql select * from queue_messages where id = :message_id; select * from clip_jobs where id = :job_id; select * from job_events where aggregate_id = :job_id order by created_at asc; ``` ### Step 4: Reset Queue Message And Job State For a controlled retry, reset: - queue message to `PENDING` - queue message `available_at` to now - queue message `claimed_at` to `null` - queue message `last_error` to `null` - clip job to `QUEUED` - clip job `finished_at` to `null` - clip job `error_message` to `null` - clip job `progress_percent` to `0` Important note on attempt count: - keep `attempt_count` history intact for auditability unless there is a specific reason to zero it out - if you keep the current attempt count, the next retry will increment again and may go straight back to `DLQ` if `max-attempts` has not been adjusted Operational recommendation: - for one-off redrive after a fix, reset `queue_messages.attempt_count` and `clip_jobs.attempt_count` to `0` - do this only when the incident has been clearly resolved and you want a clean retry window Example transactional update for one job: ```sql begin; update queue_messages set status = 'PENDING', attempt_count = 0, available_at = now(), claimed_at = null, last_error = null where id = :message_id and status = 'DLQ'; update clip_jobs set status = 'QUEUED', attempt_count = 0, progress_percent = 0, finished_at = null, error_message = null where id = :job_id and status = 'FAILED'; insert into job_events (id, aggregate_id, event_type, message, created_at) values (:event_id, :job_id, 'REDRIVE_REQUESTED', 'Operator returned DLQ message to PENDING after root-cause fix', now()); commit; ``` If any precondition fails, roll back instead of forcing the update. ### Step 5: Resume One Worker After resetting a small batch: - start one worker replica - monitor queue depth, oldest pending age, and DLQ count - inspect the redriven jobs directly ### Step 6: Validate Outcome A successful redrive should produce: - queue message status moves to `COMPLETED` - clip job status moves to `SUCCEEDED` - asset status moves to `READY` - no immediate replacement DLQ row for the same job Validate with: ```sql select id, status, attempt_count, last_error from queue_messages where id = :message_id; select id, status, attempt_count, progress_percent, error_message from clip_jobs where id = :job_id; ``` And verify job events: ```sql select event_type, message, created_at from job_events where aggregate_id = :job_id order by created_at asc; ``` ## Bulk Redrive Rules Bulk redrive is allowed only after: - sample jobs succeed - no new systemic failures appear - worker capacity is sufficient for replay - source objects still exist for the full batch For bulk batches: - redrive by tenant or failure class where possible - avoid mixing unrelated failure causes - use small waves instead of one large reset Recommended wave size: - start with `10` - increase gradually while watching queue age, API health, and DLQ count ## Rollback Plan If redriven jobs fail again: 1. stop or scale down workers 2. stop further redrive batches 3. inspect whether the failure signature matches the original incident 4. return affected rows to `DLQ` only if needed to preserve operator clarity 5. open a new incident if the root cause was believed fixed but was not ## Audit Requirements For every redrive operation, record: - incident or ticket id - operator name - time of redrive - affected `job_id` values - affected `queue_messages.id` values - root cause summary - evidence that the fix was in place At minimum, write a `REDRIVE_REQUESTED` job event for each manually reset job. ## Pass Criteria The DLQ redrive procedure is considered validated when: - at least one representative DLQ job is replayed successfully in a non-production or controlled environment - the operator can reset queue and job state without creating duplicate processing - monitoring confirms no retry storm or immediate re-DLQ pattern ## Current Repo Limitation This repository defines the procedure but does not yet provide: - an admin endpoint for redrive - a CLI tool for controlled queue replay - automated guardrails for selecting safe DLQ candidates Until those exist, DLQ redrive remains an operator-driven database procedure and should be treated as a controlled maintenance action.