# Runbooks This document captures the operational readiness artifact for plan item `Runbooks`. Scope: - API and worker incidents for `video-clipping-service` - Database-backed queue behavior - Current dashboard and alert coverage - Cleanup and retention job behavior Primary references: - Dashboard: `dashboards/video-clipping-overview-grafana.json` - Dashboard notes: `docs/dashboards.md` - Alert rules: `docs/alerts.md` - Autoscaling policy: `docs/autoscaling-policies.md` - Queue timeout tuning: `docs/queue-visibility-timeout-tuning.md` ## Service Model Current implementation assumptions: - Metrics are exposed at `/actuator/prometheus`. - Health endpoints are exposed through Spring Boot Actuator. - Queue mode may be `memory` or `db`, but operational queue metrics and retry behavior documented here assume `video-clipping.queue=db`. - Database-backed queue messages move through `PENDING`, `PROCESSING`, `COMPLETED`, or `DLQ`. - Worker retries are driven by: - `video-clipping.database-queue.max-attempts` - `video-clipping.database-queue.retry-backoff-ms` - `video-clipping.database-queue.visibility-timeout-ms` - Cleanup jobs are driven by: - `video-clipping.cleanup.enabled` - `video-clipping.cleanup.local-artifact-poll-interval-ms` - `video-clipping.cleanup.retention-poll-interval-ms` ## Quick Triage When an alert fires: 1. Identify whether the impact is API availability, worker throughput, or data cleanup. 2. Check the Grafana dashboard for: - pending queue depth - oldest pending age - DLQ count - API request rate - API p95 latency - CPU usage - heap usage 3. Check `/actuator/health` for general service status. 4. Correlate logs using `jobId`, `assetId`, `tenantId`, and `traceId` if available. 5. Determine whether the issue is isolated to one pod or systemic across API and worker replicas. ## Runbook: Queue Age High Signals: - `VideoClippingQueueOldestPendingHigh` - `VideoClippingQueueOldestPendingCritical` - Dashboard shows elevated `Oldest Pending Age` and rising `Pending Queue Depth` Likely causes: - Worker capacity is too low for current workload. - Workers are crash-looping or not polling. - Source staging, FFmpeg, or clip upload is slow. - Database queue messages are stuck until visibility timeout reclaim. What the code does: - The database queue poller claims `PENDING` messages and marks them `PROCESSING`. - Expired `PROCESSING` messages are reclaimed after `video-clipping.database-queue.visibility-timeout-ms`. - Retry scheduling uses `video-clipping.database-queue.retry-backoff-ms`. Immediate actions: 1. Confirm queue mode is `db`. This runbook does not apply as written to `memory` mode. 2. Check worker replica health and restart status. 3. Check whether CPU is saturated on workers. 4. Check whether pending depth is growing faster than processing count. 5. Check recent deploys, node pressure, object storage reachability, and FFmpeg availability. Mitigation: 1. Scale worker replicas up within the bounds in `docs/autoscaling-policies.md`. 2. If workers are healthy but backlog remains, inspect whether `EXACT` jobs or slower object storage throughput are dominating runtime. 3. If many jobs appear stuck in `PROCESSING`, compare observed runtime to `video-clipping.database-queue.visibility-timeout-ms`. 4. If the timeout is too small for production runtime, raise it before increasing concurrency aggressively. Exit criteria: - `video_clipping_queue_oldest_pending_age_seconds` trends down - pending depth stabilizes or shrinks - no new DLQ growth caused by the mitigation ## Runbook: DLQ Messages Present Signals: - `VideoClippingDlqPresent` - Dashboard `DLQ Messages` stat is nonzero What the code does: - A failed message is retried until `video-clipping.database-queue.max-attempts`. - On final failure, the queue message is marked `DLQ`. - The clip job is marked `FAILED`. - Job events include `DLQ` and `FAILED`. Likely causes: - FFmpeg processing failure - invalid or corrupt source media - object storage materialization or upload failure - persistent code regression affecting all retries Immediate actions: 1. Determine whether DLQ growth is isolated or systemic. 2. Inspect failed job records and recent job events for representative `jobId` values. 3. Check whether failures began after a deploy or configuration change. 4. Verify object storage, database, and worker runtime dependencies. Mitigation: 1. If the issue is a bad deployment, roll back or stop worker rollout. 2. If the issue is data-specific, isolate affected tenants or assets and avoid mass redrive. 3. If the issue is dependency-related, restore storage, FFmpeg, or database health first. 4. Redrive only after the root cause is fixed. Redrive guidance: - There is no dedicated automated redrive artifact yet in the repo. - Redrive must preserve idempotency and avoid replaying into an unresolved failure condition. - Use representative sample jobs first before bulk replay. Exit criteria: - DLQ count stops increasing - new jobs succeed normally - redriven sample jobs no longer fail for the same reason ## Runbook: API 5xx Rate High Signals: - `VideoClippingApi5xxRateHigh` - `VideoClippingDownloadUrl5xxRateHigh` - Rising API p95 latency or reduced request success rate Likely causes: - storage signing or retrieval issues - unexpected server exception - dependency degradation - hot path regression after deploy Immediate actions: 1. Check whether all endpoints are failing or only a specific route such as `/v1/clips/{clipId}/download-url`. 2. Review recent application logs for exception bursts. 3. Check `/actuator/health`. 4. Compare API symptoms with worker and queue signals to decide whether the issue is upstream, downstream, or local to the API. Mitigation: 1. If failures are isolated to `download-url`, verify storage adapter health and signing configuration. 2. If failures affect create/upload/job endpoints broadly, verify repository and queue dependency health. 3. If failures correlate to a new release, roll back the API deployment. 4. If latency is high before 5xx increases, consider scaling API replicas while root cause analysis continues. Exit criteria: - 5xx ratio returns below alert threshold - p95 latency returns to expected baseline - no recurring exception burst in logs ## Runbook: Instance Down Signals: - `VideoClippingInstanceDown` - missing target in Prometheus - Kubernetes readiness or liveness failures Immediate actions: 1. Determine whether the issue affects API pods, worker pods, or both. 2. Check pod phase, restart count, and last termination reason. 3. Check readiness and liveness probe failures. 4. Verify whether the issue is caused by node loss, crash loop, bad config, or scrape misconfiguration. Mitigation: 1. If the service is crash-looping after deploy, roll back. 2. If only one target is affected, replace the pod and monitor the rest of the replica set. 3. If Prometheus scraping changed but the app is healthy, fix scrape config and keep incident scope narrow. Exit criteria: - target is consistently scraped again - pod restarts stop increasing - related queue or API alerts clear ## Runbook: Process CPU High Signals: - `VideoClippingProcessCpuHigh` - dashboard CPU panel shows sustained `process_cpu_usage` Likely causes: - worker saturation due to `EXACT` jobs - insufficient worker replica count - node oversubscription - runaway request or processing loop Immediate actions: 1. Determine whether high CPU is on API or worker pods. 2. Compare CPU saturation with queue age and pending depth. 3. Check whether workload mix changed toward more CPU-intensive jobs. Mitigation: 1. If workers are saturated and queue backlog is growing, scale workers first. 2. If API pods are saturated, scale API pods and inspect hot endpoints. 3. If a specific deployment introduced the condition, roll back. 4. If the node is pressured, rebalance or move workloads. Exit criteria: - CPU drops below sustained threshold - queue age and latency normalize ## Runbook: Cleanup Or Retention Not Keeping Up Signals: - local disk usage continues to grow - old generated artifacts remain under `tmp/` - expired clips or source objects are not being removed on schedule What the code does: - `LocalArtifactCleanupJob` removes expired files under: - configured FFmpeg input directory - configured FFmpeg output directory - `tmp/in-memory-storage` - `tmp/stub-output` - `RetentionCleanupJob` deletes expired clips and expired source objects, then marks assets deleted. Immediate actions: 1. Verify `video-clipping.cleanup.enabled=true`. 2. Confirm poll intervals and retention hours are set as expected. 3. Check whether the service instance that owns scheduled jobs is running. 4. Verify object storage delete operations are succeeding. Mitigation: 1. If cleanup was disabled accidentally, re-enable it and monitor deletion progress. 2. If local artifacts are too large for current retention windows, temporarily scale down worker throughput or increase local storage while cleanup catches up. 3. If object deletion is failing, fix storage access first before manually deleting metadata. Exit criteria: - local artifact growth stops - expired assets and clips are removed on schedule ## Escalation Rules Escalate immediately if: - DLQ count is growing and new jobs continue to fail after rollback - queue age stays above critical threshold after scaling workers - API 5xx affects upload completion, job creation, or clip download across tenants - cleanup lag risks disk exhaustion or unbounded storage growth ## Known Gaps - No dedicated first-class metric exists yet for clip-job failure rate. - No dedicated FFmpeg exit-code metrics exist yet. - No dedicated object storage failure counter exists yet. - No dedicated runbook automation exists yet for DLQ redrive. These gaps are consistent with `docs/alerts.md` and should be closed before treating the operational surface as fully production-complete.