diff --git a/docs/alerts.md b/docs/alerts.md new file mode 100644 index 0000000..c85ff38 --- /dev/null +++ b/docs/alerts.md @@ -0,0 +1,151 @@ +# Alerts + +This document captures the alerting artifact for operational readiness plan item `Alerts`. + +Current metric surface: + +- Prometheus scrapes `/actuator/prometheus`. +- Database-backed queue deployments expose: +- `video_clipping_queue_pending` +- `video_clipping_queue_processing` +- `video_clipping_queue_dlq` +- `video_clipping_queue_oldest_pending_age_seconds` +- Standard Spring Boot and JVM metrics include: +- `http_server_requests_seconds_count` +- `process_cpu_usage` +- `system_cpu_usage` +- `jvm_memory_used_bytes` +- `up` + +Alert coverage implemented now: + +- Queue age too high. +- DLQ messages present. +- API 5xx rate. +- Process unavailable. +- Sustained process CPU saturation. + +Alert rules: + +```yaml +apiVersion: monitoring.coreos.com/v1 +kind: PrometheusRule +metadata: + name: video-clipping-service + labels: + app.kubernetes.io/name: video-clipping-service +spec: + groups: + - name: video-clipping-service + rules: + - alert: VideoClippingQueueOldestPendingHigh + expr: max(video_clipping_queue_oldest_pending_age_seconds) > 120 + for: 10m + labels: + severity: warning + service: video-clipping-service + annotations: + summary: Video clipping queue age is elevated + description: Oldest pending queue message has been older than 120 seconds for 10 minutes. + + - alert: VideoClippingQueueOldestPendingCritical + expr: max(video_clipping_queue_oldest_pending_age_seconds) > 300 + for: 15m + labels: + severity: critical + service: video-clipping-service + annotations: + summary: Video clipping queue age is critically high + description: Oldest pending queue message has been older than 300 seconds for 15 minutes. + + - alert: VideoClippingDlqPresent + expr: sum(video_clipping_queue_dlq) > 0 + for: 5m + labels: + severity: critical + service: video-clipping-service + annotations: + summary: Video clipping DLQ contains messages + description: At least one clip job message has remained in the dead-letter queue for 5 minutes. + + - alert: VideoClippingApi5xxRateHigh + expr: | + ( + sum(rate(http_server_requests_seconds_count{uri!~"/actuator.*",status=~"5.."}[5m])) + / + clamp_min(sum(rate(http_server_requests_seconds_count{uri!~"/actuator.*"}[5m])), 0.001) + ) > 0.05 + for: 10m + labels: + severity: warning + service: video-clipping-service + annotations: + summary: Video clipping API 5xx ratio is high + description: More than 5 percent of non-Actuator API requests have returned 5xx for 10 minutes. + + - alert: VideoClippingDownloadUrl5xxRateHigh + expr: | + ( + sum(rate(http_server_requests_seconds_count{uri=~".*/download-url",status=~"5.."}[5m])) + / + clamp_min(sum(rate(http_server_requests_seconds_count{uri=~".*/download-url"}[5m])), 0.001) + ) > 0.02 + for: 10m + labels: + severity: warning + service: video-clipping-service + annotations: + summary: Video clipping signed download URL failures are elevated + description: More than 2 percent of signed clip download URL requests have returned 5xx for 10 minutes. + + - alert: VideoClippingInstanceDown + expr: up{job=~".*video-clipping.*"} == 0 + for: 5m + labels: + severity: critical + service: video-clipping-service + annotations: + summary: Video clipping target is down + description: Prometheus has failed to scrape a video-clipping-service target for 5 minutes. + + - alert: VideoClippingProcessCpuHigh + expr: max(process_cpu_usage{application="video-clipping-service"}) > 0.85 + for: 15m + labels: + severity: warning + service: video-clipping-service + annotations: + summary: Video clipping process CPU is saturated + description: Process CPU usage has remained above 85 percent for 15 minutes. +``` + +Threshold rationale: + +- `120s` queue age aligns with the red dashboard threshold and represents backlog outside normal steady-state behavior. +- `300s` queue age is long enough to page because it risks breaching throughput expectations before the `15` minute visibility timeout becomes relevant. +- Any nonzero DLQ count is actionable because failed clip jobs require investigation or redrive. +- A sustained `5%` API 5xx ratio is large enough to avoid paging on isolated errors while still catching broken releases or dependencies. +- A stricter `2%` threshold is used for `download-url` because that endpoint should usually be lightweight and storage-signing failures directly block clip retrieval. +- `85%` sustained process CPU is a warning signal for worker saturation; it should be correlated with queue backlog before scaling policy changes. + +Operational notes: + +- If production uses SQS instead of the database queue, replace the queue-depth and queue-age queries with native cloud queue metrics while keeping the same alert semantics. +- The `up{job=~".*video-clipping.*"}` matcher should be narrowed to the actual Prometheus scrape job names used in the target environment. +- The `process_cpu_usage{application="video-clipping-service"}` selector assumes the Micrometer `application` label is present from `spring.application.name`. If the label set differs in production, adjust the selector accordingly. +- The `download-url` matcher depends on the `uri` label shape exported by Spring Boot HTTP metrics. Verify that the label resolves to the route template used by the deployment. + +Current instrumentation gaps: + +- Job failure rate above threshold is not directly alertable yet because the service does not currently export a counter for failed clip jobs. +- Worker disk usage high is not directly exposed by this service; alert it from Kubernetes or node metrics such as container filesystem or ephemeral-volume usage. +- FFmpeg exit-code spike is not directly alertable yet because FFmpeg outcome counters by exit class are not exported. +- Object storage errors are only indirectly visible through API/worker failures today; a dedicated storage operation failure counter is still needed. +- Signed URL generation failures are only approximated via the `download-url` endpoint 5xx alert; a dedicated business metric would make this more precise. + +Recommended next instrumentation when these alerts need to become first-class: + +- Add counters for clip-job outcomes such as `video.clipping.jobs.started`, `video.clipping.jobs.succeeded`, and `video.clipping.jobs.failed`. +- Add counters for storage operation failures by operation type. +- Add counters for signed URL generation success and failure. +- Add counters for FFmpeg failure categories or exit-code classes. diff --git a/docs/video-clipping-service-implementation-plan.md b/docs/video-clipping-service-implementation-plan.md index 9ab4666..bcf86b2 100644 --- a/docs/video-clipping-service-implementation-plan.md +++ b/docs/video-clipping-service-implementation-plan.md @@ -720,7 +720,7 @@ Keep the domain independent from Spring framework details. Adapters implement st ### Operational Readiness 1. [x] Dashboards. -2. [ ] Alerts. +2. [x] Alerts. 3. [ ] Runbooks. 4. [ ] Backup and restore validation. 5. [ ] DLQ redrive procedure.