From 6600ded92910e4743db9f64af63150e4eabbaf6a Mon Sep 17 00:00:00 2001 From: JSLMPR Date: Wed, 22 Jul 2026 11:43:10 +0200 Subject: [PATCH] docs: P3.5 mix/ducking review findings (calibration deferred to P3.6) Co-Authored-By: Claude Opus 4.8 Claude-Session: https://claude.ai/code/session_01Bg76sLc43Wc3j5ZcLkboYR --- docs/cinematic-highlight-poc-plan.md | 9 ++++++++- 1 file changed, 8 insertions(+), 1 deletion(-) diff --git a/docs/cinematic-highlight-poc-plan.md b/docs/cinematic-highlight-poc-plan.md index 01d2e2a..138d5dd 100644 --- a/docs/cinematic-highlight-poc-plan.md +++ b/docs/cinematic-highlight-poc-plan.md @@ -92,7 +92,14 @@ Ranked from the first render's evidence: (committed localpoc profile keeps the safe heuristic default; worker started manually). CAVEATS: yolov8n.pt is **AGPL-3.0** (production blocker on this license alone; see yolov8n.pt.license.txt); loopback HTTP is flagged non-compliant for a certified path; Haar face detector gives false positives. -- [ ] P3.5 Music/SFX creative fit + ducking review; consider longer inference / better prompts. +- [~] P3.5 Music/SFX + ducking review (measurement-based). FINDINGS: sidechain ducking IS implemented + (`[music_raw][voice]sidechaincompress` with the configured threshold/ratio/attack/release; FFmpeg + auto-splits `[voice]` so it both keys the duck and stays in the mix) and it fires (music drops during + voice). Source audio is near-silent (-54 dB RMS) -> negligible. Loudness/true-peak stay in spec. + OPEN: voice-vs-music balance cannot be validated or tuned by measurement alone — MusicGen produces + different audio each run, so cross-render A/B is confounded, and this needs LISTENING. A blind +3.5 dB + voice boost was tried and reverted (could not verify it helped; measurement suggested it did not). + Per the "do not tune audio blindly" discipline, mix calibration is deferred to P3.6 (human review). - [ ] P3.6 Freeze acceptance thresholds + blinded human creative review vs baseline before declaring success. ## Milestone log