video_editing_poc/docs/cinematic-highlight-poc-pla...

207 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Local Cinematic Highlight PoC — Plan & Milestones
Owner: (Jaden) · Started: 2026-07-21 · Machine: Intel Mac (x86_64, 16 CPU, 32 GB, no GPU)
## Objective
Produce ONE genuinely cinematic highlight from a local source video, end-to-end through the existing
pipeline, using only local models resident in the service runtime:
Piper (voiceover) + MusicGen (music) + AudioLDM2 (SFX), with deliberate visual treatment and pacing,
and explicit human approval before any render.
A valid MP4 or a passing test is NOT success. Success = representative output passes measured media QA
and a human creative review.
## Non-negotiables (still in force during the PoC)
No external AI services in the media path · no placeholder silence/tones/OS `say` as a finished asset ·
no unlicensed assets · no rendering without explicit approval for the specific project · no inference-time
network (models load from local dirs, `HF_HUB_OFFLINE=1`). NOTE: the user authorized model DOWNLOADS on
2026-07-21 (one-time provisioning), reversing the earlier no-download stance; inference stays offline.
## Proven capability baseline (2026-07-21)
| Model | Runtime | Status | Evidence |
|---|---|---|---|
| Piper `en_US-lessac-medium` | piper-tts 1.5.0 | ✅ works | real speech, 22.05 kHz |
| MusicGen `musicgen-small` | transformers 4.44.2 + torch 2.2.2 | ✅ works | 5 s / 44.7 s CPU, mean 18 dB, 32 kHz |
| AudioLDM2 `cvssp/audioldm2` | diffusers 0.30.3 | ✅ works | 3 s / 22.3 s CPU, mean 21.8 dB, 16 kHz |
Environment traps (see memory `local-model-runtime-intel-mac`): torch capped at 2.2.2 (Intel-Mac),
numpy<2, transformers must be 4.x, audiocraft/xformers unusable here, `HF_HUB_DISABLE_XET=1` for downloads.
## Plan
### Phase 1 — Generate all three assets THROUGH the pipeline
- [x] 1.1 Rewrite `tools/local_asset_worker.py`: musictransformers MusicGen, sfxdiffusers AudioLDM2
(resample to 48 kHz mono), voiceoverPiper (unchanged). audiocraft path removed. CLI + exit codes preserved.
- [x] 1.2 Materialize models into stable `models/` dirs via `save_pretrained` (`tools/provision_local_models.py`):
`models/musicgen-small` (2.2 G), `models/audioldm2` (4.2 G); `models/piper` (60 M) already present.
- [x] 1.3 License/provenance `.license.txt` sidecars for each model marker file (config.json / model_index.json /
voice .onnx). BOTH audio models are NON-COMMERCIAL (MusicGen CC-BY-NC-4.0, AudioLDM2 CC-BY-NC-SA-4.0)
flagged for production review. Honors the `AssetLicensePolicy` regular-file gate; no Java change.
- [x] 1.4 Smoke test: voiceover 3.67 s/15.8 dB, music 4.94 s/12.0 dB, sfx 3.00 s/20.6 dB all 48 kHz mono,
real signal, exit 0.
- [x] 1.5 `mvn -o verify` 245 tests / 62 classes / 0 failures / 0 errors / 0 skips (unchanged from baseline).
**Phase 1 COMPLETE (2026-07-21).**
### Phase 2 — One approved end-to-end highlight (STOP before render for explicit approval)
- [x] 2.1 Opt-in `localpoc` Spring profile added (`application-localpoc.yml`): venv python, model paths,
offline, no bootstrap auto-start, heuristic visual, isolated PoC I/O dirs, render OFF, approval REQUIRED.
- [x] 2.2 DJI source processed under `localpoc` (app started, ~4 min 4K analysis): wrote `category.json`
(GENERIC_VLOG, conf 0.25) + `highlight-candidates.json` (3 windows: 27-39 s, 36-48 s, 81-93 s).
Actual footage = blue Porsche Taycan aerial orbit (EV car reveal).
- [x] 2.3 Hand-authored `director/edit-plan.json`: 3 x 9 s highlights (opening_hook / rising_energy /
hero_payoff), grounded in the visible car; passes ALL validator constraints (containment, speed 1.0,
duration bounds, VO budget, category match). Manual + local-only (no external AI).
- [x] 2.4 **RENDERED (operator-approved 2026-07-21).** render-enabled override + approved.flag local models
generated 3 VO (Piper) + 3 music (MusicGen) + SFX (AudioLDM2) 3 highlights rendered project
`final.mp4` assembled (27.3 s, 51.7 MB). Flow ~8.6 min. All real local-model assets, no placeholders.
- [~] 2.5 QA measured (honest):
**PASS** 1920x1080 / H.264 / 30 fps / AAC 48 kHz; duration 27.3 s; integrated loudness -15.8 LUFS
(target -16 ±1); no black frames; no long silence; overlays render & are grounded ("First light.",
"Silent power."); VO intelligible.
**FAIL/weak** true peak: highlight_001 = 0.0 dBFS (clips; target -1.5 dBTP); h2/h3 = -1.4 dBFS
(loudnorm TP limiting not enforced tightly). Visual "treatment" is essentially PASSTHROUGH no visible
cinematic grade/reframe (renderer maps any treatment to a generic look); overlays are plain white
captions. LRA 16.9 (high). Pipeline's own qa-report passed because its grade/clip checks are structural.
**Verdict: genuine end-to-end local PoC, NOT yet "genuinely cinematic."** Visual grade + true-peak
limiting are the top Phase 3 items.
### Phase 3 — Iterate to quality
Ranked from the first render's evidence:
- [x] P3.1 Filmic grade in `HighlightFfmpegRenderer.cinematicVisualFilter`: S-curve + teal-orange
colorbalance + eq + unsharp + vignette (replaces the weak fixed eq). Only the highlight renderer
changed; multi-clip FfmpegEditRenderer untouched.
- [x] P3.1b **Beat-specific grading** (a0b6023): grade now varies by story beat (read from the EditPlan
style key, no EditDecision/serialization change) opening=cool/soft (linear_contrast), rising=balanced
(medium_contrast), hero=rich/warm (strong_contrast). Verified visually (calmdramatic arc) + unit test.
- [x] P3.2 True-peak safety: added `alimiter=level=disabled:limit=0.72` after loudnorm in the audio mix.
VERIFIED by re-render: project true peak 0.0 dBFS (clipping) **-1.7 dBFS**; per-highlight -1.7/-2.7/
-2.8 (all <= -1.5 dBTP gate); integrated loudness -16.3 LUFS. (First pass at limit=0.79 gave -1.3 on
h1, 0.2 dB over tightened to 0.72 for inter-sample margin.) Test added.
- [x] P3.3 Overlay styling: refined 48px caption, soft drop shadow, thin subtle border, smooth alpha
fade in/out (0.4s ramps). Verified rendering; committed 24bd1d7. (Placement/safe-area unchanged.)
- [x] P3.4 **Real CV visual analysis via YOLOv8** (resident model, NO bootstrap script). Deps installed into
an isolated `.venv-local-cv`; uvicorn worker run directly against the existing `yolov8n.pt` offline
(`YOLO_OFFLINE=True`, telemetry sync off). Fixed a latent bug: the Java CV client negotiated HTTP/2
(h2c) which the HTTP/1.1-only worker mishandled (422/empty body) pinned HTTP/1.1. Result on the DJI
source: category GENERIC_VLOG@0.25 -> **CAR_VLOG@0.95** (real `car` detection), measured blur 0.191 /
exposure 0.95 (OpenCV), method local_cv_worker_opencv_yolo. Provider stays OPT-IN via runtime override
(committed localpoc profile keeps the safe heuristic default; worker started manually).
CAVEATS: yolov8n.pt is **AGPL-3.0** (production blocker on this license alone; see yolov8n.pt.license.txt);
loopback HTTP is flagged non-compliant for a certified path; Haar face detector gives false positives.
- [~] P3.5 Music/SFX + ducking review (measurement-based). FINDINGS: sidechain ducking IS implemented
(`[music_raw][voice]sidechaincompress` with the configured threshold/ratio/attack/release; FFmpeg
auto-splits `[voice]` so it both keys the duck and stays in the mix) and it fires (music drops during
voice). Source audio is near-silent (-54 dB RMS) -> negligible. Loudness/true-peak stay in spec.
OPEN: voice-vs-music balance cannot be validated or tuned by measurement alone — MusicGen produces
different audio each run, so cross-render A/B is confounded, and this needs LISTENING. A blind +3.5 dB
voice boost was tried and reverted (could not verify it helped; measurement suggested it did not).
Per the "do not tune audio blindly" discipline, mix calibration is deferred to P3.6 (human review).
- [~] P3.6 Acceptance gate set up in `docs/cinematic-highlight-acceptance-review.md`: FROZEN technical
thresholds + creative rubric (04, avg≥3.5 / no dim<3). **Gate A (technical) measured = PASS** on the
canonical render (1080p/H.264/AAC 48k, 15.9 LUFS, TP 2.4/2.5/2.7 dBTP, no black/silence, all assets
present; one marginal note on aggregate duration +0.32 s). **Gate B (creative) requires the human
review** cannot be self-assessed; awaiting reviewer scores.
## Milestone log
- 2026-07-21: Models provisioned & individually proven (Piper, MusicGen, AudioLDM2). Plan approved. Phase 1 started.
- 2026-07-21: **Phase 1 complete.** Worker rewritten (audiocrafttransformers MusicGen + diffusers AudioLDM2),
models materialized to `models/` with license sidecars, all 3 asset kinds generate 48 kHz mono real audio
through the worker, `mvn -o verify` green (245/0/0/0). Paused for review before Phase 2.
- 2026-07-21: Phase 1 committed (a95d1fa). **Phase 2 up to approval gate:** localpoc profile added, DJI source
processed candidates + category, grounded 3-highlight director plan authored & pre-validated.
STOPPED awaiting explicit render approval for dji_21230510013241_0091_d.
- 2026-07-21: **Phase 2 rendered** (operator-approved) first full local sourcefinal highlight reel.
- 2026-07-22: **Phase 3.1 + 3.2 done & verified.** Filmic grade + true-peak limiter. Re-render measured TP
-1.7/-2.7/-2.8 dBFS (was 0.0 clipping), I -16.3 LUFS; mvn verify 245/0/0/0. Committed 588c652 (profile),
0fcfca9 (render quality).
- 2026-07-22: **Phase 3.3 done & verified.** Premium overlay styling (soft shadow, alpha fade). Committed
24bd1d7. Combined re-render (grade+limiter+overlays) verified: TP -2.5/-2.6/-2.8 dBFS, I -16.3 LUFS,
1080p/27.3s; mvn verify 246/0/0/0.
- 2026-07-22: **Beat-specific grading done & verified** (a0b6023). Grade varies per story beat; final
re-render TP -2.3/-2.5/-2.8 dBFS, I -16.3 LUFS, 1080p/27.3s; mvn verify 247/0/0/0. Phase 3 visual/audio
craft items complete.
- 2026-07-22: **P3.4 real CV (YOLOv8) done & verified.** Offline worker vs existing yolov8n.pt (no bootstrap).
Fixed latent HTTP/2 client bug. DJI source now classifies CAR_VLOG@0.95 with measured blur/exposure and a
real `car` label (was GENERIC_VLOG@0.25/filename heuristic). mvn verify 247/0/0/0. AGPL + loopback caveats
recorded. Remaining: P3.5 music/SFX fit, P3.6 frozen thresholds + blinded review.
### Phase 4 — Cinematic pass (after human review: "not genuinely cinematic")
Human Gate-B verdict on the Phase-3 output: functional, not cinematic. Direction chosen: dynamic editing,
better music, NO voiceover, bolder finish. Done so far (commit 842f2a7):
- [x] **No voiceover** validator made VO optional; new plan carries none (music + SFX driven).
- [x] **2.39:1 letterbox + film grain** in the renderer; overlays raised above the bar.
- [x] **Slow-motion hero** (plan: hero source span < target 0.67x) + beat-specific bold grade.
- [x] **Stronger music direction** driving orchestral/hybrid prompts (still `musicgen-small`).
- [ ] **Push-in / Ken Burns** prototyped via time-based `crop` zoom; INVALID (FFmpeg crop can't use `t`
for w/h). Reverted. Redo with `zoompan` as a follow-up.
- [ ] **Bigger music model** `musicgen-medium` download stalled + is very slow on CPU; deferred. Consider
medium, one continuous score across the piece, or licensed production music.
- [ ] Re-run human Gate B on the cinematic cut.
Verified: mvn 247/0; final 16.5 LUFS, TP 2.7 dBTP, 1080p, no VO. Deliverable:
`output/localpoc/highlight-projects/dji_21230510013241_0091_d/final.mp4`.
### 2026-07-23 — Second source (bowling strike) + orientation/audio/duration fixes
- New source `bowling_strike.mp4` (portrait phone footage) rendered e2e via montage; confirmed the DJI
Phase-5 finding that **source footage is the ceiling**.
- **Service-wide fixes applied** (not just PoC), all with tests, `mvn -o verify` 254/0:
- **Orientation-aware rendering.** `FfmpegClipInspector` now reads display rotation (side-data + `rotate`
tag) and stores effective width/height; `HighlightFfmpegRenderer.outputGeometry()` renders portrait
sources to a portrait frame and skips the 2.39 letterbox (landscape unchanged). Fixes portrait footage
being stretched into landscape.
- **Audio mix.** Source audio (`[0:a]`) ducked to 20 dB (24 with narration) so the generated score leads
instead of being buried; montage music cue raised to 9 dB.
- **Highlight duration constraint removed.** Validator no longer enforces min/max (only positive/finite +
sane speed + candidate containment); config bounds relaxed to soft candidate hints (min 2 / max 3600).
A highlight can be as long as the story needs.
- E2e render flow captured in memory `highlight-e2e-render-runbook` (token-saver).
### Phase 5 — Next steps (roadmap as of 2026-07-22, after montage mode)
To reach genuinely cinematic output (Gate B):
- [ ] P5.1 **Better source footage** (biggest lever; source is the real ceiling parked car / industrial lot).
- [x] P5.2 **Automate shot selection — Tier 1 (DONE 2026-07-23).** `HighlightMontageDirector` composes
`director/montage.json` from measured motion (YDIF) + audio-energy (RMS) curves: setup continuous
action/tension slow-mo payoff on the audio climax resolution button, camera-whip tail trimmed. Wired
into the scheduler behind `auto-director-enabled` (localpoc on). Verified: auto cut hand cut on bowling.
- [ ] P5.3 **Better music** musicgen-medium/large (slow) or licensed production track; one produced score.
- [ ] P5.4 Polish: proper `zoompan` in-shot push-in; beat-synced cut lengths; better/optional voice.
To make it a real system:
- [x] P5.5 **Resident local director (DONE 2026-07-23).** Two tiers, both offline + auto: Tier-1
`HighlightMontageDirector` (measured motion+audio shot list) and Tier-2 `HighlightVisionDirector` +
`tools/vision_caption.py` (moondream2 captions the payoff semantic overlay "STRIKE" + scene music).
Wired into the scheduler (`auto-director-enabled`, `vision-director-enabled`; localpoc on). Plans are no
longer manual. See `docs/cinematic-quality-rules.md` R9.
To reach production (deferred until quality passes):
- [ ] P5.6 Formal Gate B human review (rubric in cinematic-highlight-acceptance-review.md).
- [ ] P5.7 **Licensing blocker**: MusicGen CC-BY-NC, AudioLDM2 CC-BY-NC-SA, YOLOv8 AGPL are non-commercial/
copyleft -> production needs commercially-licensed models/assets.
- [ ] P5.8 The hardening backlog below (security, CI/CD, containers, PostgreSQL, no-egress cert, repro build).
### Director review of the bowling portrait cut (2026-07-23) — general craft gaps
Frame-by-frame + audio review. Verdict: technically clean (correct portrait, -15.5 LUFS, TP -2.8, no defects,
score leads) but NOT yet top-tier cinematic. Craft gaps, all fixable as GENERAL pipeline features:
- [x] P5.9 **Adaptive exposure (DONE 2026-07-23).** `probeSourceLuma` measures source frames (signalstats
YAVG); `exposureNormalizationFilter` maps the measured mean to a target band (source-driven gamma); the
grade is now exposure-preserving (black-lift, gamma≥1, brightness≥0, soft vignette). Bowling final went
63→~100 mean luma with a measured gamma of 1.098. General rules captured in `docs/cinematic-quality-rules.md`.
- [x] P5.10 **Life in held frames (DONE 2026-07-23).** `probeSegmentMotion` (YDIF) measures per-shot motion;
`pushInAmount` maps low motion → stronger in-shot `zoompan` push-in, high motion → gentle (floor keeps a
little life). Applied before `setpts` so slow-mo survives. Bowling: static opening got 0.117 push, active
celebration 0.084. See `docs/cinematic-quality-rules.md` R5.
- [ ] P5.11 **Speed-ramp + crossfades.** The cut into the slow-mo hero is an abrupt speed change and beats are
hard cuts. Ease into slow-mo and add short crossfades on beat boundaries (renderer transition upgrade).
- [ ] P5.12 **Bolder, animated overlay synced to the music hit.** Current overlay is small plain white text;
make it larger/bolder with a scale/fade punch timed to the payoff (general overlay upgrade).
- [ ] P5.13 **Music climax sync.** Align the generated score's peak to the hero beat's timeline position
(pass the payoff timestamp into music prompting/trimming). Ties to P5.3.
- [ ] P5.14 **Show the actual action.** The edit never shows the pins fall (source is on the person). General:
shot-selection (P5.2) should locate the true impact via audio-transient + motion, and source coverage
should include it.
## Deferred (until output-quality gate passes)
Production hardening: Spring Security/OIDC, PostgreSQL/Testcontainers, containers/K8s, CI/CD, distributed
ops, digest-bound authenticated approval. Recorded, not deleted.