6.9 KiB
6.9 KiB
Cinematic Highlight Quality Rules (source-adaptive)
Derived from real defects found on the DJI and bowling sources (2026-07-23). The governing principle: every rule is dynamic — it MEASURES the source (frames or audio) and adapts. Nothing is a fixed constant tuned to one clip. A fixed grade/geometry/level that looks right on one video is wrong on the next.
Legend: ✅ implemented · ⏳ planned (see cinematic-highlight-poc-plan.md P5.x).
R1 — Exposure: normalize to a target, never crush ✅
- Measure: sample source frames (
ffmpeg signalstatsYAVG, 2 fps, bounded) → mean luma.HighlightFfmpegRenderer.probeSourceLuma. - Adapt:
exposureNormalizationFilterpicks a gamma that maps the measured mean → a target band (~120/255). Dark clip → brightened; well-exposed clip → left alone; over-bright → pulled down. - Invariant: the stylistic grade must be exposure-preserving — lift the black point, keep gamma ≥ 1, brightness ≥ 0, soft vignette. A grade must never reduce mean luma the way the old one did (111 → 63).
- Evidence: bowling final went 63 → ~100 with the source-measured gamma 1.098.
R2 — Orientation: respect the source, never distort ✅
- Measure: read display rotation (
FfmpegClipInspectorside-data /rotatetag) → effective W×H. - Adapt:
outputGeometryrenders portrait sources to a portrait frame, landscape to widescreen; the 2.39 letterbox is applied to landscape only. Portrait→portrait scale is proportional, so no stretch.
R3 — Audio balance: the story audio must lead ✅
- Rule: the generated score is the bed and leads; source audio is ducked under it (−20 dB, −24 with narration); with narration, voice leads via side-chain ducking. No element is silently buried.
- Accurate mastering: single-pass
loudnormis only ~±2 LUFS accurate, so the finished file is measured (probeIntegratedLoudness) and a corrective gain (loudnessGainDb→masterLoudness) is applied to hit the target, with a brickwall limiter for true peak. No-ops when already on target (e.g. a render that lands at −15.5 needs no fix; one at −18.7 is boosted). Handles MusicGen loudness variance. - Verify: integrated loudness ≈ −16 LUFS, true peak ≤ −1.5 dBTP; confirm the score is audible, not just present.
R4 — Duration: length follows the story ✅
- No fixed min/max highlight length. The cut is as long as the beats need (validator only checks positive duration + sane playback speed + candidate containment).
R5 — No dead air: motion in held shots ✅
- Measure: per-shot motion = mean temporal luma difference (YDIF via signalstats) over the shot's source
range.
HighlightFfmpegRenderer.probeSegmentMotion. - Adapt:
pushInAmountturns low motion into a stronger in-shotzoompanpush-in and high motion into a gentle one (with a floor so every shot has a little life); applied beforesetptsso slow-motion survives. - Evidence (bowling): static opening shot measured lowest motion → strongest push (0.117); active celebration → gentlest (0.084). Confirmed visually (opening pushes in ~11% over 2s).
R6 — Transitions: ease, don't jerk — crossfades ✅ / speed-ramp ⏳
- Crossfades (DONE):
xfadeTimelineCommanddissolves montage beats (video xfade + audio acrossfade) instead of hard-cutting, incl. the cut into the slow-mo payoff. Timeline compresses by (n-1)*xf;shiftOverlayForCrossfadere-times overlays and the reported duration is adjusted so overlays/loudness/QA stay aligned. Opt-in viaediting.crossfade-seconds(0 = hard cuts default; localpoc 0.25), clamped to half the shortest beat. Verified on bowling (visible dissolve, overlay stayed on the payoff). - Speed-ramp into slow-mo (deferred): easing the playback-speed change itself conflicts with the R5
per-shot push-in (sub-segment splitting would reset the zoom); needs a time-varying
setptsor a push-in that spans sub-segments. The crossfade already softens the cut into the payoff.
R7 — Overlays: bold, animated, synced ✅
- Rule: captions are large (fontsize 84), thick-outlined + drop-shadowed for legibility on any background,
with a snappy entrance (0.18 s alpha punch + 34 px rise-up over 0.22 s) and a soft ease-out.
drawTextFilter. - Ties to Tier 2 (R9): the vision director produces the caption text ("STRIKE"); R7 makes it land. The overlay is placed on the payoff beat, so it's timed to the musical/edit accent. Verified on bowling.
R8 — Music sync: climax on the payoff ⏳ (P5.13)
- Measure: the payoff beat's timeline position.
- Adapt: prompt/trim the score so its peak lands on the payoff, not wherever the generator happened to put it.
R9 — Show the action, not just the reaction — Tier 1 ✅ / Tier 2 ⏳
- Tier 1 (measurement director, DONE):
HighlightMontageDirectormeasures a per-window motion curve (YDIF) and audio-energy curve (RMS), then composes a story-structured montage automatically: setup → continuous action/tension (release/roll/watch, never chopped) → slow-motion payoff on the audio climax → resolution button, trimming a high-motion camera-whip tail. Enabled byhighlight-scheduler.auto-director-enabled(on in localpoc); writesdirector/montage.jsonafter analysis. Verified on bowling: auto cut ≈ the hand cut. - Tier 2 (semantic VLM director, DONE 2026-07-23/24):
HighlightVisionDirector+tools/vision_caption.pyrun a local vision-language model (moondream2, offline). It now captions several beat frames (two questions each in one call: a discriminative description + a punchy label) and:- guides selection —
semanticScore/semanticCurveturn the descriptions into a per-window highlight-worthiness signal that the montage director blends with audio to place the payoff on the semantically-strongest moment (verified: on bowling the payoff moved onto moondream's detected celebration); - decorates — the payoff label becomes the bold overlay and the description flavors the music.
All generic: the scoring uses generic emotion/action/idle keywords (no content-specific terms), and it
fails soft to the measured cut. ~25s/frame CPU; enabled by
vision-director-enabled(localpoc on). Honest limit: a small VLM on distant subjects is only weakly discriminative — a terse question collapses to a constant answer (use descriptive questions); a stronger VLM or clearer framing would help. moondream2 is Apache-2.0 (commercial-friendly, unlike the CC-BY-NC audio models).
- guides selection —
- Source caveat: a director can only cut what was filmed. If the camera never shows the pins, no tier can.
Rules R1–R5 are live in HighlightFfmpegRenderer / FfmpegClipInspector / HighlightDirectorPlanValidator
and apply to every project automatically. R6–R9 are the next implementation targets; each must likewise be
driven by a source measurement, never a per-video constant.