video_editing_poc/.claude/skills/cinematic-media-engineering.../SKILL.md

25 KiB

name description
cinematic-media-engineering-reference Load when designing, debugging, or reviewing this repository's highlight detection, FFmpeg trimming/rendering, scene and visual analysis, audio mixing/mastering, voiceover, music/SFX generation, overlays/color, proxies/thumbnails, or local media-model runtime. Use it to interpret media metrics correctly and to separate current heuristics from production evidence.

Cinematic Media Engineering Reference

Last verified against the repository: 2026-07-21.

Use this skill correctly

Use this skill to reason about the media domain and this repository's implementation. Before changing behavior, load video-editing-change-control. For commands, artifact locations, and startup safety, load video-editing-run-and-operate. For a failing run, load video-editing-debugging-playbook. For release evidence and thresholds, load video-editing-validation-and-qa. For property defaults, load video-editing-config-and-flags. For architecture boundaries, load video-editing-architecture-contract.

Do not use this skill as:

  • permission to render, download a model, access a network, or change a production-facing default;
  • proof that an output is cinematic because a heuristic score or qa-report.json says so;
  • a licensing opinion or permission to use an asset/model;
  • a replacement for final-media measurement and an approved human creative review.

Context-specific operating consequences

The no-waiver policy is authoritative in video-editing-change-control. Apply these media-specific consequences before every media or model experiment:

  1. Work offline. Pre-provision every executable, Python wheel, model weight, voice, font, LUT, music track, and SFX file through an approved artifact path.
  2. Pin the artifact version and SHA-256; record its license, source, approval, target platform, and expected output contract.
  3. Fail closed when a required model or asset is missing. Never substitute silence, a sine tone, OS speech, a heuristic provider, or an unrelated library asset.
  4. Treat rendering as a behavior-changing operation. Require explicit approval; do not rely on current packaged defaults.
  5. Never use external AI services. Model/media inference must be in-process or use approved non-network IPC; loopback HTTP is network and is not a certified path.
  6. Route any code, configuration, model, threshold, or render-profile change through video-editing-change-control, then through video-editing-validation-and-qa before promotion.

Current code still violates parts of this policy. tools/run_local_cv_worker.sh installs packages and may fetch yolov8n.pt; tools/run_local_asset_worker.sh installs unpinned packages; and SourceVisualAnalyzer can fall back to filename heuristics. The 2026-07-21 working tree removed asset silence/tone/host-speech fallbacks, made strict asset readiness fail startup, and requires local model paths, but no complete resident asset-model bundle is present. Do not run the bootstrap launchers in a controlled or production environment.

Map the actual media flow

Stage Current implementation Interpret it as
Inspect FfmpegClipInspector reads format duration, codecs, width/height, and r_frame_rate Container-level summary, not a complete timing/color/channel contract
Sample ThumbnailExtractor takes 5 evenly spaced JPEGs by default; ProxyGenerator makes 640-wide H.264 CRF 28/AAC 96k proxies Review aids; neither is full-source evidence
Segment ShotSceneSegmenter uses FFmpeg select='gt(scene,0.35)',metadata=print, then rejects sub-1-second boundaries Frame-difference boundary proposals, not semantic scenes or highlights
Analyze audio SourceAudioAnalyzer chains silencedetect and volumedetect dBFS mean/maximum and thresholded silence only; not LUFS, true peak, dialogue, beat, or quality analysis
Analyze visuals LocalCvVisualAnalysisProvider calls tools/local_cv_worker.py; optional fallback is HeuristicVisualAnalysisProvider Sparse-thumbnail technical hints with major blind spots
Rank HighlightCandidateGenerator makes shot-aligned or overlapping 12-second coverage windows, ranks with source-level visual scores, scene score, duration, and non-silence, and keeps the configured top 3 by default Deterministic baseline with explicit fallback downweighting; not semantic or temporally resolved multimodal understanding
Direct HighlightDirectorPromptGenerator asks a filesystem-capable model for director/edit-plan.json; HighlightDirectorPlanValidator gates it Still an untrusted creative proposal until validated and reviewed; validation proves contract safety, not artistic quality
Prepare assets HighlightAssetPreparationService, HighlightLocalAssetWorker, LocalAssetSynthesizer File materialization; existence does not prove provenance, intelligibility, fit, or quality
Render HighlightDirectorFlowService makes one EditDecision per highlight; HighlightFfmpegRenderer renders and stream-concats A single source range per highlight, not an internally multi-shot edit
Report renderer writes render-manifest.json and qa-report.json; the highlight path probes duration, black ranges, long silence, and sample peaks Technical evidence only; loudness, true peak, A/V sync, freeze detection, raster text safety, and creative fit remain unmeasured

The two ranking paths are deliberately distinct. CinematicHighlightAnalyzer remains the older multi-clip edit-project ranker and writes root-level files. The single-source HighlightSourceScheduler now invokes HighlightCandidateGenerator, which writes analysis/category.json and analysis/highlight-candidates.json before prompt generation. Treat its category as model-supported only when independent local visual labels clear the threshold; metadata/fallback evidence stays low-confidence generic_vlog and requires media review.

Reason about time before touching cuts

Define these terms once:

  • PTS (presentation timestamp): when a decoded frame/sample is presented.
  • DTS (decode timestamp): when it is decoded; reordered codecs may have DTS different from PTS.
  • time base: seconds per timestamp tick. Wall-clock seconds are timestamp * time_base.
  • CFR/VFR: constant/variable frame rate. r_frame_rate is not proof of CFR and is not a frame clock.
  • GOP (group of pictures): frames between independently decodable keyframes; non-keyframes depend on neighbors.

Current plans round seconds to milliseconds. FfmpegClipInspector does not retain time_base, start_time, avg_frame_rate, frame count, rotation, sample aspect ratio, or VFR evidence. Do not assume frame = seconds * r_frame_rate, do not compare decimal timestamps as if they were original PTS, and do not promise frame-accurate cuts from analysis/ffprobe.json alone.

Inspect the missing timing facts without writing media:

INPUT=path/to/source.mp4
ffprobe -v error -show_entries \
  format=start_time,duration,format_name:stream=index,codec_type,codec_name,time_base,start_time,duration,r_frame_rate,avg_frame_rate,nb_frames,width,height,pix_fmt,sample_aspect_ratio,sample_rate,channel_layout \
  -of json "$INPUT"
ffprobe -v error -select_streams v:0 -show_frames \
  -show_entries frame=best_effort_timestamp_time,key_frame,pict_type \
  -of csv=p=0 "$INPUT" | head -n 40

Interpret boundaries as half-open ranges [start,end). For every decision, enforce the repository's duration equation:

rendered_duration = (source_end - source_start) / playback_speed

EditPlanValidator permits only 0.25x through 4.0x and allows 0.05 seconds of equation error. HighlightDirectorFlowService independently clamps speed into that range but does not invoke the validator.

Copy versus re-encode

Operation Boundary behavior Repository use
-c copy trim Start commonly snaps to dependency/keyframe constraints; no filter can run Folder clipping may preserve input quality; do not call this an exact creative cut without packet evidence
Decode + encode trim Decoder can discard to requested time and encoder creates a new boundary; still quantized to stream/filter time bases Both edit renderers trim each segment, filter it, and encode H.264/AAC
Concat demuxer + -c copy Requires stream-compatible segments: codec parameters, time bases, stream layout, and ordering must agree Both renderers join normalized segments this way; mixed audio-presence/channel layouts remain a risk
xfade/acrossfade Requires overlapping, timestamp-normalized streams and shortens/repositions the combined timeline Not implemented

The string crossfade is misleading today. Both renderers translate it into a fade-out on one segment plus a fade-in on the next, then concatenate them without overlap. HighlightDirectorFlowService itself emits only cut, fade-in, and fade-out.

Understand codecs, containers, and formats here

Keep three layers separate: a container (MP4/MOV/MKV) stores streams; a codec (H.264/HEVC/AAC) compresses a stream; a decoded pixel/sample format describes raw data.

The current segment contract is MP4 containing libx264 H.264 at configured resolution and frame rate, yuv420p, CRF 18, preset veryfast, plus optional AAC at the configured bitrate and 48 kHz default. The configured video-bitrate is not used by either editing renderer. ProxyGenerator separately emits H.264 CRF 28, width 640 by default, and AAC 96k. ThumbnailExtractor emits JPEG at -q:v 2.

Do not infer production compatibility from the .mp4 suffix. Before concat or delivery, compare stream count, codec/profile/level, pixel format, dimensions, time base, frame rate behavior, audio sample rate, channel layout, and color metadata. The current inspector cannot do that comparison.

The renderers scale to fit and pad to 1920x1080 by default. This preserves the whole source frame but can letterbox/pillarbox. Any non-none visual treatment also applies the same fixed contrast/saturation/brightness, unsharp, vignette, and a 4% center crop in HighlightFfmpegRenderer; the treatment text does not select a distinct grade. HighlightVisualEffectsStage writes descriptive effect names only. It does not execute LUTs, grain, animated overlays, or prompt-specific effects.

Read the visual signals without overclaiming

Signal Exact current calculation What it can and cannot mean
Scene score FFmpeg scene metadata at selected frames; threshold 0.35 Large frame difference. It can fire on flashes/camera motion and miss semantic changes. ShotSegment.sceneScore is the boundary score entering that segment, not its average quality.
Cut-density “motion” clamp(0.25 + ((segments-1)/duration*60)/20) Editing/camera-change density, not optical or subject motion. A static multicamera cut sequence can score high; a continuous action shot can score low.
Blur score Average clamp(variance(Laplacian(gray))/500) over readable thumbnails Really a sparse sharpness/edge-energy score: higher is sharper. Noise, text, and texture inflate it; intentional shallow focus and low-detail frames depress it.
Exposure score Average clamp(1 - 2*abs(mean(gray)/255 - 0.5)) Mean proximity to mid-gray. It does not detect clipped highlights, crushed shadows, contrast, skin exposure, HDR transfer, or local exposure.
Composition Average of aspect score (0.8 for 1.70-1.90, else 0.55) and thumbnail-presence score (0.7 or 0.45) An aspect/file-presence prior, not composition analysis.
Face presence OpenCV frontal-face Haar cascade; any detection yields faces_detected Sparse, frontal, binary presence only. No identity, consent, expression, tracking, screen area, or false-positive calibration.
Objects YOLO max confidence per class over sampled thumbnails, top 10 Detector confidence on sparse frames, not semantic relevance or continuous presence. Current default yolov8n.pt may auto-download.
Filename heuristic Tokens such as porsche, car, food, family in clip ID/path Metadata guess only. Confidence constants are authored values, not calibrated probabilities.

When OpenCV or thumbnails are unavailable, the worker returns 0.5 quality defaults. When YOLO yields nothing, it returns unknown at 0.2. These are sentinel-like values but are not explicitly typed as “unknown”; never aggregate them as measurements. SourceVisualAnalyzer marks a local-CV failure with local_cv_failed_fallback_..., so gate on analysisMethod, not just numeric fields.

Worked visual example

For 13 detected segments in a 120-second source, the current motion heuristic gives:

cuts_per_minute = (13 - 1) / 120 * 60 = 6
motion_score = 0.25 + 6 / 20 = 0.55

This proves only that the scene detector retained about six boundaries per minute at the configured threshold/minimum duration. It does not prove visible motion or highlight value. Discriminate the causes by reviewing frames around each boundary and a proxy interval on both sides.

Evaluate highlight candidates as evidence

Current CinematicHighlightAnalyzer:

  1. Classifies from project name/style/clip path substrings.
  2. Splits each clip into eight-second windows.
  3. Caps at 12 windows and spreads them through long clips.
  4. Scores min(1, 0.35 + 0.12*(max(.1,sharpness)+max(.1,brightness)+max(.1,motion)) + role + category).
  5. Assigns role from position: first is hook, last is hero, early third is rising action.

Because FfmpegClipInspector initializes all three quality fields to zero and no current enrichment replaces them, quality contributes exactly 0.036. A first car candidate therefore scores 0.35 + 0.036 + 0.25 + 0.20 = 0.836 without inspecting its pixels or sound. Treat the ranking as a coverage scaffold.

For a production candidate evaluator, require each candidate to carry independently measured and normalized features plus their availability masks:

Dimension Discriminating evidence Reject or penalize when
Technical usability sharpness distribution, highlight/shadow clipping, shake, occlusion, audio noise/clipping defect defeats the intended role unless narratively essential
Semantic moment local embeddings/detectors/transcript tied to sampled timestamps evidence does not support the category/action claim
Novelty distance from already selected shots in visual/audio embedding space angle/action is redundant
Story function observable setup, action, reaction, reveal, or payoff role is assigned only from source position
Editability clean handles before/after, usable continuity, compatible motion/audio cut begins/ends mid-action or cannot bridge coherently
Audio value dialogue/event/ambience confidence and intelligibility silence/noise is mistaken for low/high drama

Predict expected metric movement before changing a model or weight. Evaluate on an approved, licensed golden set using timestamp-level relevance, diversity, coverage, false-positive cost, and blinded human pairwise preference. Never tune and certify on the same sources.

Engineer sound instead of merely attaching files

Define the units:

  • dBFS: sample amplitude relative to digital full scale; 0 dBFS is the sample ceiling.
  • LUFS-I: program integrated loudness with perceptual weighting and gating.
  • dBTP: reconstructed true peak, which can exceed sample peak between samples.
  • LRA: loudness range, a statistical measure of loudness variation, not simply max minus min.

SourceAudioAnalyzer reports volumedetect mean/max dBFS and silencedetect ranges at -35 dB for 0.5 seconds by default. It does not report any LUFS, dBTP, LRA, clipping count, speech intelligibility, noise floor, music beats, or semantic events. For missing audio it reports audioPresent=false but numeric mean/max values of 0; never interpret those zeros as loud audio.

Measure a final file read-only:

FINAL=path/to/final.mp4
ffmpeg -hide_banner -nostats -i "$FINAL" \
  -af loudnorm=I=-16:TP=-1.5:LRA=11:print_format=json -f null -
ffmpeg -hide_banner -nostats -i "$FINAL" \
  -af ebur128=peak=true -f null -
ffmpeg -hide_banner -nostats -i "$FINAL" \
  -af silencedetect=n=-35dB:d=0.5 -f null -

The repository defaults target -16 LUFS-I, -1.5 dBTP, and LRA 11, but these are configuration targets, not certified acceptance thresholds. The renderer uses one-pass loudnorm; exact conformance requires measuring the encoded output, and deterministic file normalization normally requires a measured first pass followed by a second pass using the measured fields.

Current mix graph and its consequences

  • Source audio is always referenced as [0:a]; a video-only timeline can make mixing fail.
  • Music now uses the plan cue's gainDb, start, end, trim, and delay. It still needs final-media loudness measurement and a proven scene-fit generation model.
  • Each script line is a distinct voiceover asset, trimmed and delayed to its planned VoiceoverLine range, then combined into the voice bus at unity gain. Generated waveform duration and intelligibility are not yet measured against that slot.
  • Music is sidechain-compressed only when voiceover exists, using configured threshold 0.045, ratio 8, attack 20 ms, release 250 ms by default.
  • SFX are trimmed, gain-adjusted in dB, and delayed to their cue start.
  • amix=duration=first makes the source timeline define output length, then one-pass loudness normalization feeds AAC 192k/48 kHz by default.
  • Audio-mix and preview failures fail the highlight render. The mastering check proves that the successful command contained loudnorm; it does not measure encoded LUFS-I or true peak.

Therefore asset existence and a passed QA JSON are not proof of audible voiceover/music/SFX. Verify stream presence, timing, loudness, true peak, silence, and content audibly and instrumentally. A max sample peak of 0 dBFS signals risk but does not alone prove intersample clipping; use true-peak analysis.

For voice timing, derive speech duration from the generated waveform, not characters / constant. Place each line with adelay, trim/pad deliberately, leave breaths and edit handles, and fail when a line cannot fit its approved slot without unacceptable time-stretch. Duck from the aligned voice bus, not an unaligned full-file narration.

Apply color, LUTs, and overlays safely

Before grading, preserve and inspect transfer, primaries, matrix, range, bit depth, pixel format, and rotation. Current inspection and rendering do not preserve or validate this color contract explicitly. A numeric eq preset is not color management and can damage already-graded, log, HDR, or limited-range footage.

Treat a LUT as executable media data: approve its license and checksum; document expected input/output color spaces; test neutral ramps, skin/brand colors, legal range, banding, and highlight roll-off. Although a LUT folder is configured, neither renderer uses it.

Current overlays use FFmpeg drawtext, 54 px white type, 2 px dark border, and positions around 6% horizontal/8-10% vertical margins. The animation field is validated in the multi-clip path but ignored by both renderers. Fonts are not explicitly selected, so output depends on host font resolution. The highlight QA check rejects out-of-timeline and unsupported placement values, but “safe” placement names are heuristics, not proof against glyph width, multiline wrapping, platform UI, subtitles, crops, or localization.

Require a pinned licensed font, measured text bounds, explicit wrapping, title/action safe policy for target platforms, and raster review at every output aspect ratio. Do not approve an overlay because text_overlays_safe=true; it proves plan bounds, not raster safety.

Make local inference reproducible

For every model runtime, record this minimum manifest next to the approved deployment artifact, not as tribal knowledge:

Field Required content
Identity task, model family, exact revision/weight filename, SHA-256
Provenance authoritative source, retrieval date, license text/version, commercial-use decision, approver
Runtime OS/architecture, CPU/GPU/accelerator, driver, Python, native libraries, framework/wheel hashes
Inference preprocessing, image/audio sample policy, thresholds, seeds, deterministic settings, precision/device
Contract input/output schema, label map/voice identity, expected sample rate/channel/pixel format
Evidence golden-set version/hash, accuracy/quality results, latency, memory, known failure slices

Use platform-specific artifacts for macOS development, Linux/VPS, and cloud images, but prove semantic parity on the same golden inputs. Bit-identical neural outputs across different accelerators may be unrealistic; define allowed numeric/output drift before testing and compare downstream decisions as well as tensors.

Current gaps are blocking for production reproducibility: tools/local_asset_requirements.txt has no versions/hashes; local CV package versions are pinned but not hash-locked; model names are mutable aliases; no model/asset checksum or license manifest is enforced; the runtime can select OS say/espeak; AudioCraft get_pretrained and YOLO name loading can use network/cache state.

Inventory pre-provisioned artifacts without downloading anything:

find input/highlights/assets -type f -print 2>/dev/null | sort
find . -type f \( -name '*.pt' -o -name '*.pth' -o -name '*.onnx' -o -name '*.bin' \) -print
shasum -a 256 path/to/model-or-asset

On Linux, use sha256sum path/to/model-or-asset. A hash proves identity, not license or quality.

Placeholder trap: worked trace

As of 2026-07-21, unavailable AudioCraft/Piper models return failure, the worker does not emit tones, silence, or host speech, LocalAssetSynthesizer deletes failed/inaudible output, and requested missing assets block rendering. The highlight path creates script-specific per-line voice files and delays them to planned reading-time windows before combining the voice bus for ducking. Detecting nonzero PCM is still not speech intelligibility or script-faithfulness validation, so keep approved rendering blocked until semantic audio QA and licensed immutable model provenance exist.

Interpret final output with a measured checklist

Before calling a render production-ready, require all of the following:

  • Source selections trace to timestamped evidence and are diverse, technically usable, and story-relevant.
  • Every plan range is within the real stream timeline and survives decode-boundary inspection.
  • Final streams match an approved codec/container/pixel/audio/color contract.
  • Final duration is probed, not copied from plan arithmetic.
  • Every required asset has approved provenance, checksum, license, model/runtime manifest, and intended content.
  • Voiceover is intelligible, fact-grounded, aligned to its slots, and absent only by explicit creative decision.
  • Music and SFX are aligned, mixed, ducked, and measured; no placeholder or silent fallback exists.
  • LUFS-I, dBTP, LRA, silence, black frames, freeze/duplication, and A/V sync are measured against approved thresholds.
  • Overlays are rendered with pinned fonts and checked for bounds/safe areas on each target.
  • Color treatment is appropriate to the input color space and does not clip or shift protected colors.
  • A blinded human review compares against the approved baseline; heuristic scores cannot waive this review.
  • Promotion follows video-editing-change-control and evidence is retained under the process defined by video-editing-validation-and-qa.

Provenance and maintenance

The implementation claims above come from FfmpegClipInspector, ThumbnailExtractor, ProxyGenerator, WaveformGenerator, ShotSceneSegmenter, SourceAudioAnalyzer, SourceVisualAnalyzer, HeuristicVisualAnalysisProvider, LocalCvVisualAnalysisProvider, CinematicHighlightAnalyzer, HighlightSourceScheduler, HighlightDirectorPromptGenerator, HighlightDirectorFlowService, HighlightVisualEffectsStage, HighlightFfmpegRenderer, FfmpegEditRenderer, EditPlanValidator, LocalAssetRuntimeVerifier, LocalAssetSynthesizer, tools/local_cv_worker.py, tools/local_asset_worker.py, both worker shell scripts, both requirements files, and application.yml.

Re-verify volatile facts with read-only commands:

rg -n 'ffmpeg|ffprobe|sceneDetection|silenceThreshold|loudness|ducking|autoStart|strictRuntime|fallback' src/main/java/org/example/videoclips src/main/resources/application.yml
rg -n 'filter|libx264|crf|loudnorm|sidechaincompress|amix|drawtext|concat|copy_post_timeline' src/main/java/org/example/videoclips/editing/*Renderer.java
rg -n 'Laplacian|brightness|cuts_per_minute|Haar|YOLO|yolov8n|fallback|silence|tone|get_pretrained' tools src/main/java/org/example/videoclips/editing
rg -n 'category.json|highlight-candidates.json|CinematicHighlightAnalyzer' src/main/java/org/example/videoclips/editing
sed -n '1,140p' src/main/resources/application.yml
sed -n '1,120p' tools/local_cv_requirements.txt
sed -n '1,120p' tools/local_asset_requirements.txt
ffmpeg -version
ffprobe -version
ffmpeg -hide_banner -filters | rg 'scene|loudnorm|ebur128|xfade|sidechaincompress|lut3d|drawtext'

If any command, property, path, formula, or wiring changes, update this skill in the same approved change. Date-stamp the new verification and preserve the distinction between implemented behavior, measurement theory, candidate improvements, and certified evidence.