video_editing_poc/.claude/skills/video-editing-cinematic-hig.../SKILL.md

33 KiB

name description
video-editing-cinematic-highlights-campaign Load this skill when planning, implementing, debugging, or promoting the single-source cinematic highlight campaign: local-only highlight identification and editing with production visuals, music, SFX, and voiceover; especially for candidate generation, an offline director, model packaging, director-plan safety, measured media QA, resumability, source-to-final tests, or macOS/Linux parity.

Cinematic Highlights Campaign

Status verified against the repository on 2026-07-21. Treat every path or component marked [PROPOSED] as work that does not exist yet. Introduce it only through video-editing-change-control.

Use this campaign correctly

Use this skill to drive the hardest live problem: turn one source video into correctly selected, high-cinematic-quality highlights with production visuals, music, sound effects (SFX), and voiceover, using only pre-provisioned local models resident in the service runtime.

Do not use it for:

Need Use instead
A routine, already-understood patch video-editing-change-control
A symptom with no established root cause video-editing-debugging-playbook
Media theory or FFmpeg fundamentals cinematic-media-engineering-reference
General test evidence or existing test inventory video-editing-validation-and-qa
Existing measurement commands and artifact inspection video-editing-diagnostics-and-tooling
Post-triage causal proof or first-principles derivation video-editing-proof-and-analysis-toolkit
Environment setup or normal operation video-editing-build-and-env or video-editing-run-and-operate

Never use this campaign to route around approval, licensing, security, or change control.

Terms and hard constraints

Term Meaning here
Certified fixture Rights-cleared source media plus immutable annotations, checksum, provenance, and an assigned development or locked-holdout split.
Temporal IoU (tIoU) Intersection duration divided by union duration for a predicted and annotated highlight interval.
Recall@3 Fraction of annotated highlights matched by one of the top three predictions at the stated tIoU.
Macro-F1 Unweighted mean of per-category F1 scores, so a large category cannot hide a weak one.
nDCG@3 Normalized discounted cumulative gain for the top three ranked candidates; higher relevance near rank one receives more credit.
SLO Service-level objective: an approved target for a measured service-level indicator over a stated workload/window.
SBOM Software Bill of Materials: an inventory of shipped software components and their identities.
Fail closed Stop before rendering or promotion when required evidence, models, assets, approval, or QA is missing.
Local model A model whose weights, loader, tokenizer, and runtime dependencies are pre-provisioned in the deployed artifact or mounted runtime bundle; loading it performs no network operation.
Creative rubric A blinded 0-4 human score for story, selection, pacing, visual craft, sound design, voiceover, and factual grounding.

The no-waiver rules are authoritative in video-editing-change-control; these are their campaign consequences:

  • Make no automatic dependency or model downloads. Do not call external AI services.
  • Make no network connection during the certified offline run, including loopback and model resolution. The current CV worker uses loopback HTTP; replace that transport under change control before certification.
  • Use only assets with recorded licenses and permitted production use.
  • Produce no placeholder silence, tones, generic assets, or unrequested fallback render.
  • Render only after explicit approval. Do not change packaged production-facing defaults.
  • Do not treat a successful FFmpeg exit or a JSON field set to true as quality evidence.
  • Keep every experiment isolated behind an explicit, disabled-by-default campaign control. Promotion still follows video-editing-change-control.

Current truth: do not design from the plan documents

Repository fact Evidence and consequence
HighlightSourceScheduler calls HighlightCandidateGenerator after source analysis and before HighlightDirectorPromptGenerator. The single-source flow now persists analysis/category.json and analysis/highlight-candidates.json; missing files mean candidate generation failed or the project predates this change.
HighlightCandidateGenerator is a deterministic baseline; CinematicHighlightAnalyzer remains the older multi-clip ranker. The single-source generator ranks shot/coverage windows from source-level visual scores, FFmpeg silence sections, scene score, duration, and position. It explicitly downweights heuristic/fallback visual evidence; do not call it semantic highlight understanding.
No local director executes the generated highlight prompt. director/director-brief.md tells an operator to use Codex, Claude, or another agent; rendering begins only when director/edit-plan.json appears.
The highlight plan has a dedicated HighlightDirectorPlanValidator; it does not use the multi-clip EditPlanValidator. It requires persisted category/candidates, safe unique IDs, candidate-contained time ranges, configured duration bounds, and nonblank visual/music/SFX/voiceover direction before asset work.
Plan-derived highlightId and mutable asset-request paths are validated. HighlightLocalAssetWorker requires exact project/highlight identity, supported types, safe keys, bounded duration, contract-exact request/target paths, project containment, and no symlink component before materialization. Adjacent nonblank asset-license sidecars are required and preserved; authenticity, checksum, origin, allowed-use validation, and request-file integrity signatures remain open.
Standalone bootstrap scripts remain network-capable. Both worker shell scripts run pip install in auto mode; the CV script defaults to yolov8n.pt. LocalAssetRuntimeVerifier no longer invokes the asset bootstrap, but operators must still not run either launcher in a certified environment.
Strict asset readiness now fails startup. The 2026-07-21 working tree requires pre-provisioned Piper, MusicGen, and AudioGen paths and throws when strict readiness is incomplete. No such three-model bundle is present in this repository.
Placeholder asset success was removed. The Python and Java asset workers now return failure and delete invalid output instead of emitting OS speech, silence, or tones. Missing requested music, SFX, or voiceover blocks rendering.
Audio and preview failures now fail the highlight render. The renderer no longer copies the pre-mix timeline or final output after those failures. Prepared SFX keys and rendered SFX keys are shared.
Highlight QA now includes selected media probes but is not a creative-quality gate. Final duration, black ranges, long silence, and sample peaks are probed. Integrated loudness, true peak, A/V sync, freeze detection, raster safe area, asset semantics, and human creative acceptance remain open.
The source-to-final test is incomplete. HighlightDirectorFlowServiceTest uses a mocked renderer and dummy files. CinematicEditingIntegrationTest covers the separate multi-clip flow.
Processing is not stage-resumable. The scanner skips only projects with root final.mp4; there is no per-stage journal, input digest, lease, or verified checkpoint.

Reconfirm those claims before opening a campaign change:

rg -n "HighlightCandidateGenerator|category.json|highlight-candidates.json" \
  src/main/java/org/example/videoclips/editing/{HighlightSourceScheduler,HighlightCandidateGenerator,HighlightDirectorPromptGenerator}.java
rg -n "planValidator|HighlightDirectorPlanValidator|highlightId\(\)|targetPath\(\)" \
  src/main/java/org/example/videoclips/editing/{HighlightDirectorFlowService,HighlightDirectorPlanValidator,HighlightLocalAssetWorker}.java
rg -n "pip install|get_pretrained|write_silence|fallbackTone|strict_runtime_degraded|duration_matches_timeline" \
  tools src/main/java/org/example/videoclips/editing

Expected on the 2026-07-21 working tree: matches in all three commands; explicit scheduler wiring to HighlightCandidateGenerator; no placeholder-success implementation in the asset synthesizers; network-capable launcher matches remain non-zero. Request target strings remain in the code but are now checked against exact project-local paths before use. If that changes, update this campaign before proceeding.

Campaign gate map

Do not start a later phase until the preceding gate is green.

Phase Deliverable Numeric gate
0 Frozen offline baseline 245 tests in 62 test classes; the full offline Maven suite has 0 failures, 0 errors, and 0 skips
1 Certified fixture and annotation set 4 current categories represented; 2 independent annotations per source; 100% checksums/licenses/splits present
2 Single-source category and candidate wiring Macro-F1 and Recall@3 thresholds calibrated and recorded; every processed fixture emits exactly 1 category file and 1 candidate file
3 Resident local director and safe plan 0 network operations; 100% schema-valid plans; all adversarial plans rejected before filesystem/render work
4 Fail-closed model and asset runtime 100% manifest/checksum/license/canary pass; any one missing capability prevents readiness and render
5 Production render semantics Every requested effect/audio/VO action either has execution evidence or fails the render; 0 silent degradations
6 Measured technical and creative QA 0 ERROR checks; creative average >=3.5/4 and no dimension <3; paired-preference lower 95% bound >50%
7 Resumability and real source-to-final test 1 final per requested highlight, 0 duplicate side effects, and exact recovery after every injected stage interruption
8 Offline platform, security, performance, operations macOS and Linux/VPS outputs pass identical gates; 0 attempted network calls; 0 path escapes; workload SLOs met
9 Controlled promotion All change-control evidence present; human reference-architecture rubric has no score <3 and average >=3.5

Phase 0: freeze an offline baseline

  1. Work from the repository root. Record, but do not mutate, source and tool state:
git status --short
git rev-parse HEAD
java -version
mvn -version
ffmpeg -version | head -n 1
ffprobe -version | head -n 1
printf 'test_methods='; rg -n '^\s*@Test\b' src/test/java | wc -l
  1. Run the current tests that do not invoke the network-capable model/asset bootstrap paths:
mvn -q -o -Dtest=HighlightSourceSchedulerTest,HighlightDirectorPromptGeneratorTest,HighlightDirectorFlowServiceTest test

Gate 0 expectations for the current working tree: Maven exits 0. Treat generated test media as disposable, not certified evidence. A 2026-07-21 mvn -q -o verify run executed all 245 tests in 62 test classes with zero failures, errors, or skips and passed the bound JaCoCo gate. Maven offline mode does not prevent subprocess networking, so call this offline dependency-resolution evidence, not no-egress certification.

  • If offline Maven reports a missing artifact, branch to video-editing-build-and-env; pre-provision a locked artifact repository or build cache. Do not enable network.
  • If a test fails because local voice/TTS behavior differs, branch to Phase 4. The current test suite can consume an untracked local environment and is not yet clean-checkout reproducible.
  • If the worktree is dirty, preserve unrelated changes. Record their paths and isolate campaign evidence; never reset them.

Phase 1: create certified fixtures and annotations

The repository has no certified creative golden inventory. Do not promote existing input/ or output/ media to goldens by assumption.

Under change control, add [PROPOSED] fixture metadata and small rights-cleared media, or an authenticated internal fixture-bundle mount whose immutable digest is recorded. The metadata must contain: fixture ID, SHA-256, byte size, duration, codecs, category, rights/license ID, permitted use, provenance, split, and annotation version. Keep weight files and large/proprietary footage out of Git.

For each source, collect two independent annotations containing category, positive highlight intervals, excluded intervals, story role, visible facts allowed in voiceover, unsafe/blurred frames, speech intervals, and music/SFX intent. An adjudicator resolves disagreement without seeing model output. Lock the holdout split before tuning.

Starting floor (a campaign minimum, not a SOTA claim): at least 12 sources, at least 3 each for car_vlog, food_vlog, family_vlog, and generic_vlog; at least one no-audio and one low-quality negative source; 2 annotators per source. Record a larger statistically powered sample before making external quality claims.

Gate 1 [PROPOSED] command; implement a repository script only through change control:

test -x tools/validate-highlight-fixtures && tools/validate-highlight-fixtures --offline --strict

Expected: exit 0, sources>=12, categories=4, annotators_per_source>=2, missing_checksums=0, missing_licenses=0, split_overlap=0. If the script is absent, stop: do not replace it with visual inspection. If rights are unclear, quarantine the fixture and branch to legal/ asset governance.

Phase 2: wire category and candidate generation

Implement a single-source application use case between HighlightSourceAnalyzer.analyze and prompt generation. It must consume persisted source, scene, audio, and visual analysis; classify using actual local evidence; generate shot-boundary-aligned candidates; and atomically write exactly:

  • analysis/category.json
  • analysis/highlight-candidates.json

Do not merely copy the metadata-keyword classifier from the multi-clip path. First define a deterministic heuristic baseline, then compare local learned scorers against it. Candidate IDs must be stable for the same source/config/model digests. Durations must honor the configured 8-35 second range; at most the configured top 3 proceed to planning. Uncertain classification is a review state, not permission to render generic content.

Before running, predict category confusion, candidate Recall@3 at tIoU>=0.5, and ranking nDCG@3 on development fixtures. Record predicted numbers, then run once. Set final thresholds from baseline plus confidence intervals in an ADR; until that ADR exists, use the campaign floors: macro-F1 >=0.85, Recall@3 >=0.80, nDCG@3 >=0.80, and invalid/out-of-range candidates =0. These are candidate promotion thresholds, not current results.

Gate 2 [PROPOSED]:

mvn -o -Dtest=HighlightSourceCandidatePipelineTest,HighlightCandidateEvaluationTest test
test -x tools/evaluate-highlight-selection && \
  tools/evaluate-highlight-selection --fixtures certified --split holdout --offline

Expected: both exit 0; one category and one candidate artifact per source; the four numeric floors above pass. If category passes but Recall@3 fails, branch to segmentation/fusion experiments. If Recall passes but nDCG fails, keep candidates and work only on ranking. If all offline metrics pass but human selected moments remain weak, inspect annotation agreement before changing the model.

Phase 3: implement a resident local director and validate plans

Replace the operator-mediated external-agent step with an explicit application port and a local adapter. The adapter must use a pre-provisioned model path, constrained JSON decoding, fixed schema version, bounded context/output, deterministic seed where supported, timeout, cancellation, and no socket or model-hub resolution. Preserve human approval; the model may propose but never approve.

Validate HighlightDirectorPlan before creating storyboard, asset requests, or directories:

Required validation Reject when
Identity Project/source mismatch; unknown candidate; duplicate/unsafe highlight ID
Numeric safety NaN/infinity; negative/reversed/out-of-source time; target outside configured range
Structural bounds Zero highlights; more than configured maximum; empty/oversized strings or arrays
Semantics Unknown category, story role, effect, transition, render profile, voice, or asset type
Grounding Voiceover asserts a fact absent from certified visible/transcript evidence
Path safety Absolute path, .., separator, symlink escape, or caller-supplied target path
Asset closure A requested music/SFX/voice/font/LUT capability has no licensed local resolution

Derive target paths server-side from validated IDs. Resolve against the project root, normalize, verify containment, reject symlinks at trust boundaries, and create through the store. Make the JSON schema version explicit and validate it during the Maven build.

Gate 3 [PROPOSED]:

mvn -o -Dtest=LocalHighlightDirectorContractTest,HighlightDirectorPlanValidatorTest,HighlightPathSafetyTest test

Expected: exit 0; valid plans accepted =100%; malformed, unknown-field, timestamp, NaN/infinity, duplicate-ID, ../, absolute-path, backslash, symlink, oversized-text, unknown-effect, and missing-asset cases rejected =100%; filesystem writes after a rejected plan =0; network operations =0.

If constrained decoding still emits invalid JSON, do not repair it silently: record the failure and stop. If factual grounding cannot be established, omit the line or route to human review.

Phase 4: package and preflight all local capabilities

Add [PROPOSED] typed configuration for a runtime model root and a signed/immutable model manifest; do not overload the current free-form model-name properties. For every CV, ASR, director, TTS, music, and SFX model record: purpose, relative path, exact bytes, SHA-256, loader/runtime version, tokenizer/ config paths, license identifier and license-text digest, approved use, attribution, memory estimate, and a deterministic canary input/output predicate.

Pre-provision Python wheels and weights in the build/deployment supply chain. Use locked versions and hashes. Production startup must never run pip, YOLO("yolov8n.pt"), or get_pretrained("facebook/..."). A readiness preflight must verify regular-file containment, checksums, licenses, loader compatibility, capacity, and one inference canary for every enabled capability. Any failure makes readiness false and blocks project claims and rendering.

Gate 4 [PROPOSED]:

test -x tools/verify-local-model-bundle && \
  tools/verify-local-model-bundle --manifest "${VIDEO_EDITING_MODEL_MANIFEST:?set manifest}" --offline --strict
mvn -o -Dtest=LocalModelManifestTest,OfflineCapabilityPreflightTest,FailClosedAssetGenerationTest test

Expected: models_checked>=6, checksum_mismatches=0, license_failures=0, canary_failures=0, network_attempts=0; Maven exits 0. Re-run after corrupting one byte in a disposable test copy: expected preflight exit non-zero, readiness false, renders started =0. If a license is absent or ambiguous, the model/asset is unavailable, not experimental.

Remove production paths to say, espeak, tones, and silence. Keep negative tests proving these cannot satisfy a required asset. Do not edit packaged defaults during the campaign; activate new components only with explicit environment configuration in approved environments.

Phase 5: make render instructions real

Implement an allowlisted renderer capability registry. A plan may name only behavior the renderer can prove it executed. Today, any non-empty visual treatment maps to the same crop/grade and fades are per-segment. Audio and preview failures now propagate; keep regression tests on that fail-closed behavior. Do not relabel the remaining generic visual behaviors as cinematic.

Plane Required semantics and evidence
Selection/edit Multiple decisions when the story requires them; source/timeline arithmetic verified; true inter-shot transitions rather than independent fades
Visual Explicit color transform, stabilization, reframing, speed ramp, overlay/font, and transition parameters; command/filter evidence plus output measurements
Music Licensed/generated asset digest; deliberate in/out points, trim/loop policy, gain envelope, and final mix contribution
SFX One existing local asset per cue; sample-accurate or measured timing; no missing-path attempt
Voiceover One grounded script version, approved local voice/model, line timing/alignment, intelligibility measurement, and ducking evidence
Master Two-pass or otherwise measured loudness normalization; output codec, geometry, cadence, duration, and channel layout probed from the file

Make every requested production asset blocking. On FFmpeg, preview, or mix failure, mark the attempt failed and publish no final.mp4. Write to a temporary name and atomically promote only after QA. Record tool versions, complete argv, exit codes, input/output hashes, models/assets/licenses, timing, and effective config in render-manifest.json.

Gate 5 [PROPOSED]:

mvn -o -Dtest=HighlightRendererIntegrationTest,HighlightAudioMixIntegrationTest,HighlightEffectContractTest test

Expected: exit 0; requested actions = executed actions; unresolved actions =0; silent degradations =0; failure injection for each FFmpeg stage publishes final files =0. If a filter is unsupported on macOS or Linux, do not substitute one; branch to Phase 8 and either standardize FFmpeg or remove the capability through change control.

Phase 6: replace asserted QA with observed QA

Each QA check must include expected value, observed value, unit, tool/algorithm version, command or method, artifact path/digest, severity, and pass/fail. Probe the actual final file. At minimum measure: stream presence, duration delta, 1920x1080 geometry, 30 fps cadence, H.264/AAC codecs, 48 kHz audio, black/freeze spans, silence spans, clipping/true peak, integrated loudness, overlay bounds, asset contribution, voice intelligibility/alignment, and requested-effect execution.

Use the current configured mastering targets as initial technical gates: integrated loudness -16 LUFS +/-1, true peak <=-1.5 dBTP, duration error <=0.25 s, missing requested streams/assets =0, and ERROR checks =0. Calibrate black/silence/freeze tolerances by content and annotations; do not apply a universal threshold that rejects intentional black or quiet moments.

Operator spot-check commands (diagnostic only, not a replacement for automated parsing):

FINAL="${FINAL:?set final mp4}"
ffprobe -v error -show_entries format=duration:stream=index,codec_type,codec_name,width,height,r_frame_rate,sample_rate,channels -of json "$FINAL"
ffmpeg -hide_banner -i "$FINAL" -vf "blackdetect=d=0.25:pix_th=0.10" -af "silencedetect=n=-45dB:d=1.0" -f null -
ffmpeg -hide_banner -nostats -i "$FINAL" -filter_complex "ebur128=peak=true" -f null -

Gate 6 technical command is [PROPOSED]; implement the parser through change control rather than copying console text into qa-report.json:

test -x tools/verify-highlight-render && \
  tools/verify-highlight-render --final "$FINAL" --manifest "${MANIFEST:?set render manifest}" --strict

Expected: exit 0, error_checks=0, missing_requested_assets=0, duration_error_seconds<=0.25, integrated_lufs in [-17,-15], and true_peak_dbtp<=-1.5. If the tool is absent, Gate 6 is red.

Run blinded creative review against the frozen deterministic baseline. Use at least 3 reviewers per output; score the seven rubric dimensions 0-4; randomize A/B order; retain disagreements. Gate: average >=3.5, every dimension median >=3, factual errors =0, and the 95% confidence interval lower bound for preference over baseline >50%. If technical QA passes but creative QA fails, do not tune FFmpeg blindly: branch according to reviewer tags (selection -> Phase 2, story/VO -> Phase 3, craft/audio -> Phase 5).

Phase 7: add resumability and a real source-to-final test

Add a durable stage journal with state, attempt, lease owner/expiry, started/completed timestamps, code/config/model/input digests, output paths/digests, and failure classification. Required stages: claim, analyze, classify, rank, direct, validate, approve, prepare-assets, render, QA, publish. Commit each checkpoint atomically. Resume only when all input digests match; otherwise invalidate that stage and every dependent stage. Never infer completion from file existence alone.

Build a real source-to-final test using a tiny rights-cleared fixture and pre-provisioned test models/ assets. Do not mock the analyzer, director, asset generator, renderer, FFmpeg, or FFprobe. Test restart after every stage, duplicate scans, stale leases, corrupt checkpoints, changed model digest, invalid plan, missing asset, FFmpeg failure, and crash between temporary output and atomic publish.

Gate 7 [PROPOSED]:

mvn -o -Dtest=HighlightSourceToFinalIT,HighlightWorkflowResumeIT,HighlightWorkflowFailureIT test
mvn -o verify

Expected: exit 0; final outputs equal requested approved highlights; duplicate renders/uploads =0; published partial files =0; each injected interruption resumes at the first invalid/incomplete stage; changing any input digest recomputes all dependent stages. If the test uses a mock renderer or dummy text as MP4, it does not satisfy this gate.

Phase 8: prove offline parity, security, performance, and operations

Run the same immutable fixture bundle, model bundle, service artifact, config digest, and FFmpeg build on macOS development and Linux/VPS. Cloud deployment may mount the same pre-provisioned bundle but must not fetch it at application startup. Container work is [PROPOSED]; the repository has no current container definition, so do not claim --network none validation until one is approved.

Define the representative workload before measuring: source minutes, resolution/codecs, highlight count, concurrent jobs, hardware/accelerator, warm/cold state, and storage. Record per-stage p50/p95/ p99, real-time factor (wall seconds/source seconds), CPU, accelerator, RSS, disk peak, queue depth, failures, and output bytes. Establish SLOs from measured capacity and business demand; never import the 200-RPS synchronous-API example into a media pipeline.

Gate 8 [PROPOSED]:

test -x tools/run-offline-highlight-certification && \
  tools/run-offline-highlight-certification --fixtures certified --models "${VIDEO_EDITING_MODEL_MANIFEST:?}" --strict
mvn -o -Dtest=HighlightOfflineSecurityIT,HighlightCapacityIT,HighlightGracefulShutdownIT test
mvn -o verify

Expected on both platforms: technical and creative Gate 6 pass; network_attempts=0; path escapes and symlink escapes =0; checksum/license failures =0; orphan child processes =0; corrupt or partial finals =0; measured p95 and resource peaks remain within the approved workload SLO. Compare semantic decisions and QA outcomes, not bit-identical encoded video unless the codec/toolchain is fully standardized. If platform results diverge, capture tool/model/config digests and stop promotion.

Security review must threat-model malicious media, decompression/resource exhaustion, hostile model output, path traversal/symlinks, command injection, poisoned weights, unlicensed assets, prompt injection in transcripts/metadata, denial of service, and sensitive media/log leakage. FFmpeg commands remain argv lists, never shell strings. Run workers with least privilege, bounded CPU/memory/time/disk, read-only model mounts, project-scoped writable storage, and no network capability.

Operations must expose stage duration/failure, queue age/depth, active jobs, model readiness, candidate counts, approval waits, render real-time factor, QA failures, disk pressure, and worker restarts without source filenames, prompt text, or project IDs as metric labels. Define alerts, retention, cancellation, graceful shutdown, recovery, and model rollback before production.

Ranked solution menu

Choose with measured evidence, not model popularity.

Rank Candidate Use when Required derivation and disproof
1 Local multimodal director over deterministic shot/audio candidates Resident model fits capacity and materially improves grounded story decisions Predict memory, latency, Recall@3-to-plan loss, grounding error, and creative preference. Disprove with locked holdout, ablation without visual/audio inputs, adversarial transcripts, and local-only trace.
2 Local feature models + deterministic ranker + smaller constrained text director A full multimodal director is too slow or weakly grounded Derive feature normalization, fusion weights, calibration, context budget, and failure states. Ablate each feature; require one mechanism to explain positive and negative fixtures.
3 Local learned temporal scorer + template/constrained planner Director variance blocks schema safety but selection can be learned Prove scorer improves Recall/nDCG and templates improve repeatability without lowering blinded creative scores. Reject if category styles collapse to generic edits.
4 Deterministic heuristics and fixed renderer Establish the baseline and retain an emergency diagnostic comparator only Specify every weight/rule and sensitivity. Never promote as “high cinematic quality” without independently passing all creative gates.

For any chosen option, the decision record must state: requirement; selected approach; alternatives; benefits/trade-offs; operational consequences; security implications; verification; revisit condition. No new database, queue, service, framework, or model is justified merely by this campaign.

Known wrong paths: fence these off

  • Do not use filename/project-style keywords as proof of visual content category.
  • Do not ask an external agent to write the plan and call the workflow local-only.
  • Do not let model libraries resolve names such as yolov8n.pt or facebook/musicgen-small at runtime.
  • Do not interpret a non-strict readiness log as certification. Strict mode now fails startup, but its sidecar and file-presence checks do not establish provenance authenticity, checksums, allowed use, or quality.
  • Do not accept silence, sine tones, OS say/espeak, missing optional assets, or source-audio-only copies when the approved plan requests production audio.
  • Do not use EditPlanValidator unchanged for highlight projects; it reads the multi-clip store and analysis contract. Reuse validation concepts, not the wrong persistence contract.
  • Do not resolve untrusted plan IDs or request target paths directly with Path.resolve/Path.of.
  • Do not add more prompt adjectives to fix missing analysis, weak selection, or unsupported effects.
  • Do not call a fixed 4% crop a dynamic crop, per-segment fades a crossfade, or hard-coded true QA.
  • Do not tune on the locked holdout, promote unlabeled output, or judge cinematic quality by eye.
  • Do not retry permanent validation/model/license failures; do not resume from file existence alone.
  • Do not alter packaged production defaults to expose campaign behavior.

Phase 9: promotion protocol

Route every implementation and final promotion through video-editing-change-control. Promotion requires one evidence bundle containing baseline commit/tool digests; fixture/annotation version and rights; model/dependency SBOM, hashes, licenses, and signatures; predictions made before experiments; selection/director/renderer/QA results including negatives and ablations; offline network-denial proof; macOS/Linux parity; performance/capacity; threat model; recovery/graceful-shutdown evidence; human review data; runbook/alerts; rollback; and ADRs.

The project is not campaign-complete until every gate passes on the locked holdout, the Fortune 500 human-review rubric has no category below 3 and average at least 3.5, security/testing/operational readiness are each at least 3, and no critical behavior relies on an undocumented assumption. A model upgrade restarts Phases 4, 6, 7, and 8. A schema/renderer change restarts Phases 3, 5, 6, 7, and 8.

Provenance and maintenance

Primary code: HighlightSourceScheduler, HighlightSourceAnalyzer, HighlightDirectorPromptGenerator, HighlightDirectorPlanScanner, HighlightDirectorFlowService, HighlightDirectorPlan, HighlightLocalAssetWorker, LocalAssetSynthesizer, LocalAssetRuntimeVerifier, HighlightFfmpegRenderer, FileSystemHighlightProjectStore, and the two worker scripts. Planning documents are context, not implementation truth.

Re-verify volatile facts in one line each:

rg -n '^\s*@Test\b' src/test/java | wc -l
rg -n "category.json|highlight-candidates.json" src/main/java/org/example/videoclips/editing
rg -n "pip install|get_pretrained|write_silence|fallbackTone|strict_runtime_degraded" tools src/main/java/org/example/videoclips/editing
rg -n "duration_matches_timeline|required_assets_resolved|audio_mastering_applied" src/main/java/org/example/videoclips/editing/HighlightFfmpegRenderer.java
mvn -o -Dtest=HighlightSourceSchedulerTest,HighlightDirectorPromptGeneratorTest,HighlightDirectorFlowServiceTest test

Update the dated truth table, baseline counts, thresholds/ADR links, fixture inventory, model manifest contract, and platform/SLO evidence whenever those commands or campaign decisions change.