video_editing_poc/docs/cinematic-highlight-poc-pla...

4.2 KiB
Raw Blame History

Local Cinematic Highlight PoC — Plan & Milestones

Owner: (Jaden) · Started: 2026-07-21 · Machine: Intel Mac (x86_64, 16 CPU, 32 GB, no GPU)

Objective

Produce ONE genuinely cinematic highlight from a local source video, end-to-end through the existing pipeline, using only local models resident in the service runtime: Piper (voiceover) + MusicGen (music) + AudioLDM2 (SFX), with deliberate visual treatment and pacing, and explicit human approval before any render.

A valid MP4 or a passing test is NOT success. Success = representative output passes measured media QA and a human creative review.

Non-negotiables (still in force during the PoC)

No external AI services in the media path · no placeholder silence/tones/OS say as a finished asset · no unlicensed assets · no rendering without explicit approval for the specific project · no inference-time network (models load from local dirs, HF_HUB_OFFLINE=1). NOTE: the user authorized model DOWNLOADS on 2026-07-21 (one-time provisioning), reversing the earlier no-download stance; inference stays offline.

Proven capability baseline (2026-07-21)

Model Runtime Status Evidence
Piper en_US-lessac-medium piper-tts 1.5.0 works real speech, 22.05 kHz
MusicGen musicgen-small transformers 4.44.2 + torch 2.2.2 works 5 s / 44.7 s CPU, mean 18 dB, 32 kHz
AudioLDM2 cvssp/audioldm2 diffusers 0.30.3 works 3 s / 22.3 s CPU, mean 21.8 dB, 16 kHz

Environment traps (see memory local-model-runtime-intel-mac): torch capped at 2.2.2 (Intel-Mac), numpy<2, transformers must be 4.x, audiocraft/xformers unusable here, HF_HUB_DISABLE_XET=1 for downloads.

Plan

Phase 1 — Generate all three assets THROUGH the pipeline

  • 1.1 Rewrite tools/local_asset_worker.py: music→transformers MusicGen, sfx→diffusers AudioLDM2 (resample to 48 kHz mono), voiceover→Piper (unchanged). audiocraft path removed. CLI + exit codes preserved.
  • 1.2 Materialize models into stable models/ dirs via save_pretrained (tools/provision_local_models.py): models/musicgen-small (2.2 G), models/audioldm2 (4.2 G); models/piper (60 M) already present.
  • 1.3 License/provenance .license.txt sidecars for each model marker file (config.json / model_index.json / voice .onnx). BOTH audio models are NON-COMMERCIAL (MusicGen CC-BY-NC-4.0, AudioLDM2 CC-BY-NC-SA-4.0) — flagged for production review. Honors the AssetLicensePolicy regular-file gate; no Java change.
  • 1.4 Smoke test: voiceover 3.67 s/15.8 dB, music 4.94 s/12.0 dB, sfx 3.00 s/20.6 dB — all 48 kHz mono, real signal, exit 0.
  • 1.5 mvn -o verify → 245 tests / 62 classes / 0 failures / 0 errors / 0 skips (unchanged from baseline).

Phase 1 COMPLETE (2026-07-21).

Phase 2 — One approved end-to-end highlight (STOP before render for explicit approval)

  • 2.1 Opt-in localpoc Spring profile (base/production defaults untouched): venv python, model paths, offline flags, no bootstrap auto-start, heuristic fallback off, isolated PoC I/O dirs, render disabled.
  • 2.2 Reprocess the DJI source so category.json + highlight-candidates.json are written.
  • 2.3 Hand-author a director edit-plan.json grounded in persisted candidates (manual, local-only).
  • 2.4 Explicit approval for the specific project → render → final + manifest + QA report.
  • 2.5 Measure technical QA + structured human creative review.

Phase 3 — Iterate to quality

  • Improve selection, pacing, visual treatment, voice, music/SFX fit; keep experiments reproducible.

Milestone log

  • 2026-07-21: Models provisioned & individually proven (Piper, MusicGen, AudioLDM2). Plan approved. Phase 1 started.
  • 2026-07-21: Phase 1 complete. Worker rewritten (audiocraft→transformers MusicGen + diffusers AudioLDM2), models materialized to models/ with license sidecars, all 3 asset kinds generate 48 kHz mono real audio through the worker, mvn -o verify green (245/0/0/0). Paused for review before Phase 2.

Deferred (until output-quality gate passes)

Production hardening: Spring Security/OIDC, PostgreSQL/Testcontainers, containers/K8s, CI/CD, distributed ops, digest-bound authenticated approval. Recorded, not deleted.