Make the stronger Tier-2 judge the localpoc default (vision-caption-script ->
vision_caption_llamacpp.py). Guard it so it never silently degrades: resolveCaptionScript
checks the llama.cpp backend is provisioned (binary env + weights) and, if not, falls
back to moondream with a WARN (event=vision_backend_not_ready). Worker defaults the
model/mmproj to the repo's ./models/qwen2.5-vl-3b paths, so only the machine-specific
binary env is mandatory. mvn verify: 294 tests green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Beats now blend instead of hard-cutting (incl. the cut into the slow-mo payoff)
via an xfade + acrossfade chain (xfadeTimelineCommand). The timeline compresses
by (n-1)*xf, so shiftOverlayForCrossfade re-times overlays and the reported
duration is reduced to keep overlays, loudness mastering, and QA aligned. Opt-in
via editing.crossfade-seconds (0 = hard cuts default; localpoc 0.25), clamped to
half the shortest beat. Offset math + overlay shift unit-tested; verified on the
bowling cut (visible dissolve, STRIKE overlay stayed on the payoff).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MPuJXQyAeWpFcTtcnxo1UN
Rendering (whole service, driven by source measurements, not constants):
- Orientation-aware geometry: FfmpegClipInspector reads display rotation and
stores effective dims; HighlightFfmpegRenderer.outputGeometry renders portrait
sources portrait and skips the 2.39 letterbox on portrait (landscape unchanged).
- Dynamic exposure: probeSourceLuma measures the frames; exposureNormalizationFilter
maps the mean toward a target; the grade is now exposure-preserving (no crushed
subjects: bowling final went ~63 -> ~100 mean luma).
- Motion-adaptive in-shot push-in (zoompan), amount from per-shot YDIF.
- Audio mix ducks source audio under the generated score so it leads.
- Highlight duration is no longer capped (validator + config).
Automatic director (plans were hand-authored before):
- Tier 1 HighlightMontageDirector: composes the montage from measured motion (YDIF)
and audio-energy (RMS) curves -- setup, continuous action/tension, slow-mo payoff
on the audio climax, resolution button, camera-whip tail trimmed.
- Tier 2 HighlightVisionDirector + tools/vision_caption.py: a local, offline
vision-language model (moondream2) captions the payoff frame and augments the
montage with a semantic overlay ("STRIKE") and scene-informed music; fails soft.
- Wired into the scheduler behind auto-director-enabled / vision-director-enabled
(on in the localpoc profile).
Docs: cinematic-quality-rules.md (R1-R5, R9 both tiers), poc-plan milestones.
Tests: mvn -o verify -> 262 passing, 0 failures.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MPuJXQyAeWpFcTtcnxo1UN
Activated with --spring.profiles.active=localpoc. Points only at pre-provisioned
local model paths (Piper voice, MusicGen, AudioLDM2), disables bootstrap
auto-start, uses heuristic visual analysis, isolates PoC input/output dirs, and
keeps render disabled + director approval required. Base/production defaults are
untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bg76sLc43Wc3j5ZcLkboYR