The vision director now captions several beat frames (two questions per frame in one worker call: a discriminative description + a punchy label) and turns the descriptions into a per-window "highlight-worthiness" curve via generic emotion/action/idle keyword scoring (semanticScore/semanticCurve). The montage director blends that curve with audio to place the payoff on the semantically strongest moment; the payoff label becomes the bold overlay and the description flavors the music. Everything is content-agnostic and fails soft to the measured cut. Verified on bowling: the payoff moved onto moondream's detected celebration. Honest limit: a small VLM on distant subjects is only weakly discriminative; descriptive questions beat terse ones (which collapse to a constant answer). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MPuJXQyAeWpFcTtcnxo1UN |
||
|---|---|---|
| .claude | ||
| .idea | ||
| dashboards | ||
| docs | ||
| input/source | ||
| output | ||
| src | ||
| tools | ||
| .gitignore | ||
| pom.xml | ||