forked from jsl/video_editing_poc
The director's vision JUDGE is only as good as its model, and moondream2 can't perceive some actions on hard footage (R15 ceiling). Make the captioner backend a config choice so a stronger local VLM drops in with no code change: - tools/vision_caption_llamacpp.py: same manifest->JSON contract as tools/vision_caption.py, but backed by llama.cpp `llama-mtmd-cli` (GGUF). Runs on this x86 CPU via AVX and bypasses the torch==2.2.2 / transformers 4.x trap entirely (no PyTorch). Model/mmproj/binary paths come from env vars; fully offline, serverless (per-frame CLI, mmap stays warm). - editing.vision-caption-script selects the worker (default: moondream). The Java HighlightVisionDirector now reads the configured script instead of a hardcoded path -- nothing else changes. - docs/LOCAL-MODELS.md: provisioning + enablement for Qwen2.5-VL-3B (Apache-2.0) via llama.cpp; alternatives (Qwen3-VL, Gemma 3 4B). Honest note: likely improves the bowling case but unverified until tested with real weights. mvn verify: 293 tests green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| beat_detect.py | ||
| local_asset_requirements.txt | ||
| local_asset_worker.py | ||
| local_cv_requirements.txt | ||
| local_cv_worker.py | ||
| provision_local_models.py | ||
| run_local_asset_worker.sh | ||
| run_local_cv_worker.sh | ||
| subject_track.py | ||
| vision_caption.py | ||
| vision_caption_llamacpp.py | ||