The director's vision JUDGE is only as good as its model, and moondream2 can't
perceive some actions on hard footage (R15 ceiling). Make the captioner backend a
config choice so a stronger local VLM drops in with no code change:
- tools/vision_caption_llamacpp.py: same manifest->JSON contract as
tools/vision_caption.py, but backed by llama.cpp `llama-mtmd-cli` (GGUF). Runs on
this x86 CPU via AVX and bypasses the torch==2.2.2 / transformers 4.x trap
entirely (no PyTorch). Model/mmproj/binary paths come from env vars; fully
offline, serverless (per-frame CLI, mmap stays warm).
- editing.vision-caption-script selects the worker (default: moondream). The Java
HighlightVisionDirector now reads the configured script instead of a hardcoded
path -- nothing else changes.
- docs/LOCAL-MODELS.md: provisioning + enablement for Qwen2.5-VL-3B (Apache-2.0)
via llama.cpp; alternatives (Qwen3-VL, Gemma 3 4B). Honest note: likely improves
the bowling case but unverified until tested with real weights.
mvn verify: 293 tests green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Port the run-to-final flow and the local model runtime/version-trap facts from
Claude-specific auto-memory into repo Markdown so any assistant or human reading
the repo (Codex, Claude, etc.) has them. Runbook updated for the auto-director
(the plan is generated automatically now; manual authoring is an override).
Linked from the README.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MPuJXQyAeWpFcTtcnxo1UN