Commit Graph

2 Commits

Author SHA1 Message Date
JSLMPR cf8ec76430 fix(vision): Qwen2.5-VL llama.cpp backend — verified to break the moondream ceiling
Provisioned and tested the stronger Tier-2 judge on this Intel Mac (CPU-only):

- vision_caption_llamacpp.py: force CPU (-ngl 0 --no-mmproj-offload) because the
  integrated GPU times out on the vision encoder (Metal command-buffer timeout);
  extract the final assistant turn from the chat-templated output.
- docs/LOCAL-MODELS.md: the VERIFIED build+run recipe, incl. two real gotchas ->
  Command Line Tools libc++ mismatch (add -isystem <SDK>/usr/include/c++/v1 to the
  cmake flags, else ggml-base fails on <array>), and the Intel-GPU Metal timeout.

Verified result: where moondream described the bowling celebration as "standing in
a bowling alley", Qwen2.5-VL-3B says "raising their arms in a celebratory gesture"
(highlight-worthiness 1.0) and rates turn-around/anticipation low (0.2 / 0.15). End
to end, the director's judge now chooses time=10.5s score=1.0 (the celebration) vs
moondream's 13.0s score=0.15 (the turn-away). Model weights are gitignored under
models/ (Apache-2.0, provisioned offline).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-25 23:50:12 +02:00
JSLMPR 3652cefacc feat(vision): pluggable llama.cpp/GGUF Tier-2 VLM backend (stronger judge)
The director's vision JUDGE is only as good as its model, and moondream2 can't
perceive some actions on hard footage (R15 ceiling). Make the captioner backend a
config choice so a stronger local VLM drops in with no code change:

- tools/vision_caption_llamacpp.py: same manifest->JSON contract as
  tools/vision_caption.py, but backed by llama.cpp `llama-mtmd-cli` (GGUF). Runs on
  this x86 CPU via AVX and bypasses the torch==2.2.2 / transformers 4.x trap
  entirely (no PyTorch). Model/mmproj/binary paths come from env vars; fully
  offline, serverless (per-frame CLI, mmap stays warm).
- editing.vision-caption-script selects the worker (default: moondream). The Java
  HighlightVisionDirector now reads the configured script instead of a hardcoded
  path -- nothing else changes.
- docs/LOCAL-MODELS.md: provisioning + enablement for Qwen2.5-VL-3B (Apache-2.0)
  via llama.cpp; alternatives (Qwen3-VL, Gemma 3 4B). Honest note: likely improves
  the bowling case but unverified until tested with real weights.

mvn verify: 293 tests green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-25 19:11:54 +02:00