fix(vision): Qwen2.5-VL llama.cpp backend — verified to break the moondream ceiling

Provisioned and tested the stronger Tier-2 judge on this Intel Mac (CPU-only):

- vision_caption_llamacpp.py: force CPU (-ngl 0 --no-mmproj-offload) because the
  integrated GPU times out on the vision encoder (Metal command-buffer timeout);
  extract the final assistant turn from the chat-templated output.
- docs/LOCAL-MODELS.md: the VERIFIED build+run recipe, incl. two real gotchas ->
  Command Line Tools libc++ mismatch (add -isystem <SDK>/usr/include/c++/v1 to the
  cmake flags, else ggml-base fails on <array>), and the Intel-GPU Metal timeout.

Verified result: where moondream described the bowling celebration as "standing in
a bowling alley", Qwen2.5-VL-3B says "raising their arms in a celebratory gesture"
(highlight-worthiness 1.0) and rates turn-around/anticipation low (0.2 / 0.15). End
to end, the director's judge now chooses time=10.5s score=1.0 (the celebration) vs
moondream's 13.0s score=0.15 (the turn-away). Model weights are gitignored under
models/ (Apache-2.0, provisioned offline).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
JSLMPR 2026-07-25 23:50:12 +02:00
parent 3652cefacc
commit cf8ec76430
2 changed files with 41 additions and 20 deletions

View File

@ -50,15 +50,22 @@ transformers 4.x / no-xformers trap entirely** (no PyTorch involved).
**Recommended model:** `Qwen2.5-VL-3B-Instruct` (Apache-2.0 — commercial-friendly; verify the model card),
`Q4_K_M` GGUF + its `mmproj` vision projector. `Qwen3-VL-2B/4B` (official GGUF) or `Gemma 3 4B` are alternatives.
Provision **offline** (on a networked machine, then copy the files over — nothing downloads at service runtime):
Provision **offline** (on a networked machine, then copy the files over — nothing downloads at service runtime).
The recipe below is the **verified** one for this reference machine (x86_64 macOS, Command Line Tools, no usable
GPU):
```bash
# 1) Build llama.cpp with the multimodal CLI (one-time, needs a compiler; NOT in a certified/offline env)
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp && cmake -B build && cmake --build build -j
# -> produces build/bin/llama-mtmd-cli
# 1) Build llama.cpp with the multimodal CLI (one-time; needs cmake + a C++ compiler; NOT in a certified env).
# GOTCHA (Command Line Tools, not full Xcode): clang can't find libc++ headers (<array> not found) because
# they live only under the SDK. Point CMAKE at them, or the ggml-base compile fails:
CXXV1=$(xcrun --show-sdk-path)/usr/include/c++/v1
git clone --depth 1 https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build -DGGML_NATIVE=ON -DLLAMA_CURL=OFF \
-DCMAKE_CXX_FLAGS="-isystem $CXXV1" -DCMAKE_C_FLAGS="-isystem $CXXV1"
cmake --build build --config Release -j --target llama-mtmd-cli # -> build/bin/llama-mtmd-cli
# 2) Fetch the GGUF weights + mmproj (e.g. from a bartowski/Mungert/official Qwen GGUF repo)
# Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf and mmproj-Qwen2.5-VL-3B-Instruct-f16.gguf
# 2) Fetch the GGUF weights + mmproj (verified: unsloth/Qwen2.5-VL-3B-Instruct-GGUF)
# Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf (~1.9GB) and mmproj-F16.gguf (~1.3GB)
# Record license/SHA-256 next to each file, per the asset-provenance policy.
```
@ -71,12 +78,17 @@ vision-caption-script: ./tools/vision_caption_llamacpp.py
```bash
export LLAMACPP_MTMD_BIN=/abs/llama.cpp/build/bin/llama-mtmd-cli
export LLAMACPP_VLM_MODEL=/abs/models/qwen2.5-vl-3b/Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf
export LLAMACPP_VLM_MMPROJ=/abs/models/qwen2.5-vl-3b/mmproj-Qwen2.5-VL-3B-Instruct-f16.gguf
export LLAMACPP_VLM_MMPROJ=/abs/models/qwen2.5-vl-3b/mmproj-F16.gguf
# LLAMACPP_VLM_NTOKENS=64 (optional)
```
Notes: the worker invokes `llama-mtmd-cli` per frame (serverless, fully offline; the mmap'd model stays warm in
the OS cache across frames). A persistent `llama-server` backend would be faster for large batches — a future
optimisation, not required. It is **unverified whether Qwen2.5-VL breaks the specific bowling case** — it is
substantially more capable than moondream, so it likely improves discrimination, but that is a hypothesis to
test, not a guarantee.
**Runtime GOTCHA — force CPU on an Intel Mac.** The build enables Metal by default, but this machine's
integrated GPU times out on the vision encoder (`ggml_metal_synchronize: command buffer failed … GPU Timeout`).
`tools/vision_caption_llamacpp.py` therefore always passes `-ngl 0 --no-mmproj-offload` (pure CPU/AVX). Expect
~13 min per frame on CPU; ~7 frames per clip (candidates + overlay) ≈ 1020 min, fine for offline batch.
A persistent `llama-server` backend would avoid per-frame reloads — a future optimisation.
**Verified:** Qwen2.5-VL-3B **does** break the moondream ceiling. On the bowling celebration frame moondream
said "standing in a bowling alley"; Qwen2.5-VL says *"The person is raising their arms in a celebratory
gesture"* (highlight-worthiness 1.0), and correctly rates the turn-around/anticipation frames low (0.2 / 0.15)
— so the director's vision judge picks the celebration. This is the recommended production Tier-2 model.

View File

@ -30,21 +30,30 @@ import sys
def _clean(text: str) -> str:
# Strip llama.cpp control/log noise and whitespace; keep the model's answer text.
text = text.replace("\x1b[0m", "").strip()
# Drop obvious log lines (timings, "llama_", "main:", "mtmd_", "encoding image") if any leak to stdout.
lines = [ln for ln in text.splitlines()
if ln.strip() and not re.match(r"^(llama_|main:|mtmd_|clip_|ggml_|build:|load|encoding image|\s*[\d.]+\s*(ms|tokens))",
ln.strip(), re.IGNORECASE)]
# The CLI echoes the chat-templated conversation; the real answer is the LAST assistant turn.
text = text.replace("\x1b[0m", "")
marker = "assistant"
idx = text.rfind(marker)
if idx >= 0:
text = text[idx + len(marker):]
text = re.sub(r"<\|[^>]*\|>", " ", text) # strip chat control tokens like <|im_end|>
text = text.replace("[end of text]", " ")
lines = [ln.strip() for ln in text.splitlines() if ln.strip()]
# Drop any leaked log/timing lines; keep the natural-language answer.
lines = [ln for ln in lines
if not re.match(r"^(llama_|main:|mtmd_|clip_|ggml_|build:|load|encoding|decoding|\d[\d.:]*\s)",
ln, re.IGNORECASE)]
return " ".join(lines).strip()
def _caption(binary: str, model: str, mmproj: str, image: str, question: str, ntokens: int) -> str:
command = [
binary, "-m", model, "--mmproj", mmproj, "--image", image,
"-p", question, "--temp", "0", "-n", str(ntokens),
"-p", question, "--temp", "0", "-n", str(ntokens), "-t", "4",
# Force CPU: this machine's integrated GPU times out (Metal command-buffer) on the vision encoder.
"-ngl", "0", "--no-mmproj-offload",
]
proc = subprocess.run(command, capture_output=True, text=True, timeout=600)
proc = subprocess.run(command, capture_output=True, text=True, timeout=900)
if proc.returncode != 0:
sys.stderr.write("vision_caption_llamacpp: cli failed (%s): %s\n" % (proc.returncode, proc.stderr[-500:]))
return ""