fix(vision): Qwen2.5-VL llama.cpp backend — verified to break the moondream ceiling
Provisioned and tested the stronger Tier-2 judge on this Intel Mac (CPU-only): - vision_caption_llamacpp.py: force CPU (-ngl 0 --no-mmproj-offload) because the integrated GPU times out on the vision encoder (Metal command-buffer timeout); extract the final assistant turn from the chat-templated output. - docs/LOCAL-MODELS.md: the VERIFIED build+run recipe, incl. two real gotchas -> Command Line Tools libc++ mismatch (add -isystem <SDK>/usr/include/c++/v1 to the cmake flags, else ggml-base fails on <array>), and the Intel-GPU Metal timeout. Verified result: where moondream described the bowling celebration as "standing in a bowling alley", Qwen2.5-VL-3B says "raising their arms in a celebratory gesture" (highlight-worthiness 1.0) and rates turn-around/anticipation low (0.2 / 0.15). End to end, the director's judge now chooses time=10.5s score=1.0 (the celebration) vs moondream's 13.0s score=0.15 (the turn-away). Model weights are gitignored under models/ (Apache-2.0, provisioned offline). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
parent
3652cefacc
commit
cf8ec76430
|
|
@ -50,15 +50,22 @@ transformers 4.x / no-xformers trap entirely** (no PyTorch involved).
|
||||||
**Recommended model:** `Qwen2.5-VL-3B-Instruct` (Apache-2.0 — commercial-friendly; verify the model card),
|
**Recommended model:** `Qwen2.5-VL-3B-Instruct` (Apache-2.0 — commercial-friendly; verify the model card),
|
||||||
`Q4_K_M` GGUF + its `mmproj` vision projector. `Qwen3-VL-2B/4B` (official GGUF) or `Gemma 3 4B` are alternatives.
|
`Q4_K_M` GGUF + its `mmproj` vision projector. `Qwen3-VL-2B/4B` (official GGUF) or `Gemma 3 4B` are alternatives.
|
||||||
|
|
||||||
Provision **offline** (on a networked machine, then copy the files over — nothing downloads at service runtime):
|
Provision **offline** (on a networked machine, then copy the files over — nothing downloads at service runtime).
|
||||||
|
The recipe below is the **verified** one for this reference machine (x86_64 macOS, Command Line Tools, no usable
|
||||||
|
GPU):
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# 1) Build llama.cpp with the multimodal CLI (one-time, needs a compiler; NOT in a certified/offline env)
|
# 1) Build llama.cpp with the multimodal CLI (one-time; needs cmake + a C++ compiler; NOT in a certified env).
|
||||||
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp && cmake -B build && cmake --build build -j
|
# GOTCHA (Command Line Tools, not full Xcode): clang can't find libc++ headers (<array> not found) because
|
||||||
# -> produces build/bin/llama-mtmd-cli
|
# they live only under the SDK. Point CMAKE at them, or the ggml-base compile fails:
|
||||||
|
CXXV1=$(xcrun --show-sdk-path)/usr/include/c++/v1
|
||||||
|
git clone --depth 1 https://github.com/ggml-org/llama.cpp && cd llama.cpp
|
||||||
|
cmake -B build -DGGML_NATIVE=ON -DLLAMA_CURL=OFF \
|
||||||
|
-DCMAKE_CXX_FLAGS="-isystem $CXXV1" -DCMAKE_C_FLAGS="-isystem $CXXV1"
|
||||||
|
cmake --build build --config Release -j --target llama-mtmd-cli # -> build/bin/llama-mtmd-cli
|
||||||
|
|
||||||
# 2) Fetch the GGUF weights + mmproj (e.g. from a bartowski/Mungert/official Qwen GGUF repo)
|
# 2) Fetch the GGUF weights + mmproj (verified: unsloth/Qwen2.5-VL-3B-Instruct-GGUF)
|
||||||
# Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf and mmproj-Qwen2.5-VL-3B-Instruct-f16.gguf
|
# Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf (~1.9GB) and mmproj-F16.gguf (~1.3GB)
|
||||||
# Record license/SHA-256 next to each file, per the asset-provenance policy.
|
# Record license/SHA-256 next to each file, per the asset-provenance policy.
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
@ -71,12 +78,17 @@ vision-caption-script: ./tools/vision_caption_llamacpp.py
|
||||||
```bash
|
```bash
|
||||||
export LLAMACPP_MTMD_BIN=/abs/llama.cpp/build/bin/llama-mtmd-cli
|
export LLAMACPP_MTMD_BIN=/abs/llama.cpp/build/bin/llama-mtmd-cli
|
||||||
export LLAMACPP_VLM_MODEL=/abs/models/qwen2.5-vl-3b/Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf
|
export LLAMACPP_VLM_MODEL=/abs/models/qwen2.5-vl-3b/Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf
|
||||||
export LLAMACPP_VLM_MMPROJ=/abs/models/qwen2.5-vl-3b/mmproj-Qwen2.5-VL-3B-Instruct-f16.gguf
|
export LLAMACPP_VLM_MMPROJ=/abs/models/qwen2.5-vl-3b/mmproj-F16.gguf
|
||||||
# LLAMACPP_VLM_NTOKENS=64 (optional)
|
# LLAMACPP_VLM_NTOKENS=64 (optional)
|
||||||
```
|
```
|
||||||
|
|
||||||
Notes: the worker invokes `llama-mtmd-cli` per frame (serverless, fully offline; the mmap'd model stays warm in
|
**Runtime GOTCHA — force CPU on an Intel Mac.** The build enables Metal by default, but this machine's
|
||||||
the OS cache across frames). A persistent `llama-server` backend would be faster for large batches — a future
|
integrated GPU times out on the vision encoder (`ggml_metal_synchronize: command buffer failed … GPU Timeout`).
|
||||||
optimisation, not required. It is **unverified whether Qwen2.5-VL breaks the specific bowling case** — it is
|
`tools/vision_caption_llamacpp.py` therefore always passes `-ngl 0 --no-mmproj-offload` (pure CPU/AVX). Expect
|
||||||
substantially more capable than moondream, so it likely improves discrimination, but that is a hypothesis to
|
~1–3 min per frame on CPU; ~7 frames per clip (candidates + overlay) ≈ 10–20 min, fine for offline batch.
|
||||||
test, not a guarantee.
|
A persistent `llama-server` backend would avoid per-frame reloads — a future optimisation.
|
||||||
|
|
||||||
|
**Verified:** Qwen2.5-VL-3B **does** break the moondream ceiling. On the bowling celebration frame moondream
|
||||||
|
said "standing in a bowling alley"; Qwen2.5-VL says *"The person is raising their arms in a celebratory
|
||||||
|
gesture"* (highlight-worthiness 1.0), and correctly rates the turn-around/anticipation frames low (0.2 / 0.15)
|
||||||
|
— so the director's vision judge picks the celebration. This is the recommended production Tier-2 model.
|
||||||
|
|
|
||||||
|
|
@ -30,21 +30,30 @@ import sys
|
||||||
|
|
||||||
|
|
||||||
def _clean(text: str) -> str:
|
def _clean(text: str) -> str:
|
||||||
# Strip llama.cpp control/log noise and whitespace; keep the model's answer text.
|
# The CLI echoes the chat-templated conversation; the real answer is the LAST assistant turn.
|
||||||
text = text.replace("\x1b[0m", "").strip()
|
text = text.replace("\x1b[0m", "")
|
||||||
# Drop obvious log lines (timings, "llama_", "main:", "mtmd_", "encoding image") if any leak to stdout.
|
marker = "assistant"
|
||||||
lines = [ln for ln in text.splitlines()
|
idx = text.rfind(marker)
|
||||||
if ln.strip() and not re.match(r"^(llama_|main:|mtmd_|clip_|ggml_|build:|load|encoding image|\s*[\d.]+\s*(ms|tokens))",
|
if idx >= 0:
|
||||||
ln.strip(), re.IGNORECASE)]
|
text = text[idx + len(marker):]
|
||||||
|
text = re.sub(r"<\|[^>]*\|>", " ", text) # strip chat control tokens like <|im_end|>
|
||||||
|
text = text.replace("[end of text]", " ")
|
||||||
|
lines = [ln.strip() for ln in text.splitlines() if ln.strip()]
|
||||||
|
# Drop any leaked log/timing lines; keep the natural-language answer.
|
||||||
|
lines = [ln for ln in lines
|
||||||
|
if not re.match(r"^(llama_|main:|mtmd_|clip_|ggml_|build:|load|encoding|decoding|\d[\d.:]*\s)",
|
||||||
|
ln, re.IGNORECASE)]
|
||||||
return " ".join(lines).strip()
|
return " ".join(lines).strip()
|
||||||
|
|
||||||
|
|
||||||
def _caption(binary: str, model: str, mmproj: str, image: str, question: str, ntokens: int) -> str:
|
def _caption(binary: str, model: str, mmproj: str, image: str, question: str, ntokens: int) -> str:
|
||||||
command = [
|
command = [
|
||||||
binary, "-m", model, "--mmproj", mmproj, "--image", image,
|
binary, "-m", model, "--mmproj", mmproj, "--image", image,
|
||||||
"-p", question, "--temp", "0", "-n", str(ntokens),
|
"-p", question, "--temp", "0", "-n", str(ntokens), "-t", "4",
|
||||||
|
# Force CPU: this machine's integrated GPU times out (Metal command-buffer) on the vision encoder.
|
||||||
|
"-ngl", "0", "--no-mmproj-offload",
|
||||||
]
|
]
|
||||||
proc = subprocess.run(command, capture_output=True, text=True, timeout=600)
|
proc = subprocess.run(command, capture_output=True, text=True, timeout=900)
|
||||||
if proc.returncode != 0:
|
if proc.returncode != 0:
|
||||||
sys.stderr.write("vision_caption_llamacpp: cli failed (%s): %s\n" % (proc.returncode, proc.stderr[-500:]))
|
sys.stderr.write("vision_caption_llamacpp: cli failed (%s): %s\n" % (proc.returncode, proc.stderr[-500:]))
|
||||||
return ""
|
return ""
|
||||||
|
|
|
||||||
Loading…
Reference in New Issue