6.7 KiB
Local model runtime & provisioning
The pipeline runs entirely on local models. This documents what works and the version traps, so the runtime
can be reproduced (reference environment: x86_64 macOS, no GPU; a Linux/GPU host is easier). Python venv:
./.venv-local-asset/bin/python (py3.12). Nothing here downloads at service runtime — models load offline.
Models (each needs a provenance sidecar under models/)
| Model | Role | Runtime | Notes | License |
|---|---|---|---|---|
Piper en_US-lessac-medium |
voiceover | piper-tts 1.5.0 + onnxruntime |
real speech; CLI takes --model/--output_file |
MIT / Blizzard dataset |
| MusicGen small | music | transformers 4.44.2 |
facebook/musicgen-small, CPU ~9× realtime, 32 kHz |
CC-BY-NC |
| AudioLDM2 | SFX | diffusers==0.30.3 |
cvssp/audioldm2, CPU ~7× realtime, 16 kHz → resample |
CC-BY-NC-SA |
| moondream2 | Tier-2 vision director | transformers + torchvision==0.17.2 |
vikhyatk/moondream2 rev 2024-08-26, trust_remote_code, offline; ~25 s/frame CPU |
Apache-2.0 |
| YOLOv8n (optional CV) | visual analysis | .venv-local-cv (ultralytics) |
./yolov8n.pt loads offline; runs behind loopback HTTP :8091 |
AGPL-3.0 |
Version traps (x86_64 macOS — all real)
- PyTorch caps at
torch==2.2.2/torchaudio==2.2.2(last x86_64 macOS wheels). numpy must be < 2 (pinnednumpy==1.26.4,numba==0.60.0,llvmlite==0.43.0,scipy==1.13.1). - transformers must be 4.x (pinned
4.44.2). transformers 5.x silently disables PyTorch (needs torch ≥ 2.4) → models unavailable. - audiocraft does NOT work here (hard top-level
from xformers import ops; xformers has no cp312 x86_64 wheel/sdist). AudioGen is audiocraft-only → SFX uses AudioLDM2 (diffusers) instead. MusicGen runs viatransformers, not audiocraft. - HF downloads: set
HF_HUB_DISABLE_XET=1(the xet CDN times out on some networks). HF may 429 after many pulls — retry resumes from cache.
Wiring notes
tools/local_asset_worker.pysynthesizes voice/music/SFX;tools/vision_caption.pyruns moondream (batch, load-once) for the Tier-2 director.LocalAssetSynthesizerrequires each model path to be a licensed regular file (adjacent.license.txt, non-blank, notUNTRACKED) —AssetLicensePolicy.- Do not run
tools/run_local_cv_worker.sh/tools/run_local_asset_worker.shin a certified environment: theirautomodespip installand can fetch a named YOLO model. - Licensing blocker for commercial use: MusicGen (CC-BY-NC), AudioLDM2 (CC-BY-NC-SA) and YOLOv8 (AGPL) are non-commercial/copyleft. Swap in commercially-licensed models/assets before any commercial release. moondream2 (Apache-2.0) and Piper are fine.
Optional stronger Tier-2 VLM: Qwen2.5-VL via llama.cpp (recommended upgrade)
moondream2 is small and, on hard footage (distant, portrait), cannot reliably tell a highlight from an
aftermath — which caps the director's vision judge (see docs/cinematic-quality-rules.md R15). A stronger
local VLM is a drop-in: the director calls a captioner script with a fixed manifest→JSON contract, so only
the configured script path changes. tools/vision_caption_llamacpp.py implements that contract against
llama.cpp (GGUF) — which runs on this x86 CPU via AVX SIMD and bypasses the torch==2.2.2 /
transformers 4.x / no-xformers trap entirely (no PyTorch involved).
Recommended model: Qwen2.5-VL-3B-Instruct (Apache-2.0 — commercial-friendly; verify the model card),
Q4_K_M GGUF + its mmproj vision projector. Qwen3-VL-2B/4B (official GGUF) or Gemma 3 4B are alternatives.
Provision offline (on a networked machine, then copy the files over — nothing downloads at service runtime). The recipe below is the verified one for this reference machine (x86_64 macOS, Command Line Tools, no usable GPU):
# 1) Build llama.cpp with the multimodal CLI (one-time; needs cmake + a C++ compiler; NOT in a certified env).
# GOTCHA (Command Line Tools, not full Xcode): clang can't find libc++ headers (<array> not found) because
# they live only under the SDK. Point CMAKE at them, or the ggml-base compile fails:
CXXV1=$(xcrun --show-sdk-path)/usr/include/c++/v1
git clone --depth 1 https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build -DGGML_NATIVE=ON -DLLAMA_CURL=OFF \
-DCMAKE_CXX_FLAGS="-isystem $CXXV1" -DCMAKE_C_FLAGS="-isystem $CXXV1"
cmake --build build --config Release -j --target llama-mtmd-cli # -> build/bin/llama-mtmd-cli
# 2) Fetch the GGUF weights + mmproj (verified: unsloth/Qwen2.5-VL-3B-Instruct-GGUF)
# Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf (~1.9GB) and mmproj-F16.gguf (~1.3GB)
# Record license/SHA-256 next to each file, per the asset-provenance policy.
Enable it (no code change — config + env). localpoc already defaults vision-caption-script to the
llama.cpp worker; if the binary env or weights are missing it falls back to moondream with a loud WARN
(event=vision_backend_not_ready), so the default never silently degrades. The model/mmproj default to the
repo paths below, so only the (machine-specific) binary env is mandatory:
# FAST (recommended): resident server — model loads ONCE, ~3x faster per clip.
export LLAMACPP_SERVER_BIN=/abs/llama.cpp/build/bin/llama-server
# or SIMPLE: per-frame CLI (reloads the model each frame)
# export LLAMACPP_MTMD_BIN=/abs/llama.cpp/build/bin/llama-mtmd-cli
# Optional — default to ./models/qwen2.5-vl-3b/ if unset:
# export LLAMACPP_VLM_MODEL=/abs/.../Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf
# export LLAMACPP_VLM_MMPROJ=/abs/.../mmproj-F16.gguf
# export LLAMACPP_VLM_NTOKENS=64 LLAMACPP_SERVER_PORT=8123
To force moondream instead: set vision-caption-script: ./tools/vision_caption.py.
Runtime GOTCHA — force CPU on an Intel Mac. The build enables Metal by default, but this machine's
integrated GPU times out on the vision encoder (ggml_metal_synchronize: command buffer failed … GPU Timeout).
tools/vision_caption_llamacpp.py therefore always passes -ngl 0 --no-mmproj-offload (pure CPU/AVX). Expect
~1–3 min per frame on CPU; ~7 frames per clip (candidates + overlay) ≈ 10–20 min, fine for offline batch.
A persistent llama-server backend would avoid per-frame reloads — a future optimisation.
Verified: Qwen2.5-VL-3B does break the moondream ceiling. On the bowling celebration frame moondream said "standing in a bowling alley"; Qwen2.5-VL says "The person is raising their arms in a celebratory gesture" (highlight-worthiness 1.0), and correctly rates the turn-around/anticipation frames low (0.2 / 0.15) — so the director's vision judge picks the celebration. This is the recommended production Tier-2 model.