5.2 KiB
Local model runtime & provisioning
The pipeline runs entirely on local models. This documents what works and the version traps, so the runtime
can be reproduced (reference environment: x86_64 macOS, no GPU; a Linux/GPU host is easier). Python venv:
./.venv-local-asset/bin/python (py3.12). Nothing here downloads at service runtime — models load offline.
Models (each needs a provenance sidecar under models/)
| Model | Role | Runtime | Notes | License |
|---|---|---|---|---|
Piper en_US-lessac-medium |
voiceover | piper-tts 1.5.0 + onnxruntime |
real speech; CLI takes --model/--output_file |
MIT / Blizzard dataset |
| MusicGen small | music | transformers 4.44.2 |
facebook/musicgen-small, CPU ~9× realtime, 32 kHz |
CC-BY-NC |
| AudioLDM2 | SFX | diffusers==0.30.3 |
cvssp/audioldm2, CPU ~7× realtime, 16 kHz → resample |
CC-BY-NC-SA |
| moondream2 | Tier-2 vision director | transformers + torchvision==0.17.2 |
vikhyatk/moondream2 rev 2024-08-26, trust_remote_code, offline; ~25 s/frame CPU |
Apache-2.0 |
| YOLOv8n (optional CV) | visual analysis | .venv-local-cv (ultralytics) |
./yolov8n.pt loads offline; runs behind loopback HTTP :8091 |
AGPL-3.0 |
Version traps (x86_64 macOS — all real)
- PyTorch caps at
torch==2.2.2/torchaudio==2.2.2(last x86_64 macOS wheels). numpy must be < 2 (pinnednumpy==1.26.4,numba==0.60.0,llvmlite==0.43.0,scipy==1.13.1). - transformers must be 4.x (pinned
4.44.2). transformers 5.x silently disables PyTorch (needs torch ≥ 2.4) → models unavailable. - audiocraft does NOT work here (hard top-level
from xformers import ops; xformers has no cp312 x86_64 wheel/sdist). AudioGen is audiocraft-only → SFX uses AudioLDM2 (diffusers) instead. MusicGen runs viatransformers, not audiocraft. - HF downloads: set
HF_HUB_DISABLE_XET=1(the xet CDN times out on some networks). HF may 429 after many pulls — retry resumes from cache.
Wiring notes
tools/local_asset_worker.pysynthesizes voice/music/SFX;tools/vision_caption.pyruns moondream (batch, load-once) for the Tier-2 director.LocalAssetSynthesizerrequires each model path to be a licensed regular file (adjacent.license.txt, non-blank, notUNTRACKED) —AssetLicensePolicy.- Do not run
tools/run_local_cv_worker.sh/tools/run_local_asset_worker.shin a certified environment: theirautomodespip installand can fetch a named YOLO model. - Licensing blocker for commercial use: MusicGen (CC-BY-NC), AudioLDM2 (CC-BY-NC-SA) and YOLOv8 (AGPL) are non-commercial/copyleft. Swap in commercially-licensed models/assets before any commercial release. moondream2 (Apache-2.0) and Piper are fine.
Optional stronger Tier-2 VLM: Qwen2.5-VL via llama.cpp (recommended upgrade)
moondream2 is small and, on hard footage (distant, portrait), cannot reliably tell a highlight from an
aftermath — which caps the director's vision judge (see docs/cinematic-quality-rules.md R15). A stronger
local VLM is a drop-in: the director calls a captioner script with a fixed manifest→JSON contract, so only
the configured script path changes. tools/vision_caption_llamacpp.py implements that contract against
llama.cpp (GGUF) — which runs on this x86 CPU via AVX SIMD and bypasses the torch==2.2.2 /
transformers 4.x / no-xformers trap entirely (no PyTorch involved).
Recommended model: Qwen2.5-VL-3B-Instruct (Apache-2.0 — commercial-friendly; verify the model card),
Q4_K_M GGUF + its mmproj vision projector. Qwen3-VL-2B/4B (official GGUF) or Gemma 3 4B are alternatives.
Provision offline (on a networked machine, then copy the files over — nothing downloads at service runtime):
# 1) Build llama.cpp with the multimodal CLI (one-time, needs a compiler; NOT in a certified/offline env)
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp && cmake -B build && cmake --build build -j
# -> produces build/bin/llama-mtmd-cli
# 2) Fetch the GGUF weights + mmproj (e.g. from a bartowski/Mungert/official Qwen GGUF repo)
# Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf and mmproj-Qwen2.5-VL-3B-Instruct-f16.gguf
# Record license/SHA-256 next to each file, per the asset-provenance policy.
Enable it (no code change — config + env only):
# application-localpoc.yml (video-clipping.editing)
vision-caption-script: ./tools/vision_caption_llamacpp.py
export LLAMACPP_MTMD_BIN=/abs/llama.cpp/build/bin/llama-mtmd-cli
export LLAMACPP_VLM_MODEL=/abs/models/qwen2.5-vl-3b/Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf
export LLAMACPP_VLM_MMPROJ=/abs/models/qwen2.5-vl-3b/mmproj-Qwen2.5-VL-3B-Instruct-f16.gguf
# LLAMACPP_VLM_NTOKENS=64 (optional)
Notes: the worker invokes llama-mtmd-cli per frame (serverless, fully offline; the mmap'd model stays warm in
the OS cache across frames). A persistent llama-server backend would be faster for large batches — a future
optimisation, not required. It is unverified whether Qwen2.5-VL breaks the specific bowling case — it is
substantially more capable than moondream, so it likely improves discrimination, but that is a hypothesis to
test, not a guarantee.