video_editing_poc/docs/LOCAL-MODELS.md

6.1 KiB
Raw Blame History

Local model runtime & provisioning

The pipeline runs entirely on local models. This documents what works and the version traps, so the runtime can be reproduced (reference environment: x86_64 macOS, no GPU; a Linux/GPU host is easier). Python venv: ./.venv-local-asset/bin/python (py3.12). Nothing here downloads at service runtime — models load offline.

Models (each needs a provenance sidecar under models/)

Model Role Runtime Notes License
Piper en_US-lessac-medium voiceover piper-tts 1.5.0 + onnxruntime real speech; CLI takes --model/--output_file MIT / Blizzard dataset
MusicGen small music transformers 4.44.2 facebook/musicgen-small, CPU ~9× realtime, 32 kHz CC-BY-NC
AudioLDM2 SFX diffusers==0.30.3 cvssp/audioldm2, CPU ~7× realtime, 16 kHz → resample CC-BY-NC-SA
moondream2 Tier-2 vision director transformers + torchvision==0.17.2 vikhyatk/moondream2 rev 2024-08-26, trust_remote_code, offline; ~25 s/frame CPU Apache-2.0
YOLOv8n (optional CV) visual analysis .venv-local-cv (ultralytics) ./yolov8n.pt loads offline; runs behind loopback HTTP :8091 AGPL-3.0

Version traps (x86_64 macOS — all real)

  • PyTorch caps at torch==2.2.2 / torchaudio==2.2.2 (last x86_64 macOS wheels). numpy must be < 2 (pinned numpy==1.26.4, numba==0.60.0, llvmlite==0.43.0, scipy==1.13.1).
  • transformers must be 4.x (pinned 4.44.2). transformers 5.x silently disables PyTorch (needs torch ≥ 2.4) → models unavailable.
  • audiocraft does NOT work here (hard top-level from xformers import ops; xformers has no cp312 x86_64 wheel/sdist). AudioGen is audiocraft-only → SFX uses AudioLDM2 (diffusers) instead. MusicGen runs via transformers, not audiocraft.
  • HF downloads: set HF_HUB_DISABLE_XET=1 (the xet CDN times out on some networks). HF may 429 after many pulls — retry resumes from cache.

Wiring notes

  • tools/local_asset_worker.py synthesizes voice/music/SFX; tools/vision_caption.py runs moondream (batch, load-once) for the Tier-2 director.
  • LocalAssetSynthesizer requires each model path to be a licensed regular file (adjacent .license.txt, non-blank, not UNTRACKED) — AssetLicensePolicy.
  • Do not run tools/run_local_cv_worker.sh / tools/run_local_asset_worker.sh in a certified environment: their auto modes pip install and can fetch a named YOLO model.
  • Licensing blocker for commercial use: MusicGen (CC-BY-NC), AudioLDM2 (CC-BY-NC-SA) and YOLOv8 (AGPL) are non-commercial/copyleft. Swap in commercially-licensed models/assets before any commercial release. moondream2 (Apache-2.0) and Piper are fine.

moondream2 is small and, on hard footage (distant, portrait), cannot reliably tell a highlight from an aftermath — which caps the director's vision judge (see docs/cinematic-quality-rules.md R15). A stronger local VLM is a drop-in: the director calls a captioner script with a fixed manifest→JSON contract, so only the configured script path changes. tools/vision_caption_llamacpp.py implements that contract against llama.cpp (GGUF) — which runs on this x86 CPU via AVX SIMD and bypasses the torch==2.2.2 / transformers 4.x / no-xformers trap entirely (no PyTorch involved).

Recommended model: Qwen2.5-VL-3B-Instruct (Apache-2.0 — commercial-friendly; verify the model card), Q4_K_M GGUF + its mmproj vision projector. Qwen3-VL-2B/4B (official GGUF) or Gemma 3 4B are alternatives.

Provision offline (on a networked machine, then copy the files over — nothing downloads at service runtime). The recipe below is the verified one for this reference machine (x86_64 macOS, Command Line Tools, no usable GPU):

# 1) Build llama.cpp with the multimodal CLI (one-time; needs cmake + a C++ compiler; NOT in a certified env).
#    GOTCHA (Command Line Tools, not full Xcode): clang can't find libc++ headers (<array> not found) because
#    they live only under the SDK. Point CMAKE at them, or the ggml-base compile fails:
CXXV1=$(xcrun --show-sdk-path)/usr/include/c++/v1
git clone --depth 1 https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build -DGGML_NATIVE=ON -DLLAMA_CURL=OFF \
      -DCMAKE_CXX_FLAGS="-isystem $CXXV1" -DCMAKE_C_FLAGS="-isystem $CXXV1"
cmake --build build --config Release -j --target llama-mtmd-cli   # -> build/bin/llama-mtmd-cli

# 2) Fetch the GGUF weights + mmproj (verified: unsloth/Qwen2.5-VL-3B-Instruct-GGUF)
#    Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf (~1.9GB)  and  mmproj-F16.gguf (~1.3GB)
#    Record license/SHA-256 next to each file, per the asset-provenance policy.

Enable it (no code change — config + env only):

# application-localpoc.yml  (video-clipping.editing)
vision-caption-script: ./tools/vision_caption_llamacpp.py
export LLAMACPP_MTMD_BIN=/abs/llama.cpp/build/bin/llama-mtmd-cli
export LLAMACPP_VLM_MODEL=/abs/models/qwen2.5-vl-3b/Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf
export LLAMACPP_VLM_MMPROJ=/abs/models/qwen2.5-vl-3b/mmproj-F16.gguf
# LLAMACPP_VLM_NTOKENS=64 (optional)

Runtime GOTCHA — force CPU on an Intel Mac. The build enables Metal by default, but this machine's integrated GPU times out on the vision encoder (ggml_metal_synchronize: command buffer failed … GPU Timeout). tools/vision_caption_llamacpp.py therefore always passes -ngl 0 --no-mmproj-offload (pure CPU/AVX). Expect ~13 min per frame on CPU; ~7 frames per clip (candidates + overlay) ≈ 1020 min, fine for offline batch. A persistent llama-server backend would avoid per-frame reloads — a future optimisation.

Verified: Qwen2.5-VL-3B does break the moondream ceiling. On the bowling celebration frame moondream said "standing in a bowling alley"; Qwen2.5-VL says "The person is raising their arms in a celebratory gesture" (highlight-worthiness 1.0), and correctly rates the turn-around/anticipation frames low (0.2 / 0.15) — so the director's vision judge picks the celebration. This is the recommended production Tier-2 model.