video_editing_poc/docs/LOCAL-MODELS.md

101 lines
6.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Local model runtime & provisioning
The pipeline runs entirely on local models. This documents what works and the version traps, so the runtime
can be reproduced (reference environment: **x86_64 macOS, no GPU**; a Linux/GPU host is easier). Python venv:
`./.venv-local-asset/bin/python` (py3.12). Nothing here downloads at service runtime — models load offline.
## Models (each needs a provenance sidecar under `models/`)
| Model | Role | Runtime | Notes | License |
|---|---|---|---|---|
| Piper `en_US-lessac-medium` | voiceover | `piper-tts` 1.5.0 + onnxruntime | real speech; CLI takes `--model`/`--output_file` | MIT / Blizzard dataset |
| MusicGen small | music | `transformers` 4.44.2 | `facebook/musicgen-small`, CPU ~9× realtime, 32 kHz | **CC-BY-NC** |
| AudioLDM2 | SFX | `diffusers==0.30.3` | `cvssp/audioldm2`, CPU ~7× realtime, **16 kHz → resample** | **CC-BY-NC-SA** |
| moondream2 | Tier-2 vision director | `transformers` + `torchvision==0.17.2` | `vikhyatk/moondream2` rev `2024-08-26`, `trust_remote_code`, offline; ~25 s/frame CPU | Apache-2.0 |
| YOLOv8n (optional CV) | visual analysis | `.venv-local-cv` (ultralytics) | `./yolov8n.pt` loads offline; runs behind loopback HTTP `:8091` | **AGPL-3.0** |
## Version traps (x86_64 macOS — all real)
- **PyTorch caps at `torch==2.2.2` / `torchaudio==2.2.2`** (last x86_64 macOS wheels). numpy must be **< 2**
(pinned `numpy==1.26.4`, `numba==0.60.0`, `llvmlite==0.43.0`, `scipy==1.13.1`).
- **transformers must be 4.x** (pinned `4.44.2`). transformers 5.x silently disables PyTorch (needs torch 2.4)
models unavailable.
- **audiocraft does NOT work here** (hard top-level `from xformers import ops`; xformers has no cp312 x86_64
wheel/sdist). AudioGen is audiocraft-only **SFX uses AudioLDM2 (diffusers)** instead. MusicGen runs via
`transformers`, not audiocraft.
- **HF downloads:** set `HF_HUB_DISABLE_XET=1` (the xet CDN times out on some networks). HF may **429** after
many pulls retry resumes from cache.
## Wiring notes
- `tools/local_asset_worker.py` synthesizes voice/music/SFX; `tools/vision_caption.py` runs moondream (batch,
load-once) for the Tier-2 director.
- `LocalAssetSynthesizer` requires each model path to be a licensed **regular file** (adjacent `.license.txt`,
non-blank, not `UNTRACKED`) `AssetLicensePolicy`.
- Do **not** run `tools/run_local_cv_worker.sh` / `tools/run_local_asset_worker.sh` in a certified environment:
their `auto` modes `pip install` and can fetch a named YOLO model.
- **Licensing blocker for commercial use:** MusicGen (CC-BY-NC), AudioLDM2 (CC-BY-NC-SA) and YOLOv8 (AGPL) are
non-commercial/copyleft. Swap in commercially-licensed models/assets before any commercial release. moondream2
(Apache-2.0) and Piper are fine.
## Optional stronger Tier-2 VLM: Qwen2.5-VL via llama.cpp (recommended upgrade)
moondream2 is small and, on hard footage (distant, portrait), cannot reliably tell a highlight from an
aftermath which caps the director's vision **judge** (see `docs/cinematic-quality-rules.md` R15). A stronger
local VLM is a **drop-in**: the director calls a captioner script with a fixed manifestJSON contract, so only
the configured script path changes. **`tools/vision_caption_llamacpp.py`** implements that contract against
**llama.cpp** (GGUF) which runs on this **x86 CPU via AVX SIMD** and **bypasses the torch==2.2.2 /
transformers 4.x / no-xformers trap entirely** (no PyTorch involved).
**Recommended model:** `Qwen2.5-VL-3B-Instruct` (Apache-2.0 commercial-friendly; verify the model card),
`Q4_K_M` GGUF + its `mmproj` vision projector. `Qwen3-VL-2B/4B` (official GGUF) or `Gemma 3 4B` are alternatives.
Provision **offline** (on a networked machine, then copy the files over nothing downloads at service runtime).
The recipe below is the **verified** one for this reference machine (x86_64 macOS, Command Line Tools, no usable
GPU):
```bash
# 1) Build llama.cpp with the multimodal CLI (one-time; needs cmake + a C++ compiler; NOT in a certified env).
# GOTCHA (Command Line Tools, not full Xcode): clang can't find libc++ headers (<array> not found) because
# they live only under the SDK. Point CMAKE at them, or the ggml-base compile fails:
CXXV1=$(xcrun --show-sdk-path)/usr/include/c++/v1
git clone --depth 1 https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build -DGGML_NATIVE=ON -DLLAMA_CURL=OFF \
-DCMAKE_CXX_FLAGS="-isystem $CXXV1" -DCMAKE_C_FLAGS="-isystem $CXXV1"
cmake --build build --config Release -j --target llama-mtmd-cli # -> build/bin/llama-mtmd-cli
# 2) Fetch the GGUF weights + mmproj (verified: unsloth/Qwen2.5-VL-3B-Instruct-GGUF)
# Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf (~1.9GB) and mmproj-F16.gguf (~1.3GB)
# Record license/SHA-256 next to each file, per the asset-provenance policy.
```
Enable it (no code change config + env). **localpoc already defaults `vision-caption-script` to the
llama.cpp worker**; if the binary env or weights are missing it **falls back to moondream with a loud WARN**
(`event=vision_backend_not_ready`), so the default never silently degrades. The model/mmproj default to the
repo paths below, so only the (machine-specific) **binary env** is mandatory:
```bash
# FAST (recommended): resident server — model loads ONCE, ~3x faster per clip.
export LLAMACPP_SERVER_BIN=/abs/llama.cpp/build/bin/llama-server
# or SIMPLE: per-frame CLI (reloads the model each frame)
# export LLAMACPP_MTMD_BIN=/abs/llama.cpp/build/bin/llama-mtmd-cli
# Optional — default to ./models/qwen2.5-vl-3b/ if unset:
# export LLAMACPP_VLM_MODEL=/abs/.../Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf
# export LLAMACPP_VLM_MMPROJ=/abs/.../mmproj-F16.gguf
# export LLAMACPP_VLM_NTOKENS=64 LLAMACPP_SERVER_PORT=8123
```
To force moondream instead: set `vision-caption-script: ./tools/vision_caption.py`.
**Runtime GOTCHA — force CPU on an Intel Mac.** The build enables Metal by default, but this machine's
integrated GPU times out on the vision encoder (`ggml_metal_synchronize: command buffer failed GPU Timeout`).
`tools/vision_caption_llamacpp.py` therefore always passes `-ngl 0 --no-mmproj-offload` (pure CPU/AVX). Expect
~13 min per frame on CPU; ~7 frames per clip (candidates + overlay) 1020 min, fine for offline batch.
A persistent `llama-server` backend would avoid per-frame reloads a future optimisation.
**Verified:** Qwen2.5-VL-3B **does** break the moondream ceiling. On the bowling celebration frame moondream
said "standing in a bowling alley"; Qwen2.5-VL says *"The person is raising their arms in a celebratory
gesture"* (highlight-worthiness 1.0), and correctly rates the turn-around/anticipation frames low (0.2 / 0.15)
so the director's vision judge picks the celebration. This is the recommended production Tier-2 model.