video_editing_poc/docs/LOCAL-MODELS.md

40 lines
2.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Local model runtime & provisioning
The pipeline runs entirely on local models. This documents what works and the version traps, so the runtime
can be reproduced (reference environment: **x86_64 macOS, no GPU**; a Linux/GPU host is easier). Python venv:
`./.venv-local-asset/bin/python` (py3.12). Nothing here downloads at service runtime — models load offline.
## Models (each needs a provenance sidecar under `models/`)
| Model | Role | Runtime | Notes | License |
|---|---|---|---|---|
| Piper `en_US-lessac-medium` | voiceover | `piper-tts` 1.5.0 + onnxruntime | real speech; CLI takes `--model`/`--output_file` | MIT / Blizzard dataset |
| MusicGen small | music | `transformers` 4.44.2 | `facebook/musicgen-small`, CPU ~9× realtime, 32 kHz | **CC-BY-NC** |
| AudioLDM2 | SFX | `diffusers==0.30.3` | `cvssp/audioldm2`, CPU ~7× realtime, **16 kHz → resample** | **CC-BY-NC-SA** |
| moondream2 | Tier-2 vision director | `transformers` + `torchvision==0.17.2` | `vikhyatk/moondream2` rev `2024-08-26`, `trust_remote_code`, offline; ~25 s/frame CPU | Apache-2.0 |
| YOLOv8n (optional CV) | visual analysis | `.venv-local-cv` (ultralytics) | `./yolov8n.pt` loads offline; runs behind loopback HTTP `:8091` | **AGPL-3.0** |
## Version traps (x86_64 macOS — all real)
- **PyTorch caps at `torch==2.2.2` / `torchaudio==2.2.2`** (last x86_64 macOS wheels). numpy must be **< 2**
(pinned `numpy==1.26.4`, `numba==0.60.0`, `llvmlite==0.43.0`, `scipy==1.13.1`).
- **transformers must be 4.x** (pinned `4.44.2`). transformers 5.x silently disables PyTorch (needs torch 2.4)
models unavailable.
- **audiocraft does NOT work here** (hard top-level `from xformers import ops`; xformers has no cp312 x86_64
wheel/sdist). AudioGen is audiocraft-only **SFX uses AudioLDM2 (diffusers)** instead. MusicGen runs via
`transformers`, not audiocraft.
- **HF downloads:** set `HF_HUB_DISABLE_XET=1` (the xet CDN times out on some networks). HF may **429** after
many pulls retry resumes from cache.
## Wiring notes
- `tools/local_asset_worker.py` synthesizes voice/music/SFX; `tools/vision_caption.py` runs moondream (batch,
load-once) for the Tier-2 director.
- `LocalAssetSynthesizer` requires each model path to be a licensed **regular file** (adjacent `.license.txt`,
non-blank, not `UNTRACKED`) `AssetLicensePolicy`.
- Do **not** run `tools/run_local_cv_worker.sh` / `tools/run_local_asset_worker.sh` in a certified environment:
their `auto` modes `pip install` and can fetch a named YOLO model.
- **Licensing blocker for commercial use:** MusicGen (CC-BY-NC), AudioLDM2 (CC-BY-NC-SA) and YOLOv8 (AGPL) are
non-commercial/copyleft. Swap in commercially-licensed models/assets before any commercial release. moondream2
(Apache-2.0) and Piper are fine.