# Local model runtime & provisioning The pipeline runs entirely on local models. This documents what works and the version traps, so the runtime can be reproduced (reference environment: **x86_64 macOS, no GPU**; a Linux/GPU host is easier). Python venv: `./.venv-local-asset/bin/python` (py3.12). Nothing here downloads at service runtime — models load offline. ## Models (each needs a provenance sidecar under `models/`) | Model | Role | Runtime | Notes | License | |---|---|---|---|---| | Piper `en_US-lessac-medium` | voiceover | `piper-tts` 1.5.0 + onnxruntime | real speech; CLI takes `--model`/`--output_file` | MIT / Blizzard dataset | | MusicGen small | music | `transformers` 4.44.2 | `facebook/musicgen-small`, CPU ~9× realtime, 32 kHz | **CC-BY-NC** | | AudioLDM2 | SFX | `diffusers==0.30.3` | `cvssp/audioldm2`, CPU ~7× realtime, **16 kHz → resample** | **CC-BY-NC-SA** | | moondream2 | Tier-2 vision director | `transformers` + `torchvision==0.17.2` | `vikhyatk/moondream2` rev `2024-08-26`, `trust_remote_code`, offline; ~25 s/frame CPU | Apache-2.0 | | YOLOv8n (optional CV) | visual analysis | `.venv-local-cv` (ultralytics) | `./yolov8n.pt` loads offline; runs behind loopback HTTP `:8091` | **AGPL-3.0** | ## Version traps (x86_64 macOS — all real) - **PyTorch caps at `torch==2.2.2` / `torchaudio==2.2.2`** (last x86_64 macOS wheels). numpy must be **< 2** (pinned `numpy==1.26.4`, `numba==0.60.0`, `llvmlite==0.43.0`, `scipy==1.13.1`). - **transformers must be 4.x** (pinned `4.44.2`). transformers 5.x silently disables PyTorch (needs torch ≥ 2.4) → models unavailable. - **audiocraft does NOT work here** (hard top-level `from xformers import ops`; xformers has no cp312 x86_64 wheel/sdist). AudioGen is audiocraft-only → **SFX uses AudioLDM2 (diffusers)** instead. MusicGen runs via `transformers`, not audiocraft. - **HF downloads:** set `HF_HUB_DISABLE_XET=1` (the xet CDN times out on some networks). HF may **429** after many pulls — retry resumes from cache. ## Wiring notes - `tools/local_asset_worker.py` synthesizes voice/music/SFX; `tools/vision_caption.py` runs moondream (batch, load-once) for the Tier-2 director. - `LocalAssetSynthesizer` requires each model path to be a licensed **regular file** (adjacent `.license.txt`, non-blank, not `UNTRACKED`) — `AssetLicensePolicy`. - Do **not** run `tools/run_local_cv_worker.sh` / `tools/run_local_asset_worker.sh` in a certified environment: their `auto` modes `pip install` and can fetch a named YOLO model. - **Licensing blocker for commercial use:** MusicGen (CC-BY-NC), AudioLDM2 (CC-BY-NC-SA) and YOLOv8 (AGPL) are non-commercial/copyleft. Swap in commercially-licensed models/assets before any commercial release. moondream2 (Apache-2.0) and Piper are fine.