# GPU / AI-Separator Upgrade — Planning Notes (for me, the assistant) Status: **planned, not built.** Last updated: 2026-08-23. Context: VoiceGrab core (`voicegrab.py`) currently isolates voice with the `noisereduce` STFT pipeline (DSP). This note covers adding a deep-learning separator (Demucs / UVR / RoFormer family) as an optional "heavy isolation" mode for clips with music or competing speech. --- ## 1. Problem the DSP path doesn't solve `noisereduce` (spectral gating) is great for **stationary / broadband noise**: fan, hum, room tone, light traffic. It **cannot** separate: - background music (guitars, drums, other instruments overlap speech band) - a second speaker talking - intermittent noise bursts that contain voice-like formant structure Spectral gate will either leave the music or carve holes in the target voice (pumping artifacts). For AI-voice-training samples, a *source separator* (speech vs. non-speech) is the right tool. ## 2. Candidate models (ranked for our use case) | Model | Size (ckpt) | What it does | Notes | |---|---|---|---| | **Demucs `htdemucs_ft`** | ~92 MB | 4-stem (vocals/drums/bass/other) | MIT code + weights, maintained, CLI-ready, best "default choice". `htdemucs` (80 MB) is the original. Fine-tuned `_ft` preserves voice quality better — prefer it. | | **BS-RoFormer / Mel-Band RoFormer** (vocal separation & speech enhancement) | ~50–100 MB | SOTA speech/music separation quality | Checkpoint licensing is **research-leaning** — must verify before bundling/redistributing. | | **UVR5 MDX-Net / VR Arch** checkpoints | ~70–100 MB each | Community vocal-sep models, many flavors | The *app* is a GUI wrapper; the checkpoints are community-made with mixed licenses — risky to ship. Good for *reference* only. | | **Demucs `htdemucs_6s`** | ~100 MB | 6-stem incl. a dedicated "speech" stem | Heavier; only worth it if 4-stem vocals bleed. | Decision (revisit at build time): **`htdemucs_ft` as the default AI mode** — permissive license, single dependency (`demucs` pulls `torch`), and its "vocals" stem is exactly what we want. Keep an eye on RoFormer-family license clarity; if permissive weights appear, they become the recommended quality option. **License gate: never bundle a checkpoint we can't point to a permissive license for.** Demucs/`htdemucs(_ft)` = MIT. Everything else: verify first. ## 3. Hardware & performance reality check Local machine (`voicegrab2` host): **no NVIDIA GPU** (no `nvidia-smi`), 8 cores, 14 GB RAM. So: - Development/testing here will be **CPU-only**. That's fine for *correctness* and pipeline integration; all real speed numbers below for "what the user experiences" assume an RTX 30/40/50-class GPU. - CPU fallback must still work (Windows users without GPU, Macs). Estimates (5 s clip → 60 s clip): | Backend | 5 s | 60 s | VRAM | Notes | |---|---|---|---|---| | DSP (current) | <1 s | ~10–20 s | ~200 MB RAM | Always works, CPU | | Demucs CPU | ~10–25 s | ~2–6 min | ~1–2 GB RAM | usable for short samples | | Demucs CUDA (4–6 GB VRAM) | ~1–3 s | ~15–40 s | ~2 GB peak | comfortable at 6 GB+ | | 6-stem / RoFormer | similar +30–50% | similar +30–50% | ~2–4 GB | longer context = more VRAM | Practical guidance to surface in UI copy: - Target clip length for AI voice samples is 30–120 s → **fine on 4 GB VRAM**. - First run downloads model (~92 MB) to user data dir. Show this explicitly. - CPU mode: set expectations ("a few minutes for a 60 s clip"). ## 4. Architecture plan ### 4.1 Backend abstraction (replace direct `VoiceReducer` calls) ```python # voicegrab.py (sketch) class Separator: name: str # "None" | "Noise reduction (fast)" | "AI separation (Demucs)" def available(self) -> tuple[bool, str]: ... # (ok, reason) def separate(self, y: np.ndarray, sr: int, progress) -> np.ndarray: ... class DSSeparator(Separator): # current noisereduce path, unchanged behavior ... class DemucsSeparator(Separator): # lazy import; GPU-aware def __init__(self): import torch, demucs.apply, demucs.pretrained ... def separate(self, y, sr, progress): wav = y[None, :, :] # demucs expects (1, C, T) float32 model = self._model() # cached singleton sources = demucs.apply.apply_model(model, wav, device=self.device, shifts=1) vocals = sources[model.sources.index("vocals")] return vocals[0, 0].cpu().numpy() # mono speech ``` Pipeline order becomes: **extract range → optional separation (DSP or AI) → loudnorm → MP3 320k**. The AI path *replaces* DSP, not stacks on it — running `noisereduce` after Demucs is usually counterproductive (double-processing artifacts). Keep DSP as an independent mode, not a post-filter. ### 4.2 Device selection ```python def pick_device(): if torch.cuda.is_available(): return "cuda" if getattr(torch.backends, "mps", None) and torch.backends.mps.is_available(): return "mps" # Mac bonus, free return "cpu" ``` Surface the active device in status text ("Using GPU (CUDA)" / "CPU mode — will be slower"). ### 4.3 GUI changes - Replace the single "Isolate voice" checkbox with a **mode selector**: 1. `No processing` (just loudness-normalize — current non-isolate path) 2. `Noise reduction (fast)` — current DSP, default, no downloads 3. `AI voice separation (Demucs)` — shows a note: "First use downloads ~92 MB of models. GPU strongly recommended." - Progress: `demucs.apply.apply_model` supports a callback; route it into the existing `Worker.log` signal. On first use, `pretrained.get_model("htdemucs_ft")` download should report progress too (it goes through `torch.hub`/`urllib` — wrap or at least show "Downloading model… (x MB)" in the log line). - Model cache: `appdata_dir()/models/` (already the pattern for the old profile idea; reuse `appdata_dir()`). ### 4.4 Threading / UX - Keep the existing `Worker(QThread)`; add a **busy spinner + cancel button**. Demucs on CPU can run minutes — `subprocess`-style cancellation isn't available for in-process torch, so implement a soft cancel: worker checks a flag between `apply_model` shifts (pass `shifts=2`, check between) — or accept no-cancel for v1 and say so. - Disable the Export button while running (already the pattern). ## 5. Packaging — the hard part (be honest about it) `torch` + `demucs` is the weight problem. Options, ranked: **A. "AI Pack" as a separate installer component (recommended).** Core app stays small & instant-start (PyInstaller onefile, ~40–80 MB with PySide6+ffmpeg). Ship a second Inno Setup *component* ("AI voice separation (Demucs), ~1.5–2.5 GB after install") that drops pre-wheeled `torch`+`demucs` wheels into `%APPDATA%\VoiceGrab\ai\wheels`. On first AI-mode use, the app builds a venv from those wheels (`python -m venv ~/.local/share/VoiceGrab/ai_venv`) and imports from it (or runs separation as a **sidecar process** — see B). Pros: core users pay nothing; opt-in size; works offline after download. Cons: most moving parts; needs careful Windows testing. **B. Separation as a sidecar executable (cleanest runtime boundary).** Build `separate.py` → `demucs-cli.exe` with PyInstaller (`--collect-all torch` → big, but it's an *isolated* binary, e.g. `VoiceGrab-AI.exe` next to the core app). Core app shells out: `VoiceGrab-AI.exe --in clip.wav --out vocals.wav --device cuda`. Pros: crash/firewall/GPU driver issues can't take down the GUI; no venv bootstrapping; trivial to update independently; the `demucs` CLI already does exactly this job. Cons: one more build artifact. → **I'd pick B for Windows, A-or-B for Linux.** In practice B generalizes: it *is* the demucs CLI wrapped in a branded exe. **C. Single onefile with torch bundled.** ❌ Don't. 8 GB+ onefile, 60–120 s extraction on every start, memory-mapped junk in `%TEMP%`. Only acceptable as a "portable full build" niche option. ### Linux specifics - AppImage with torch = same bloat problem → ship `VoiceGrab` (core, as today) + `VoiceGrab-AI` (sidecar binary, ~2 GB) or a `pip install -e .[ai]` path. - A second `.deb` component with the wheels in `/usr/share/voicegrab-ai` is also reasonable. ## 6. Correctness / QA plan 1. **Unit:** `separate()` returns float32 mono, shape `(T,)`, dtype invariance, no NaN, SR unchanged (Demucs resamples to 44.1 k internally — **re-resample back to 48 k before loudnorm**, or just feed 44.1 through to MP3; decide and document. Voice training tools accept 44.1k fine → **simplest: feed Demucs output at 44.1k straight to MP3**, skip the 48k round-trip). 2. **Regression:** existing `kage.mp4` 1–6 s test must produce a valid 320 kbps MP3 in both DSP and AI modes; assert `ffprobe` duration ≈ 5.0 s ± 0.15 s (MP3 encoder padding). 3. **A/B listening test:** pick 3 clips (clean speech; speech+fan; speech+music) — currently only have `kage.mp4`; need 2 more fixtures. Record: artifacts? "chipmunk" consonants (classic Demucs failure)? residual music? 4. **GPU smoke:** on a CUDA machine: `torch.cuda.is_available()`, model load, 60 s clip timing, VRAM peak via `torch.cuda.max_memory_allocated()`. Acceptance: < 2× realtime on ≥ 6 GB VRAM. 5. **CPU timeout sanity:** 60 s clip on 8-core CPU must finish < 8 min; surface estimated time in UI ("≈ 2–4 min on CPU"). 6. **No-GPU Windows box** (test VM): CPU path + model download + offline second run (no re-download). ## 7. Risks / open questions - **Model redistribution license** — only ship MIT (Demucs). Re-check checkpoint license at build time; put a `THIRD_PARTY.md` in the installer. - **Torch × Python 3.14** — current host is 3.14; verify wheel availability at target versions (torch tracks newest Pythons well, but pin in CI). - **CUDA wheel size on Windows** — `torch` with cu126 wheels ≈ 2.5 GB (nvidia-cublas etc. are separate wheels). Sidecar build must `--collect-all torch nvidia.*`. - **AMD/Intel users** — ROCm wheels exist (Windows + Linux) but are finicky; IPEX for Intel Arc. v1: NVIDIA + CPU only; document the rest as "may work". - **Demucs drift** — `demucs` is PyTorch-Lightning-free now but its API (`apply_model`, `pretrained.get_model`) has shifted across versions. **Pin `demucs==4.1.x` in the AI-pack wheels.** - **Voice fidelity tradeoff:** aggressive separation can thin out sibilance / breath — exactly what TTS trainers want to *keep*. `htdemucs_ft` is the gentler choice; also consider post-gain before loudnorm (separated vocals sit ~6–10 dB lower than the mix). - **Cancel semantics** for CPU-mode runs (see 4.4) — decide before build. ## 8. Suggested build order (if I do this) 1. `SeparationBackend` interface + move existing DSP behind it (small, no deps) — ✅ safe, half a day. 2. Add `demucs` to a **dev-only** venv extra; wire `DemucsSeparator` with lazy import + `pick_device()`; test end-to-end on this CPU box with `kage.mp4` (5 s, expect ~10–30 s). 3. UI: mode selector + progress + device line. (Half a day.) 4. Packaging: prototype **B** (sidecar) with PyInstaller on this box (CPU build) to shake out `--collect-all` issues; finalize Windows build script. 5. QA pass from §6; write `THIRD_PARTY.md`; bump version to 1.1.0. Rough total: **2–4 focused days**, packaging being the long pole. ## 9. Files this would touch - `voicegrab.py` — backend interface, `DemucsSeparator`, UI mode selector - `installer/build_win.ps1` — AI sidecar build step - `installer/VoiceGrab.iss` — new `[Components]`/`[Files]` entries - `installer/build_linux.sh` — sidecar variant - `requirements-ai.txt` (new) — `demucs==4.1.*`, `torch==2.*` - `README.md` — AI mode docs + GPU guidance - `tests/test_separators.py` (new) — §6 automation --- *Reminder: the "big model" option is a **product feature behind a checkbox**, never a hard dependency. Core app must stay <100 MB and start in <2 s.*