12 KiB
GPU / AI-Separator Upgrade — Planning Notes (for me, the assistant)
Status: planned, not built. Last updated: 2026-08-23.
Context: VoiceGrab core (voicegrab.py) currently isolates voice with the
noisereduce STFT pipeline (DSP). This note covers adding a deep-learning
separator (Demucs / UVR / RoFormer family) as an optional "heavy isolation"
mode for clips with music or competing speech.
1. Problem the DSP path doesn't solve
noisereduce (spectral gating) is great for stationary / broadband noise:
fan, hum, room tone, light traffic. It cannot separate:
- background music (guitars, drums, other instruments overlap speech band)
- a second speaker talking
- intermittent noise bursts that contain voice-like formant structure
Spectral gate will either leave the music or carve holes in the target voice (pumping artifacts). For AI-voice-training samples, a source separator (speech vs. non-speech) is the right tool.
2. Candidate models (ranked for our use case)
| Model | Size (ckpt) | What it does | Notes |
|---|---|---|---|
Demucs htdemucs_ft |
~92 MB | 4-stem (vocals/drums/bass/other) | MIT code + weights, maintained, CLI-ready, best "default choice". htdemucs (80 MB) is the original. Fine-tuned _ft preserves voice quality better — prefer it. |
| BS-RoFormer / Mel-Band RoFormer (vocal separation & speech enhancement) | ~50–100 MB | SOTA speech/music separation quality | Checkpoint licensing is research-leaning — must verify before bundling/redistributing. |
| UVR5 MDX-Net / VR Arch checkpoints | ~70–100 MB each | Community vocal-sep models, many flavors | The app is a GUI wrapper; the checkpoints are community-made with mixed licenses — risky to ship. Good for reference only. |
Demucs htdemucs_6s |
~100 MB | 6-stem incl. a dedicated "speech" stem | Heavier; only worth it if 4-stem vocals bleed. |
Decision (revisit at build time): htdemucs_ft as the default AI mode —
permissive license, single dependency (demucs pulls torch), and its
"vocals" stem is exactly what we want. Keep an eye on RoFormer-family
license clarity; if permissive weights appear, they become the recommended
quality option.
License gate: never bundle a checkpoint we can't point to a permissive
license for. Demucs/htdemucs(_ft) = MIT. Everything else: verify first.
3. Hardware & performance reality check
Local machine (voicegrab2 host): no NVIDIA GPU (no nvidia-smi),
8 cores, 14 GB RAM. So:
- Development/testing here will be CPU-only. That's fine for correctness and pipeline integration; all real speed numbers below for "what the user experiences" assume an RTX 30/40/50-class GPU.
- CPU fallback must still work (Windows users without GPU, Macs).
Estimates (5 s clip → 60 s clip):
| Backend | 5 s | 60 s | VRAM | Notes |
|---|---|---|---|---|
| DSP (current) | <1 s | ~10–20 s | ~200 MB RAM | Always works, CPU |
| Demucs CPU | ~10–25 s | ~2–6 min | ~1–2 GB RAM | usable for short samples |
| Demucs CUDA (4–6 GB VRAM) | ~1–3 s | ~15–40 s | ~2 GB peak | comfortable at 6 GB+ |
| 6-stem / RoFormer | similar +30–50% | similar +30–50% | ~2–4 GB | longer context = more VRAM |
Practical guidance to surface in UI copy:
- Target clip length for AI voice samples is 30–120 s → fine on 4 GB VRAM.
- First run downloads model (~92 MB) to user data dir. Show this explicitly.
- CPU mode: set expectations ("a few minutes for a 60 s clip").
4. Architecture plan
4.1 Backend abstraction (replace direct VoiceReducer calls)
# voicegrab.py (sketch)
class Separator:
name: str # "None" | "Noise reduction (fast)" | "AI separation (Demucs)"
def available(self) -> tuple[bool, str]: ... # (ok, reason)
def separate(self, y: np.ndarray, sr: int, progress) -> np.ndarray: ...
class DSSeparator(Separator): # current noisereduce path, unchanged behavior
...
class DemucsSeparator(Separator): # lazy import; GPU-aware
def __init__(self):
import torch, demucs.apply, demucs.pretrained
...
def separate(self, y, sr, progress):
wav = y[None, :, :] # demucs expects (1, C, T) float32
model = self._model() # cached singleton
sources = demucs.apply.apply_model(model, wav, device=self.device, shifts=1)
vocals = sources[model.sources.index("vocals")]
return vocals[0, 0].cpu().numpy() # mono speech
Pipeline order becomes: extract range → optional separation (DSP or AI) →
loudnorm → MP3 320k. The AI path replaces DSP, not stacks on it — running
noisereduce after Demucs is usually counterproductive (double-processing
artifacts). Keep DSP as an independent mode, not a post-filter.
4.2 Device selection
def pick_device():
if torch.cuda.is_available():
return "cuda"
if getattr(torch.backends, "mps", None) and torch.backends.mps.is_available():
return "mps" # Mac bonus, free
return "cpu"
Surface the active device in status text ("Using GPU (CUDA)" / "CPU mode — will be slower").
4.3 GUI changes
- Replace the single "Isolate voice" checkbox with a mode selector:
No processing(just loudness-normalize — current non-isolate path)Noise reduction (fast)— current DSP, default, no downloadsAI voice separation (Demucs)— shows a note: "First use downloads ~92 MB of models. GPU strongly recommended."
- Progress:
demucs.apply.apply_modelsupports a callback; route it into the existingWorker.logsignal. On first use,pretrained.get_model("htdemucs_ft")download should report progress too (it goes throughtorch.hub/urllib— wrap or at least show "Downloading model… (x MB)" in the log line). - Model cache:
appdata_dir()/models/(already the pattern for the old profile idea; reuseappdata_dir()).
4.4 Threading / UX
- Keep the existing
Worker(QThread); add a busy spinner + cancel button. Demucs on CPU can run minutes —subprocess-style cancellation isn't available for in-process torch, so implement a soft cancel: worker checks a flag betweenapply_modelshifts (passshifts=2, check between) — or accept no-cancel for v1 and say so. - Disable the Export button while running (already the pattern).
5. Packaging — the hard part (be honest about it)
torch + demucs is the weight problem. Options, ranked:
A. "AI Pack" as a separate installer component (recommended).
Core app stays small & instant-start (PyInstaller onefile, ~40–80 MB with
PySide6+ffmpeg). Ship a second Inno Setup component ("AI voice separation
(Demucs), ~1.5–2.5 GB after install") that drops pre-wheeled torch+demucs
wheels into %APPDATA%\VoiceGrab\ai\wheels. On first AI-mode use, the app
builds a venv from those wheels (python -m venv ~/.local/share/VoiceGrab/ai_venv)
and imports from it (or runs separation as a sidecar process — see B).
Pros: core users pay nothing; opt-in size; works offline after download.
Cons: most moving parts; needs careful Windows testing.
B. Separation as a sidecar executable (cleanest runtime boundary).
Build separate.py → demucs-cli.exe with PyInstaller (--collect-all torch
→ big, but it's an isolated binary, e.g. VoiceGrab-AI.exe next to the
core app). Core app shells out: VoiceGrab-AI.exe --in clip.wav --out vocals.wav --device cuda. Pros: crash/firewall/GPU driver issues can't take down the GUI;
no venv bootstrapping; trivial to update independently; the demucs CLI
already does exactly this job. Cons: one more build artifact.
→ I'd pick B for Windows, A-or-B for Linux. In practice B generalizes:
it is the demucs CLI wrapped in a branded exe.
C. Single onefile with torch bundled. ❌ Don't. 8 GB+ onefile, 60–120 s
extraction on every start, memory-mapped junk in %TEMP%. Only acceptable as
a "portable full build" niche option.
Linux specifics
- AppImage with torch = same bloat problem → ship
VoiceGrab(core, as today)VoiceGrab-AI(sidecar binary, ~2 GB) or apip install -e .[ai]path.
- A second
.debcomponent with the wheels in/usr/share/voicegrab-aiis also reasonable.
6. Correctness / QA plan
- Unit:
separate()returns float32 mono, shape(T,), dtype invariance, no NaN, SR unchanged (Demucs resamples to 44.1 k internally — re-resample back to 48 k before loudnorm, or just feed 44.1 through to MP3; decide and document. Voice training tools accept 44.1k fine → simplest: feed Demucs output at 44.1k straight to MP3, skip the 48k round-trip). - Regression: existing
kage.mp41–6 s test must produce a valid 320 kbps MP3 in both DSP and AI modes; assertffprobeduration ≈ 5.0 s ± 0.15 s (MP3 encoder padding). - A/B listening test: pick 3 clips (clean speech; speech+fan;
speech+music) — currently only have
kage.mp4; need 2 more fixtures. Record: artifacts? "chipmunk" consonants (classic Demucs failure)? residual music? - GPU smoke: on a CUDA machine:
torch.cuda.is_available(), model load, 60 s clip timing, VRAM peak viatorch.cuda.max_memory_allocated(). Acceptance: < 2× realtime on ≥ 6 GB VRAM. - CPU timeout sanity: 60 s clip on 8-core CPU must finish < 8 min; surface estimated time in UI ("≈ 2–4 min on CPU").
- No-GPU Windows box (test VM): CPU path + model download + offline second run (no re-download).
7. Risks / open questions
- Model redistribution license — only ship MIT (Demucs). Re-check
checkpoint license at build time; put a
THIRD_PARTY.mdin the installer. - Torch × Python 3.14 — current host is 3.14; verify wheel availability at target versions (torch tracks newest Pythons well, but pin in CI).
- CUDA wheel size on Windows —
torchwith cu126 wheels ≈ 2.5 GB (nvidia-cublas etc. are separate wheels). Sidecar build must--collect-all torch nvidia.*. - AMD/Intel users — ROCm wheels exist (Windows + Linux) but are finicky; IPEX for Intel Arc. v1: NVIDIA + CPU only; document the rest as "may work".
- Demucs drift —
demucsis PyTorch-Lightning-free now but its API (apply_model,pretrained.get_model) has shifted across versions. Pindemucs==4.1.xin the AI-pack wheels. - Voice fidelity tradeoff: aggressive separation can thin out sibilance /
breath — exactly what TTS trainers want to keep.
htdemucs_ftis the gentler choice; also consider post-gain before loudnorm (separated vocals sit ~6–10 dB lower than the mix). - Cancel semantics for CPU-mode runs (see 4.4) — decide before build.
8. Suggested build order (if I do this)
SeparationBackendinterface + move existing DSP behind it (small, no deps) — ✅ safe, half a day.- Add
demucsto a dev-only venv extra; wireDemucsSeparatorwith lazy import +pick_device(); test end-to-end on this CPU box withkage.mp4(5 s, expect ~10–30 s). - UI: mode selector + progress + device line. (Half a day.)
- Packaging: prototype B (sidecar) with PyInstaller on this box (CPU
build) to shake out
--collect-allissues; finalize Windows build script. - QA pass from §6; write
THIRD_PARTY.md; bump version to 1.1.0.
Rough total: 2–4 focused days, packaging being the long pole.
9. Files this would touch
voicegrab.py— backend interface,DemucsSeparator, UI mode selectorinstaller/build_win.ps1— AI sidecar build stepinstaller/VoiceGrab.iss— new[Components]/[Files]entriesinstaller/build_linux.sh— sidecar variantrequirements-ai.txt(new) —demucs==4.1.*,torch==2.*README.md— AI mode docs + GPU guidancetests/test_separators.py(new) — §6 automation
Reminder: the "big model" option is a product feature behind a checkbox, never a hard dependency. Core app must stay <100 MB and start in <2 s.