Add AI voice separation (Demucs) as toggleable mode, v1.1.0

Implements the GPU-upgrade.md plan:

- Separator backend interface (DSSeparator = existing noisereduce path,
  DemucsSeparator = htdemucs_ft AI separation) — one mode at a time,
  never stacked
- DemucsSeparator: lazy imports (core app stays dep-free), device
  auto-pick cuda -> mps -> cpu, model cache under appdata/models,
  chunk progress reporting, soft cancel between chunks
- GUI: 'Isolate voice' checkbox replaced by a 3-way mode selector
  (fast DSP / AI Demucs / none), DSP strength controls follow the
  selection, AI radio disabled with guidance when torch+demucs are
  missing, Cancel button for in-flight AI runs
- CLI: --mode none|dsp|ai (default dsp, unchanged behavior)
- requirements-ai.txt (torch 2.x + demucs 4.1.x), THIRD_PARTY.md,
  README updates, optional AI pack flags in installer scripts,
  version bumped to 1.1.0
- tests/test_separators.py: 10 tests incl. graceful-degradation and
  full-pipeline regression (5.0 s ± 0.15 s, ~320 kbps) in all modes
This commit is contained in:
VoiceGrab 2026-08-25 20:19:18 -06:00
parent 62c4659486
commit c2126f2045
10 changed files with 741 additions and 55 deletions

6
.gitignore vendored Normal file
View File

@ -0,0 +1,6 @@
.venv/
__pycache__/
*.pyc
dist/
build/
third_party/

View File

@ -6,25 +6,39 @@ Cut a time range out of any video file and export a clean, loudness-normalized
``` ```
Voice / audio file: [ kage.mp4 ] [Open…] Voice / audio file: [ kage.mp4 ] [Open…]
Start: 0:10.00 Stop: 0:35.00 [Load preview] [▶ Play range] Start: 0:10.00 Stop: 0:35.00 [Load preview] [▶ Play range]
[x] Isolate voice (reduce background noise) strength: 80% Voice processing (choose one):
(•) Fast noise reduction (CPU) — fan / hum / room tone; no downloads
( ) AI voice separation (Demucs) — music & second speakers; GPU if available
( ) No processing — just loudness-normalize
strength: 80% [x] stationary noise
Output MP3: [ .../kage_voicegrab.mp3 ] [Browse…] Output MP3: [ .../kage_voicegrab.mp3 ] [Browse…]
────────────────────────────────────────────────── ──────────────────────────────────────────────────
[ Export MP3 ] [ Cancel ] [ Export MP3 ]
``` ```
## Features ## Features
- Works with any container ffmpeg understands: mp4, mkv, mov, webm, avi, flv, m4a… - Works with any container ffmpeg understands: mp4, mkv, mov, webm, avi, flv, m4a…
- **Isolate voice** checkbox: STFT-based noise reduction (runs on CPU, no - **Voice processing — pick one mode** (toggle, mutually exclusive):
downloads, no GPU). Strength slider + a "stationary noise" mode for - **Fast noise reduction (CPU, default)** — STFT spectral gating via
constant fans/hum. `noisereduce`. Great for fan, hum, room tone. No downloads, no GPU;
strength slider + a "stationary noise" mode for constant fans/hum.
- **AI voice separation (Demucs)** — deep-learning 4-stem separation
(we keep the *vocals* stem). Handles background music and a second
speaker, which spectral gating cannot. Optional install; uses an
NVIDIA GPU (CUDA) or Apple GPU if present, otherwise CPU (slower).
- **No processing** — just loudness-normalize.
- Export is **320 kbps MP3**, loudness-normalized to 16 LUFS (broadcast - Export is **320 kbps MP3**, loudness-normalized to 16 LUFS (broadcast
reference — good levels for voice-training datasets). reference — good levels for voice-training datasets).
- "▶ Play range" lets you audition the exact segment before saving. - "▶ Play range" lets you audition the exact segment before saving.
- AI runs report chunk progress in the status bar and can be **cancelled**
between chunks.
- Also a headless CLI (see below) for scripting/batch work. - Also a headless CLI (see below) for scripting/batch work.
## Requirements ## Requirements
- Python 3.10+ - Python 3.10+
- ffmpeg on your PATH (bundled automatically in installer builds) - ffmpeg on your PATH (bundled automatically in installer builds)
- Optional, only for AI voice separation: `pip install -r requirements-ai.txt`
(pulls `torch` + `demucs`, ~1.52.5 GB with CUDA wheels)
## Run from source ## Run from source
```bash ```bash
@ -34,11 +48,44 @@ python3 -m venv .venv
.venv/bin/python voicegrab.py .venv/bin/python voicegrab.py
``` ```
## AI voice separation (optional — GPU recommended)
```bash
.venv/bin/pip install -r requirements-ai.txt
```
- Pick **AI voice separation (Demucs)** in the GUI. The mode is disabled
(with an explanation) until `torch` + `demucs` are installed — the core
app never requires them.
- **First use downloads the `htdemucs_ft` model (~90 MB)** to the app data
dir (`<appdata>/VoiceGrab/models/`) and caches it — later runs are
instant. GPU (NVIDIA CUDA / Apple) is used automatically when present,
otherwise CPU.
- Performance expectations (5 s → 60 s clip):
- GPU (≥ 6 GB VRAM): ~13 s → ~1540 s
- CPU (8 cores): ~1025 s → a few minutes
- **Cancel** stops an AI run at the next chunk boundary. Fast DSP runs
finish in seconds and are not cancellable (by design).
- Quality note: aggressive separation can thin sibilance/breath — for voice
*training* samples the fast DSP mode is usually enough; reach for AI mode
when there's music or a competing speaker.
- Licensing: Demucs code is MIT and the `htdemucs_ft` weights are MIT
(see `THIRD_PARTY.md`).
## CLI mode ## CLI mode
```bash ```bash
voicegrab.py --cli input.mp4 10 35 output.mp3 voicegrab.py --cli input.mp4 10 35 output.mp3 # default: fast DSP
voicegrab.py --cli input.mp4 10 35 output.mp3 --mode dsp # fast noise reduction
voicegrab.py --cli input.mp4 10 35 output.mp3 --mode ai # Demucs (requirements-ai.txt)
voicegrab.py --cli input.mp4 10 35 output.mp3 --mode none # just loudness-normalize
``` ```
## Tests
```bash
.venv/bin/python tests/test_separators.py # or: .venv/bin/pytest tests/
```
Covers: DSP shape/dtype/NaN invariants, AI graceful degradation without
torch, device selection, AI unit separation (44.1 kHz output), and the full
CLI pipeline in all three modes (valid 320 kbps MP3, duration 5.0 s ± 0.15 s).
## Windows installer (single .exe + installer) ## Windows installer (single .exe + installer)
1. `pip install pyinstaller` 1. `pip install pyinstaller`
2. `python voicegrab.py.spec` is not needed — run: 2. `python voicegrab.py.spec` is not needed — run:
@ -55,17 +102,26 @@ voicegrab.py --cli input.mp4 10 35 output.mp3
(Copy `dist\VoiceGrab.exe` to `installer\` first — the script assumes (Copy `dist\VoiceGrab.exe` to `installer\` first — the script assumes
`dist\VoiceGrab.exe` relative to the project root.) `dist\VoiceGrab.exe` relative to the project root.)
To ship the AI mode inside the frozen exe (bigger, ~2 GB), install
`requirements-ai.txt` into the venv first and add
`--collect-all demucs --collect-all torch` to the PyInstaller command.
## Linux ## Linux
- **AppImage**: `installer/build_linux.sh` (needs PyInstaller + - **AppImage**: `installer/build_linux.sh` (needs PyInstaller +
`linuxdeploy` or `appimagetool` + your ffmpeg in PATH). `linuxdeploy` or `appimagetool` + your ffmpeg in PATH).
- **Debian/Ubuntu**: just run the PyInstaller binary, or install ffmpeg via - **Debian/Ubuntu**: just run the PyInstaller binary, or install ffmpeg via
apt and run from source. A `.desktop` entry template is in `installer/`. apt and run from source. A `.desktop` entry template is in `installer/`.
- **Flatpak** is also a fine route if you want it in your store. - **Flatpak** is also a fine route if you want it in your store.
- AI mode: `WITH_AI=1 ./installer/build_linux.sh` bundles `demucs`/`torch`
into the binary (bigger build).
## Notes on voice-isolation quality ## Notes on voice-isolation quality
The built-in reducer is classic DSP (spectral gating / STFT noise - **Fast mode** is classic DSP (spectral gating / STFT noise estimation) —
estimation) — great for fan, hum, light room tone. For music or heavy great for fan, hum, light room tone; it can't remove music or a second
background speech, a deep-learning model (e.g. Demucs / UVR) gives speaker.
better separation but costs ~12 GB of downloads and runs far slower - **AI mode** (Demucs `htdemucs_ft`) is a source separator — the right tool
without a GPU. For voice *training* samples, the DSP route is usually for speech-over-music or competing voices. It's a product feature behind a
more than enough, and it keeps the installer small and startup instant. radio button, never a hard dependency: the core app stays small and starts
in < 2 s without `torch`/`demucs` installed.
- The AI path *replaces* the DSP path — one or the other, never stacked
(double-processing adds artifacts).

27
THIRD_PARTY.md Normal file
View File

@ -0,0 +1,27 @@
# Third-Party Components
## Core (always)
| Component | License | Notes |
|---|---|---|
| PySide6 (Qt for Python) | LGPL v3 / GPL | GUI toolkit |
| numpy, scipy, soundfile | BSD / BSD-3 / 3-clause BSD | audio math & IO |
| noisereduce | Unlicense / MIT (see repo) | DSP voice isolation ("fast" mode) |
| ffmpeg / ffprobe | LGPL v2.1+ (GPL builds exist) | media I/O; bundled binary must match license |
## Optional — AI voice separation (only if the user installs `requirements-ai.txt`)
| Component | License | Notes |
|---|---|---|
| demucs 4.1.x (Meta/Facebook) | Code: MIT | source separation engine |
| PyTorch (torch 2.x) | BSD-3 | deep-learning runtime |
| htdemucs_ft weights (~90 MB) | **MIT** (per the model card of `adefossez/HTDemucs-ft` on Hugging Face) | downloaded on first use from Hugging Face, cached in the user data dir — never bundled |
| huggingface_hub | Apache-2.0 | model download plumbing |
### License gate (project policy)
We only use checkpoints whose license we can point to. Demucs code = MIT,
`htdemucs_ft` weights = MIT (model card). They are only ever *downloaded by*
the *user* on first use — never bundled by us. RoFormer / UVR community
checkpoints have mixed or research-only licensing and are therefore **not**
used.

View File

@ -3,7 +3,7 @@
; Compile: ISCC.exe installer\VoiceGrab.iss (from project root, after build_win.ps1) ; Compile: ISCC.exe installer\VoiceGrab.iss (from project root, after build_win.ps1)
#define MyAppName "VoiceGrab" #define MyAppName "VoiceGrab"
#define MyAppVersion "1.0.0" #define MyAppVersion "1.1.0"
#define MyAppExeName "VoiceGrab.exe" #define MyAppExeName "VoiceGrab.exe"
[Setup] [Setup]

View File

@ -1,5 +1,6 @@
#!/usr/bin/env bash #!/usr/bin/env bash
# VoiceGrab — Linux build (produces dist/VoiceGrab-<arch> binary; AppImage if linuxdeploy available) # VoiceGrab — Linux build (produces dist/VoiceGrab-<arch> binary; AppImage if linuxdeploy available)
# WITH_AI=1 ./installer/build_linux.sh → also bundle torch+demucs (bigger build)
set -euo pipefail set -euo pipefail
cd "$(dirname "$0")/.." cd "$(dirname "$0")/.."
@ -8,6 +9,13 @@ PY="${PYTHON:-.venv/bin/python}"
command -v ffmpeg >/dev/null || { echo "ffmpeg not found in PATH (needed to resolve the binary)"; exit 1; } command -v ffmpeg >/dev/null || { echo "ffmpeg not found in PATH (needed to resolve the binary)"; exit 1; }
"$PY" -m pip install pyinstaller "$PY" -m pip install pyinstaller
if [ "${WITH_AI:-0}" = "1" ]; then
echo "Building with AI voice separation (torch+demucs)…"
"$PY" -m pip install -r requirements-ai.txt
AI_FLAGS=(--collect-all demucs --collect-all torch)
else
AI_FLAGS=()
fi
FFMPEG_BIN="$(command -v ffmpeg)" FFMPEG_BIN="$(command -v ffmpeg)"
FFPROBE_BIN="$(command -v ffprobe || true)" FFPROBE_BIN="$(command -v ffprobe || true)"
@ -20,6 +28,7 @@ fi
"$PY" -m PyInstaller --noconfirm --onefile --windowed --name VoiceGrab \ "$PY" -m PyInstaller --noconfirm --onefile --windowed --name VoiceGrab \
--add-binary "$FFMPEG_BIN:." \ --add-binary "$FFMPEG_BIN:." \
"${EXTRA[@]}" \ "${EXTRA[@]}" \
"${AI_FLAGS[@]}" \
--add-data "assets/icon.png:assets" \ --add-data "assets/icon.png:assets" \
--icon assets/icon.ico \ --icon assets/icon.ico \
voicegrab.py voicegrab.py

View File

@ -1,5 +1,10 @@
# VoiceGrab — Windows build script # VoiceGrab — Windows build script
# Usage (from project root): powershell -ExecutionPolicy Bypass -File installer\build_win.ps1 # Usage (from project root): powershell -ExecutionPolicy Bypass -File installer\build_win.ps1 [-BuildAI]
# -BuildAI also bundle the optional AI voice separation (torch+demucs;
# exe grows by ~2 GB). Core users don't need it.
param(
[switch]$BuildAI
)
$ErrorActionPreference = "Stop" $ErrorActionPreference = "Stop"
$root = Split-Path -Parent $PSScriptRoot $root = Split-Path -Parent $PSScriptRoot
Set-Location $root Set-Location $root
@ -10,6 +15,16 @@ if (-not (Test-Path ".venv\Scripts\python.exe")) {
} }
& ".venv\Scripts\python.exe" -m pip install --upgrade pip & ".venv\Scripts\python.exe" -m pip install --upgrade pip
& ".venv\Scripts\python.exe" -m pip install -r requirements.txt pyinstaller & ".venv\Scripts\python.exe" -m pip install -r requirements.txt pyinstaller
if ($BuildAI) {
Write-Host "Building with AI voice separation (torch+demucs)…"
& ".venv\Scripts\python.exe" -m pip install -r requirements-ai.txt
}
# AI-mode collection flags (no-op when not building the AI pack)
$AI_FLAGS = @()
if ($BuildAI) {
$AI_FLAGS = @("--collect-all", "demucs", "--collect-all", "torch")
}
# 2. Grab a static ffmpeg build (BtbN build — includes libmp3lame) # 2. Grab a static ffmpeg build (BtbN build — includes libmp3lame)
$ffdir = "third_party\ffmpeg\bin" $ffdir = "third_party\ffmpeg\bin"
@ -33,6 +48,7 @@ if (-not (Test-Path "$ffdir\ffmpeg.exe")) {
--add-binary "$ffdir\ffprobe.exe;." ` --add-binary "$ffdir\ffprobe.exe;." `
--add-data "assets\icon.png;assets" ` --add-data "assets\icon.png;assets" `
--icon "assets\icon.ico" ` --icon "assets\icon.ico" `
@AI_FLAGS `
voicegrab.py voicegrab.py
Write-Host "Built dist\VoiceGrab.exe" Write-Host "Built dist\VoiceGrab.exe"

19
requirements-ai.txt Normal file
View File

@ -0,0 +1,19 @@
# VoiceGrab — OPTIONAL AI voice separation (heavy mode)
#
# The core app does NOT need this file; it only enables the
# "AI voice separation (Demucs)" mode.
#
# Install:
# pip install -r requirements-ai.txt
#
# Notes:
# - `torch` from PyPI on Windows/Linux ships CUDA wheels (~2 GB with the
# nvidia-* deps). CPU-only machines can install the smaller CPU build:
# pip install torch --index-url https://download.pytorch.org/whl/cpu
# - The Demucs `htdemucs_ft` model (~90 MB) downloads on first use and is
# cached under the app data dir (<appdata>/VoiceGrab/models/).
# - Licenses: see THIRD_PARTY.md (code MIT; weights usable per Meta's
# model license — fine for personal use).
torch>=2.4,<3
demucs>=4.1,<4.2

252
tests/test_separators.py Normal file
View File

@ -0,0 +1,252 @@
#!/usr/bin/env python3
"""VoiceGrab separator tests (GPU-upgrade.md §6 automation).
Run with the project venv:
.venv/bin/python tests/test_separators.py
or with pytest:
.venv/bin/pytest tests/
The AI (Demucs) end-to-end test needs the model cached or a network
connection (~90 MB, first run only) and takes ~15 s on CPU.
"""
from __future__ import annotations
import math
import os
import subprocess
import sys
import tempfile
import wave
HERE = os.path.dirname(os.path.abspath(__file__))
ROOT = os.path.dirname(HERE)
sys.path.insert(0, ROOT)
import numpy as np # noqa: E402
import voicegrab as vg # noqa: E402
# ---------------------------------------------------------------------------
# fixtures / helpers
# ---------------------------------------------------------------------------
def _synthetic_voice(n: int, sr: int = 48000) -> np.ndarray:
"""Speech-like signal (130 Hz fundamental + harmonics, syllable AM)."""
t = np.arange(n) / sr
sig = np.zeros_like(t)
for h in range(1, 31):
sig += (0.6 / h) * np.sin(2 * np.pi * 130.0 * h * t + h * 0.7)
sig *= 0.35 + 0.65 * (0.5 + 0.5 * np.sin(2 * np.pi * 4.5 * t))
sig /= np.abs(sig).max()
rng = np.random.default_rng(7)
noise = 0.5 * np.cumsum(rng.normal(0, 1, len(t)))
noise /= np.abs(noise).max()
return np.clip(0.7 * sig + 0.5 * noise, -1, 1).astype(np.float32)
def _write_test_mp3(path: str, seconds: int = 5) -> str:
"""Create a seconds-long mp3 with ffmpeg (voice-like tone + noise)."""
cmd = [
"ffmpeg", "-hide_banner", "-v", "error", "-y",
"-f", "lavfi", "-i", f"sine=frequency=220:duration={seconds}",
"-f", "lavfi", "-i", f"anoisesrc=d={seconds}:c=pink:a=0.3",
"-filter_complex",
"[0][1]amix=inputs=2:weights=1 0.5,atrim=0:" + str(seconds) + ",asetpts=PTS-STARTPTS[out]",
"-map", "[out]", "-c:a", "libmp3lame", path,
]
subprocess.run(cmd, check=True, capture_output=True)
return path
def _mp3_duration(path: str) -> float:
out = subprocess.run(
["ffprobe", "-v", "error", "-show_entries", "format=duration",
"-of", "default=nw=1:nk=1", path],
capture_output=True, text=True, check=True).stdout.strip()
return float(out)
def _mp3_bitrate(path: str) -> int:
out = subprocess.run(
["ffprobe", "-v", "error", "-show_entries", "format=bit_rate",
"-of", "default=nw=1:nk=1", path],
capture_output=True, text=True, check=True).stdout.strip()
return int(out)
def _ai_available() -> bool:
return vg._ai_deps_present()
# ---------------------------------------------------------------------------
# unit tests
# ---------------------------------------------------------------------------
def test_modes_constants():
assert vg.MODE_NONE == "none"
assert vg.MODE_DSP == "dsp"
assert vg.MODE_AI == "ai"
assert vg.AI_MODEL_NAME == "htdemucs_ft"
def test_dsp_available_and_shape():
sep = vg.DSSeparator(intensity=0.8, stationary=True)
ok, reason = sep.available()
assert ok, f"DSP should be available in this venv: {reason}"
y = _synthetic_voice(48000)
out, sr = sep.separate(y, 48000, lambda m: None)
assert sr == 48000, "DSP must not change the sample rate"
assert out.dtype == np.float32
assert out.ndim == 1 and len(out) == len(y)
assert np.isfinite(out).all(), "no NaN/Inf allowed"
assert float(np.sqrt(np.mean(out**2))) > 0, "must not be all-silence"
def test_ai_available_report():
sep = vg.DemucsSeparator()
ok, reason = sep.available()
if _ai_available():
assert ok
else:
assert not ok
assert "requirements-ai.txt" in reason
def test_ai_graceful_without_torch():
"""DemucsSeparator.available() must degrade gracefully when torch is absent."""
code = (
"import sys, types\n"
"class Blocker:\n"
" def find_spec(self, name, path=None, target=None):\n"
" if name == 'torch' or name.startswith('torch.'):\n"
" raise ModuleNotFoundError('blocked: ' + name)\n"
" return None\n"
"sys.meta_path.insert(0, Blocker())\n"
"sys.path.insert(0, %r)\n"
"import voicegrab as vg\n"
"ok, reason = vg.DemucsSeparator().available()\n"
"assert not ok, 'should report unavailable when torch is blocked'\n"
"assert 'requirements-ai.txt' in reason\n"
"print('graceful-degradation OK')\n"
) % ROOT
r = subprocess.run([sys.executable, "-c", code],
capture_output=True, text=True, timeout=120)
assert r.returncode == 0, r.stderr
assert "graceful-degradation OK" in r.stdout
def test_ai_pick_device_order():
if not _ai_available():
print(" (skipped: torch/demucs not installed)")
return
sep = vg.DemucsSeparator(device="cuda")
assert sep._pick_device() == "cuda", "explicit device must be honored"
auto = vg.DemucsSeparator()._pick_device()
assert auto in ("cpu", "cuda", "mps")
def test_ai_separate_unit():
if not _ai_available():
print(" (skipped: torch/demucs not installed)")
return
sep = vg.DemucsSeparator()
y = _synthetic_voice(3 * 48000) # 3 s @ 48 kHz
out, sr = sep.separate(y, 48000, lambda m: None)
assert sr == 44100, "Demucs runs at 44.1 kHz"
assert out.dtype == np.float32 and out.ndim == 1
assert np.isfinite(out).all()
expected = 3 * 44100
assert abs(len(out) - expected) <= 44100 * 0.02, f"duration drift: {len(out)} vs {expected}"
assert float(np.sqrt(np.mean(out**2))) > 0, "must not be all-silence"
# ---------------------------------------------------------------------------
# regression: full CLI pipeline (GPU-upgrade.md §6.2)
# ---------------------------------------------------------------------------
def test_cli_dsp_regression():
with tempfile.TemporaryDirectory(prefix="voicegrab-test-") as wd:
src = os.path.join(wd, "in.mp3")
out = os.path.join(wd, "out.mp3")
_write_test_mp3(src, 5)
r = subprocess.run(
[sys.executable, os.path.join(ROOT, "voicegrab.py"),
"--cli", src, "0", "5", out, "--mode", "dsp"],
capture_output=True, text=True, timeout=300)
assert r.returncode == 0, r.stderr
assert os.path.exists(out)
dur = _mp3_duration(out)
assert abs(dur - 5.0) <= 0.15, f"duration {dur} not within 5.0 ± 0.15 s"
br = _mp3_bitrate(out)
assert 300_000 <= br <= 350_000, f"expected ~320 kbps, got {br}"
def test_cli_ai_regression():
if not _ai_available():
print(" (skipped: torch/demucs not installed)")
return
with tempfile.TemporaryDirectory(prefix="voicegrab-test-") as wd:
src = os.path.join(wd, "in.mp3")
out = os.path.join(wd, "out.mp3")
_write_test_mp3(src, 5)
r = subprocess.run(
[sys.executable, os.path.join(ROOT, "voicegrab.py"),
"--cli", src, "0", "5", out, "--mode", "ai"],
capture_output=True, text=True, timeout=600)
assert r.returncode == 0, r.stderr
assert os.path.exists(out)
dur = _mp3_duration(out)
assert abs(dur - 5.0) <= 0.15, f"duration {dur} not within 5.0 ± 0.15 s"
br = _mp3_bitrate(out)
assert 300_000 <= br <= 350_000, f"expected ~320 kbps, got {br}"
def test_cli_none_mode():
with tempfile.TemporaryDirectory(prefix="voicegrab-test-") as wd:
src = os.path.join(wd, "in.mp3")
out = os.path.join(wd, "out.mp3")
_write_test_mp3(src, 5)
r = subprocess.run(
[sys.executable, os.path.join(ROOT, "voicegrab.py"),
"--cli", src, "0", "5", out, "--mode", "none"],
capture_output=True, text=True, timeout=300)
assert r.returncode == 0, r.stderr
assert abs(_mp3_duration(out) - 5.0) <= 0.15
def test_cli_rejects_bad_mode():
with tempfile.TemporaryDirectory(prefix="voicegrab-test-") as wd:
src = os.path.join(wd, "in.mp3")
_write_test_mp3(src, 5)
r = subprocess.run(
[sys.executable, os.path.join(ROOT, "voicegrab.py"),
"--cli", src, "0", "5", os.path.join(wd, "o.mp3"), "--mode", "bogus"],
capture_output=True, text=True, timeout=120)
assert r.returncode == 2
assert "none | dsp | ai" in r.stdout
# ---------------------------------------------------------------------------
def main() -> int:
tests = [v for k, v in sorted(globals().items()) if k.startswith("test_")]
failed = 0
for t in tests:
name = t.__name__
try:
t()
print(f"PASS {name}")
except AssertionError as exc:
failed += 1
print(f"FAIL {name}: {exc}")
except Exception as exc: # noqa: BLE001
failed += 1
print(f"ERROR {name}: {type(exc).__name__}: {exc}")
total = len(tests)
print(f"\n{total - failed}/{total} passed")
return 1 if failed else 0
if __name__ == "__main__":
raise SystemExit(main())

View File

@ -3,22 +3,27 @@
- Pick an input video (mp4, mkv, webm, avi, mov, ...) - Pick an input video (mp4, mkv, webm, avi, mov, ...)
- Set start / stop timestamps - Set start / stop timestamps
- Optional: isolate voice with noise reduction (no GPU or big models needed) - Optional voice processing (one of, mutually exclusive):
* fast noise reduction (CPU DSP via noisereduce no downloads)
* AI voice separation (Demucs htdemucs_ft GPU if available, optional install)
* no processing (loudness-normalize only)
- Export MP3 (320 kbps, loudness-normalized to -16 LUFS good for AI voice training) - Export MP3 (320 kbps, loudness-normalized to -16 LUFS good for AI voice training)
- Waveform preview with the selected range highlighted - Waveform preview with the selected range highlighted
""" """
from __future__ import annotations from __future__ import annotations
import math
import os import os
import platform import platform
import shutil import shutil
import subprocess import subprocess
import sys import sys
import tempfile import tempfile
import threading
import wave import wave
APP_NAME = "VoiceGrab" APP_NAME = "VoiceGrab"
APP_VERSION = "1.0.0" APP_VERSION = "1.1.0"
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# Resource / executable helpers # Resource / executable helpers
@ -33,7 +38,7 @@ def app_dir() -> str:
def appdata_dir() -> str: def appdata_dir() -> str:
"""Per-user directory for saved model noise profiles.""" """Per-user directory for cached AI models (and any saved profiles)."""
if platform.system() == "Windows": if platform.system() == "Windows":
base = os.environ.get("APPDATA", os.path.expanduser("~")) base = os.environ.get("APPDATA", os.path.expanduser("~"))
elif platform.system() == "Darwin": elif platform.system() == "Darwin":
@ -179,6 +184,203 @@ class VoiceReducer:
return nr.reduce_noise(**kwargs) return nr.reduce_noise(**kwargs)
# ---------------------------------------------------------------------------
# Separation backends — one mode at a time (DSP and AI are alternatives, not
# stacked: running noisereduce after Demucs adds double-processing artifacts)
# ---------------------------------------------------------------------------
MODE_NONE = "none" # loudness-normalize only
MODE_DSP = "dsp" # fast STFT noise reduction (noisereduce)
MODE_AI = "ai" # deep-learning separation (Demucs htdemucs_ft)
AI_MODEL_NAME = "htdemucs_ft" # MIT-licensed weights; we keep its "vocals" stem
class Separator:
"""Common interface for the voice-processing modes."""
name = "base"
def available(self) -> tuple[bool, str]:
"""Cheap check (no heavy imports). Returns (ok, human-readable reason)."""
raise NotImplementedError
def separate(self, y: "np.ndarray", sr: int, log=print) -> tuple["np.ndarray", int]:
"""Return (float32 mono audio, sample_rate)."""
raise NotImplementedError
def cancel(self):
"""Best-effort cancel of a running separation (no-op for fast modes)."""
def _ai_deps_present() -> bool:
"""Fast presence check for torch+demucs without importing them (importing
torch takes ~1-2 s and ~400 MB RAM too expensive for a startup UI check)."""
import importlib.util
try:
return (importlib.util.find_spec("torch") is not None
and importlib.util.find_spec("demucs") is not None)
except Exception:
return False
class DSSeparator(Separator):
"""Current noisereduce STFT path — unchanged behavior."""
name = MODE_DSP
def __init__(self, intensity: float = 0.8, stationary: bool = False):
self.intensity = float(intensity)
self.stationary = stationary
def available(self) -> tuple[bool, str]:
if not HAVE_NR:
return False, (f"noisereduce is not installed "
f"({NR_IMPORT_ERROR or 'unknown import error'})")
return True, ""
def separate(self, y, sr, log=print) -> tuple["np.ndarray", int]:
log("Isolating voice — fast noise reduction (CPU)…")
red = VoiceReducer(self.intensity, self.stationary)
return red.reduce(y, sr), sr
class DemucsSeparator(Separator):
"""AI voice separation with Demucs (htdemucs_ft). Lazy imports; GPU-aware.
Only imported/used when the AI mode is actually selected, so the core app
stays a hard-dependency-free ~100 MB tool.
"""
name = MODE_AI
def __init__(self, model_name: str = AI_MODEL_NAME, device: str | None = None):
self.model_name = model_name
self.device = device # None → auto-pick: cuda → mps → cpu
self._api = None # demucs.api.Separator (lazy, worker thread)
self._device_used = None
self._cancel = threading.Event()
# -- cheap checks (safe to call from the GUI thread) -------------------
def available(self) -> tuple[bool, str]:
if not _ai_deps_present():
return False, ("demucs/torch not installed — "
"pip install -r requirements-ai.txt "
"(pulls torch, ~2 GB with CUDA wheels)")
return True, ""
def cancel(self):
self._cancel.set()
# -- heavy work (worker thread only) ------------------------------------
def _pick_device(self) -> str:
if self.device:
return self.device
import torch
if torch.cuda.is_available():
return "cuda"
mps = getattr(torch.backends, "mps", None)
if mps is not None and mps.is_available():
return "mps"
return "cpu"
def _get_api(self, log):
if self._api is None:
# Cache downloads under <appdata>/VoiceGrab/models. Must be set
# before torch/huggingface_hub are first imported in this thread.
root = os.path.join(appdata_dir(), "models")
os.makedirs(root, exist_ok=True)
os.environ.setdefault("TORCH_HOME", os.path.join(root, "torch"))
os.environ.setdefault("HF_HOME", os.path.join(root, "huggingface"))
from demucs.api import Separator as _DemucsApi
self._device_used = self._pick_device()
log(f"Loading AI model '{self.model_name}' "
f"(first use downloads ~90 MB, cached afterwards)…")
self._api = _DemucsApi(self.model_name, device=self._device_used)
log("AI model ready.")
return self._api
@staticmethod
def _model_segment(model) -> float:
try:
for sub in model.models: # BagOfModels
seg = getattr(sub, "segment", None)
if seg:
return float(seg)
except AttributeError:
pass
seg = getattr(model, "segment", None)
return float(seg) if seg else 8.0
@staticmethod
def _estimate_total_chunks(d: dict, api) -> int:
"""Expected 'end' callback count = submodels × shifts × ceil(len/stride)."""
try:
length = int(d["audio_length"])
seg = DemucsSeparator._model_segment(api.model)
n_subs = int(d.get("models", 1) or 1)
stride = int(0.75 * seg * api.samplerate) # overlap=0.25
if stride <= 0:
return 0
return max(1, math.ceil(length / stride)) * n_subs
except Exception:
return 0 # unknown → log raw chunk counts instead
def separate(self, y, sr, log=print) -> tuple["np.ndarray", int]:
import numpy as np
import torch
api = self._get_api(log)
device = self._device_used
if device == "cuda":
log("AI voice separation on GPU (CUDA)…")
elif device == "mps":
log("AI voice separation on GPU (Apple MPS)…")
else:
log("AI voice separation on CPU (no GPU detected — expect a wait)…")
# (C, T) float32 at the original rate; demucs resamples to 44.1 kHz
# and duplicates mono→stereo for us (see demucs.audio.convert_audio).
wav = torch.from_numpy(np.ascontiguousarray(y, dtype=np.float32))
if wav.ndim == 1:
wav = wav[None, :]
state = {"done": 0, "total": None, "last_pct": -1}
def _progress(d):
if self._cancel.is_set():
raise KeyboardInterrupt() # demucs' documented way to abort
if state["total"] is None:
state["total"] = self._estimate_total_chunks(d, api)
if d.get("state") == "end":
state["done"] += 1
total = state["total"]
if total:
pct = min(99, int(state["done"] * 100 / total))
if state["done"] >= total or pct >= state["last_pct"] + 10:
state["last_pct"] = pct
log(f"AI separation: {state['done']}/{total} chunks ({pct}%)…")
else:
log(f"AI separation: chunk {state['done']} done…")
self._cancel.clear()
try:
api.update_parameter(callback=_progress)
_ref, stems = api.separate_tensor(wav, sr)
except KeyboardInterrupt:
raise FfmpegError("Cancelled by user.")
key = ("vocals" if "vocals" in stems
else "speech" if "speech" in stems
else next(iter(stems)))
v = stems[key]
while v.dim() > 1:
v = v[0] # first channel → mono
out = v.detach().cpu().numpy().astype(np.float32)
out = np.where(np.isfinite(out), out, 0.0) # defensive: never emit NaN/Inf
return out, int(api.samplerate)
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# Audio IO helpers (numpy/soundfile) # Audio IO helpers (numpy/soundfile)
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
@ -199,21 +401,24 @@ def write_wav_f32(path: str, y, sr: int):
# GUI # GUI
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
def _import_pyside(): def _import_pyside():
from PySide6 import QtCore, QtGui, QtWidgets # noqa: F401 from PySide6 import QtCore, QtGui, QtWidgets # noqa: F401
return QtCore, QtGui, QtWidgets return QtCore, QtGui, QtWidgets
def run_gui() -> int: def _peak_normalize(y):
QtCore, QtGui, QtWidgets = _import_pyside() """Peak-normalize (99.5th percentile) to ~ -1 dBFS reference."""
def _normalize(y):
import numpy as _np import numpy as _np
p = _np.percentile(_np.abs(y), 99.5) p = _np.percentile(_np.abs(y), 99.5)
if p < 1e-6: if p < 1e-6:
return y return y
return (y / p * 0.891) # ~ -1 dBFS reference peak return (y / p * 0.891) # ~ -1 dBFS reference peak
def run_gui() -> int:
QtCore, QtGui, QtWidgets = _import_pyside()
class Worker(QtCore.QThread): class Worker(QtCore.QThread):
log = QtCore.Signal(str) log = QtCore.Signal(str)
done = QtCore.Signal(object, str) # (success, message) done = QtCore.Signal(object, str) # (success, message)
@ -221,6 +426,11 @@ def run_gui() -> int:
def __init__(self, job): def __init__(self, job):
super().__init__() super().__init__()
self.job = job self.job = job
self._sep: Separator | None = None
def cancel(self):
if self._sep is not None:
self._sep.cancel()
def run(self): def run(self):
import numpy as _np import numpy as _np
@ -230,15 +440,26 @@ def run_gui() -> int:
wav = extract_wav(self.job["input"], self.job["start"], self.job["stop"], wd) wav = extract_wav(self.job["input"], self.job["start"], self.job["stop"], wd)
y, sr = read_wav_f32(wav) y, sr = read_wav_f32(wav)
if self.job["isolate"]: mode = self.job.get("mode", MODE_DSP)
if not HAVE_NR: if mode == MODE_AI:
raise FfmpegError(f"Voice isolation unavailable: {NR_IMPORT_ERROR}") self._sep = DemucsSeparator()
self.log.emit("Isolating voice (noise reduction)\u2026") elif mode == MODE_DSP:
red = VoiceReducer(self.job["intensity"], self.job["stationary"]) self._sep = DSSeparator(self.job["intensity"], self.job["stationary"])
y = red.reduce(y, sr) else:
self._sep = None
if self._sep is not None:
ok, reason = self._sep.available()
if not ok:
raise FfmpegError(f"Voice processing unavailable: {reason}")
y, sr = self._sep.separate(y, sr, self.log.emit)
if mode == MODE_AI:
# Separated vocals sit a few dB below the mix; bring
# peaks up before loudnorm (post-gain step).
self.log.emit("Leveling separated vocals…")
y = _peak_normalize(y)
else: else:
self.log.emit("Normalizing loudness…") self.log.emit("Normalizing loudness…")
y = _normalize(y) y = _peak_normalize(y)
work_wav = os.path.join(wd, "work.wav") work_wav = os.path.join(wd, "work.wav")
write_wav_f32(work_wav, y, sr) write_wav_f32(work_wav, y, sr)
@ -257,10 +478,10 @@ def run_gui() -> int:
self.done.emit(True, self.job["output"]) self.done.emit(True, self.job["output"])
except Exception as exc: except Exception as exc:
self.done.emit(False, str(exc)) self.done.emit(False, str(exc))
return _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) return _main_loop(QtCore, QtGui, QtWidgets, Worker)
def _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) -> int: def _main_loop(QtCore, QtGui, QtWidgets, Worker) -> int:
app = QtWidgets.QApplication(sys.argv) app = QtWidgets.QApplication(sys.argv)
app.setApplicationName(APP_NAME) app.setApplicationName(APP_NAME)
app.setApplicationVersion(APP_VERSION) app.setApplicationVersion(APP_VERSION)
@ -321,11 +542,26 @@ def _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) -> int:
row2.addWidget(btn_listen) row2.addWidget(btn_listen)
lay.addLayout(row2) lay.addLayout(row2)
# Row 3: isolation + output # Row 3: processing mode (one of: fast DSP / AI / none)
self.chk_isolate = QtWidgets.QCheckBox("Isolate voice (reduce background noise)") proc_box = QtWidgets.QGroupBox("Voice processing (choose one)")
self.chk_isolate.setChecked(True) proc_lay = QtWidgets.QVBoxLayout(proc_box)
self.chk_isolate.toggled.connect(self._isolate_toggled) self.rb_dsp = QtWidgets.QRadioButton(
lay.addWidget(self.chk_isolate) "Fast noise reduction (CPU) — fan / hum / room tone; no downloads")
self.rb_ai = QtWidgets.QRadioButton(
"AI voice separation (Demucs) — background music & second speakers; "
"uses GPU if available")
self.rb_none = QtWidgets.QRadioButton(
"No processing — just loudness-normalize")
self.rb_dsp.setChecked(True)
self.rb_ai.setToolTip(
"Deep-learning vocal separation. First use downloads ~90 MB of models. "
"GPU (NVIDIA CUDA / Apple) is strongly recommended; on CPU a 60 s clip "
"takes a few minutes. Long AI runs can be cancelled between chunks.")
for rb in (self.rb_dsp, self.rb_ai, self.rb_none):
proc_lay.addWidget(rb)
lay.addWidget(proc_box)
self.rb_dsp.toggled.connect(self._proc_mode_changed)
self.rb_ai.toggled.connect(self._proc_mode_changed)
iso_row = QtWidgets.QHBoxLayout() iso_row = QtWidgets.QHBoxLayout()
iso_row.addWidget(QtWidgets.QLabel("Noise reduction strength:")) iso_row.addWidget(QtWidgets.QLabel("Noise reduction strength:"))
@ -339,11 +575,12 @@ def _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) -> int:
"Stationary noise (fan / hum — better if constant)") "Stationary noise (fan / hum — better if constant)")
iso_row.addWidget(self.chk_stationary) iso_row.addWidget(self.chk_stationary)
iso_row.addStretch(1) iso_row.addStretch(1)
self.lbl_iso_status = QtWidgets.QLabel("")
self.lbl_iso_status.setStyleSheet("color:#64748b;")
iso_row.addWidget(self.lbl_iso_status)
lay.addLayout(iso_row) lay.addLayout(iso_row)
self.iso_row_widget = iso_row
self.lbl_proc_status = QtWidgets.QLabel("")
self.lbl_proc_status.setStyleSheet("color:#64748b;")
self.lbl_proc_status.setWordWrap(True)
lay.addWidget(self.lbl_proc_status)
out_row = QtWidgets.QHBoxLayout() out_row = QtWidgets.QHBoxLayout()
out_row.addWidget(QtWidgets.QLabel("Output MP3:")) out_row.addWidget(QtWidgets.QLabel("Output MP3:"))
@ -363,6 +600,14 @@ def _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) -> int:
# Bottom bar # Bottom bar
bar = QtWidgets.QHBoxLayout() bar = QtWidgets.QHBoxLayout()
self.btn_cancel = QtWidgets.QPushButton("Cancel")
self.btn_cancel.setMinimumHeight(40)
self.btn_cancel.setEnabled(False)
self.btn_cancel.setToolTip(
"Cancels an in-progress AI separation at the next chunk boundary. "
"Fast DSP runs finish in seconds and are not cancelable.")
self.btn_cancel.clicked.connect(self._cancel)
bar.addWidget(self.btn_cancel)
self.btn_export = QtWidgets.QPushButton("Export MP3") self.btn_export = QtWidgets.QPushButton("Export MP3")
self.btn_export.setMinimumHeight(40) self.btn_export.setMinimumHeight(40)
self.btn_export.setStyleSheet( self.btn_export.setStyleSheet(
@ -375,15 +620,32 @@ def _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) -> int:
bar.addWidget(self.lbl_status, 1) bar.addWidget(self.lbl_status, 1)
lay.addLayout(bar) lay.addLayout(bar)
if not HAVE_NR: # ----- availability / default mode (cheap checks only) -----
self.lbl_iso_status.setText(f"⚠ voice isolation unavailable ({NR_IMPORT_ERROR})") dsp_ok, dsp_reason = DSSeparator().available()
self.chk_isolate.setEnabled(False) ai_ok, ai_reason = DemucsSeparator().available()
self.spn_intensity.setEnabled(dsp_ok)
self.chk_stationary.setEnabled(dsp_ok)
if not dsp_ok:
self.rb_dsp.setEnabled(False)
self.rb_none.setChecked(True)
self.rb_ai.setEnabled(ai_ok)
self._proc_mode_changed(self.rb_dsp.isChecked())
hints = []
if not ai_ok:
hints.append(f"⚠ AI separation unavailable — {ai_reason}")
else:
hints.append(
"AI mode ready — GPU (NVIDIA/Apple) used if present, otherwise CPU "
"(slower); first use downloads ~90 MB of models, cached afterwards.")
if not dsp_ok:
hints.append(f"⚠ fast noise reduction unavailable — {dsp_reason}")
self.lbl_proc_status.setText(" ".join(hints))
# ---------- helpers ---------- # ---------- helpers ----------
def _isolate_toggled(self, on): def _proc_mode_changed(self, *_):
for w in self.iso_row_widget.items(): dsp = self.rb_dsp.isChecked()
if isinstance(w, QtWidgets.QWidget) and w not in (self.lbl_iso_status,): self.spn_intensity.setEnabled(dsp)
w.setEnabled(on) self.chk_stationary.setEnabled(dsp)
@staticmethod @staticmethod
def _parse_time(s: str) -> float: def _parse_time(s: str) -> float:
@ -515,11 +777,18 @@ def _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) -> int:
out = os.path.join(os.path.dirname(os.path.abspath(self.input_path)), out = os.path.join(os.path.dirname(os.path.abspath(self.input_path)),
base + "_voicegrab.mp3") base + "_voicegrab.mp3")
self.btn_export.setEnabled(False) self.btn_export.setEnabled(False)
self.btn_cancel.setEnabled(True)
self.status("Working… (see log in status area)") self.status("Working… (see log in status area)")
if self.rb_ai.isChecked():
mode = MODE_AI
elif self.rb_none.isChecked():
mode = MODE_NONE
else:
mode = MODE_DSP
self.worker = Worker({ self.worker = Worker({
"input": self.input_path, "start": s, "stop": e, "input": self.input_path, "start": s, "stop": e,
"output": out, "output": out,
"isolate": self.chk_isolate.isChecked(), "mode": mode,
"intensity": self.spn_intensity.value() / 100.0, "intensity": self.spn_intensity.value() / 100.0,
"stationary": self.chk_stationary.isChecked(), "stationary": self.chk_stationary.isChecked(),
"lufts": -16, "lufts": -16,
@ -528,8 +797,15 @@ def _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) -> int:
self.worker.done.connect(self._export_done) self.worker.done.connect(self._export_done)
self.worker.start() self.worker.start()
def _cancel(self):
if self.worker is None:
return
self.worker.cancel()
self.status("Cancelling… (AI separation stops at the next chunk boundary)")
def _export_done(self, ok, msg): def _export_done(self, ok, msg):
self.btn_export.setEnabled(True) self.btn_export.setEnabled(True)
self.btn_cancel.setEnabled(False)
self.worker = None self.worker = None
if ok: if ok:
self.status(f"✔ Saved: {msg}") self.status(f"✔ Saved: {msg}")
@ -580,10 +856,23 @@ def _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) -> int:
if __name__ == "__main__": if __name__ == "__main__":
if "--cli" in sys.argv: if "--cli" in sys.argv:
# Simple headless mode: voicegrab --cli <input> <start> <stop> [output.mp3] # Simple headless mode:
# voicegrab --cli <input> <start> <stop> [output.mp3] [--mode none|dsp|ai]
sys.argv = [a for a in sys.argv if a != "--cli"] sys.argv = [a for a in sys.argv if a != "--cli"]
mode = MODE_DSP
if "--mode" in sys.argv:
i = sys.argv.index("--mode")
try:
mode = sys.argv[i + 1]
except IndexError:
print("usage: --mode needs a value: none | dsp | ai")
raise SystemExit(2)
del sys.argv[i:i + 2]
if mode not in (MODE_NONE, MODE_DSP, MODE_AI):
print(f"unknown --mode {mode!r} (use: none | dsp | ai)")
raise SystemExit(2)
if len(sys.argv) < 4: if len(sys.argv) < 4:
print("usage: voicegrab --cli <input> <start> <stop> [output.mp3]") print("usage: voicegrab --cli <input> <start> <stop> [output.mp3] [--mode none|dsp|ai]")
raise SystemExit(2) raise SystemExit(2)
inp, s, e = sys.argv[1], float(sys.argv[2]), float(sys.argv[3]) inp, s, e = sys.argv[1], float(sys.argv[2]), float(sys.argv[3])
out = sys.argv[4] if len(sys.argv) > 4 else os.path.join( out = sys.argv[4] if len(sys.argv) > 4 else os.path.join(
@ -593,9 +882,21 @@ if __name__ == "__main__":
with tempfile.TemporaryDirectory(prefix="voicegrab-") as wd: with tempfile.TemporaryDirectory(prefix="voicegrab-") as wd:
wav = extract_wav(inp, s, e, wd) wav = extract_wav(inp, s, e, wd)
y, sr = read_wav_f32(wav) y, sr = read_wav_f32(wav)
if HAVE_NR: if mode == MODE_AI:
red = VoiceReducer(0.8, stationary=True) sep = DemucsSeparator()
y = red.reduce(y, sr) ok, reason = sep.available()
if not ok:
print(f"AI separation unavailable: {reason}")
raise SystemExit(3)
y, sr = sep.separate(y, sr, print)
y = _peak_normalize(y) # post-gain before loudnorm
elif mode == MODE_DSP:
sep = DSSeparator(intensity=0.8, stationary=True)
ok, reason = sep.available()
if not ok:
print(f"Fast noise reduction unavailable: {reason}")
raise SystemExit(3)
y, sr = sep.separate(y, sr, print)
red_wav = os.path.join(wd, "reduced.wav") red_wav = os.path.join(wd, "reduced.wav")
write_wav_f32(red_wav, y, sr) write_wav_f32(red_wav, y, sr)
subprocess.run([ subprocess.run([