Add AI voice separation (Demucs) as toggleable mode, v1.1.0
Implements the GPU-upgrade.md plan: - Separator backend interface (DSSeparator = existing noisereduce path, DemucsSeparator = htdemucs_ft AI separation) — one mode at a time, never stacked - DemucsSeparator: lazy imports (core app stays dep-free), device auto-pick cuda -> mps -> cpu, model cache under appdata/models, chunk progress reporting, soft cancel between chunks - GUI: 'Isolate voice' checkbox replaced by a 3-way mode selector (fast DSP / AI Demucs / none), DSP strength controls follow the selection, AI radio disabled with guidance when torch+demucs are missing, Cancel button for in-flight AI runs - CLI: --mode none|dsp|ai (default dsp, unchanged behavior) - requirements-ai.txt (torch 2.x + demucs 4.1.x), THIRD_PARTY.md, README updates, optional AI pack flags in installer scripts, version bumped to 1.1.0 - tests/test_separators.py: 10 tests incl. graceful-degradation and full-pipeline regression (5.0 s ± 0.15 s, ~320 kbps) in all modes
This commit is contained in:
parent
62c4659486
commit
c2126f2045
|
|
@ -0,0 +1,6 @@
|
||||||
|
.venv/
|
||||||
|
__pycache__/
|
||||||
|
*.pyc
|
||||||
|
dist/
|
||||||
|
build/
|
||||||
|
third_party/
|
||||||
80
README.md
80
README.md
|
|
@ -6,25 +6,39 @@ Cut a time range out of any video file and export a clean, loudness-normalized
|
||||||
```
|
```
|
||||||
Voice / audio file: [ kage.mp4 ] [Open…]
|
Voice / audio file: [ kage.mp4 ] [Open…]
|
||||||
Start: 0:10.00 Stop: 0:35.00 [Load preview] [▶ Play range]
|
Start: 0:10.00 Stop: 0:35.00 [Load preview] [▶ Play range]
|
||||||
[x] Isolate voice (reduce background noise) strength: 80%
|
Voice processing (choose one):
|
||||||
|
(•) Fast noise reduction (CPU) — fan / hum / room tone; no downloads
|
||||||
|
( ) AI voice separation (Demucs) — music & second speakers; GPU if available
|
||||||
|
( ) No processing — just loudness-normalize
|
||||||
|
strength: 80% [x] stationary noise
|
||||||
Output MP3: [ .../kage_voicegrab.mp3 ] [Browse…]
|
Output MP3: [ .../kage_voicegrab.mp3 ] [Browse…]
|
||||||
──────────────────────────────────────────────────
|
──────────────────────────────────────────────────
|
||||||
[ Export MP3 ]
|
[ Cancel ] [ Export MP3 ]
|
||||||
```
|
```
|
||||||
|
|
||||||
## Features
|
## Features
|
||||||
- Works with any container ffmpeg understands: mp4, mkv, mov, webm, avi, flv, m4a…
|
- Works with any container ffmpeg understands: mp4, mkv, mov, webm, avi, flv, m4a…
|
||||||
- **Isolate voice** checkbox: STFT-based noise reduction (runs on CPU, no
|
- **Voice processing — pick one mode** (toggle, mutually exclusive):
|
||||||
downloads, no GPU). Strength slider + a "stationary noise" mode for
|
- **Fast noise reduction (CPU, default)** — STFT spectral gating via
|
||||||
constant fans/hum.
|
`noisereduce`. Great for fan, hum, room tone. No downloads, no GPU;
|
||||||
|
strength slider + a "stationary noise" mode for constant fans/hum.
|
||||||
|
- **AI voice separation (Demucs)** — deep-learning 4-stem separation
|
||||||
|
(we keep the *vocals* stem). Handles background music and a second
|
||||||
|
speaker, which spectral gating cannot. Optional install; uses an
|
||||||
|
NVIDIA GPU (CUDA) or Apple GPU if present, otherwise CPU (slower).
|
||||||
|
- **No processing** — just loudness-normalize.
|
||||||
- Export is **320 kbps MP3**, loudness-normalized to −16 LUFS (broadcast
|
- Export is **320 kbps MP3**, loudness-normalized to −16 LUFS (broadcast
|
||||||
reference — good levels for voice-training datasets).
|
reference — good levels for voice-training datasets).
|
||||||
- "▶ Play range" lets you audition the exact segment before saving.
|
- "▶ Play range" lets you audition the exact segment before saving.
|
||||||
|
- AI runs report chunk progress in the status bar and can be **cancelled**
|
||||||
|
between chunks.
|
||||||
- Also a headless CLI (see below) for scripting/batch work.
|
- Also a headless CLI (see below) for scripting/batch work.
|
||||||
|
|
||||||
## Requirements
|
## Requirements
|
||||||
- Python 3.10+
|
- Python 3.10+
|
||||||
- ffmpeg on your PATH (bundled automatically in installer builds)
|
- ffmpeg on your PATH (bundled automatically in installer builds)
|
||||||
|
- Optional, only for AI voice separation: `pip install -r requirements-ai.txt`
|
||||||
|
(pulls `torch` + `demucs`, ~1.5–2.5 GB with CUDA wheels)
|
||||||
|
|
||||||
## Run from source
|
## Run from source
|
||||||
```bash
|
```bash
|
||||||
|
|
@ -34,11 +48,44 @@ python3 -m venv .venv
|
||||||
.venv/bin/python voicegrab.py
|
.venv/bin/python voicegrab.py
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## AI voice separation (optional — GPU recommended)
|
||||||
|
```bash
|
||||||
|
.venv/bin/pip install -r requirements-ai.txt
|
||||||
|
```
|
||||||
|
- Pick **AI voice separation (Demucs)** in the GUI. The mode is disabled
|
||||||
|
(with an explanation) until `torch` + `demucs` are installed — the core
|
||||||
|
app never requires them.
|
||||||
|
- **First use downloads the `htdemucs_ft` model (~90 MB)** to the app data
|
||||||
|
dir (`<appdata>/VoiceGrab/models/`) and caches it — later runs are
|
||||||
|
instant. GPU (NVIDIA CUDA / Apple) is used automatically when present,
|
||||||
|
otherwise CPU.
|
||||||
|
- Performance expectations (5 s → 60 s clip):
|
||||||
|
- GPU (≥ 6 GB VRAM): ~1–3 s → ~15–40 s
|
||||||
|
- CPU (8 cores): ~10–25 s → a few minutes
|
||||||
|
- **Cancel** stops an AI run at the next chunk boundary. Fast DSP runs
|
||||||
|
finish in seconds and are not cancellable (by design).
|
||||||
|
- Quality note: aggressive separation can thin sibilance/breath — for voice
|
||||||
|
*training* samples the fast DSP mode is usually enough; reach for AI mode
|
||||||
|
when there's music or a competing speaker.
|
||||||
|
- Licensing: Demucs code is MIT and the `htdemucs_ft` weights are MIT
|
||||||
|
(see `THIRD_PARTY.md`).
|
||||||
|
|
||||||
## CLI mode
|
## CLI mode
|
||||||
```bash
|
```bash
|
||||||
voicegrab.py --cli input.mp4 10 35 output.mp3
|
voicegrab.py --cli input.mp4 10 35 output.mp3 # default: fast DSP
|
||||||
|
voicegrab.py --cli input.mp4 10 35 output.mp3 --mode dsp # fast noise reduction
|
||||||
|
voicegrab.py --cli input.mp4 10 35 output.mp3 --mode ai # Demucs (requirements-ai.txt)
|
||||||
|
voicegrab.py --cli input.mp4 10 35 output.mp3 --mode none # just loudness-normalize
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## Tests
|
||||||
|
```bash
|
||||||
|
.venv/bin/python tests/test_separators.py # or: .venv/bin/pytest tests/
|
||||||
|
```
|
||||||
|
Covers: DSP shape/dtype/NaN invariants, AI graceful degradation without
|
||||||
|
torch, device selection, AI unit separation (44.1 kHz output), and the full
|
||||||
|
CLI pipeline in all three modes (valid 320 kbps MP3, duration 5.0 s ± 0.15 s).
|
||||||
|
|
||||||
## Windows installer (single .exe + installer)
|
## Windows installer (single .exe + installer)
|
||||||
1. `pip install pyinstaller`
|
1. `pip install pyinstaller`
|
||||||
2. `python voicegrab.py.spec` is not needed — run:
|
2. `python voicegrab.py.spec` is not needed — run:
|
||||||
|
|
@ -55,17 +102,26 @@ voicegrab.py --cli input.mp4 10 35 output.mp3
|
||||||
(Copy `dist\VoiceGrab.exe` to `installer\` first — the script assumes
|
(Copy `dist\VoiceGrab.exe` to `installer\` first — the script assumes
|
||||||
`dist\VoiceGrab.exe` relative to the project root.)
|
`dist\VoiceGrab.exe` relative to the project root.)
|
||||||
|
|
||||||
|
To ship the AI mode inside the frozen exe (bigger, ~2 GB), install
|
||||||
|
`requirements-ai.txt` into the venv first and add
|
||||||
|
`--collect-all demucs --collect-all torch` to the PyInstaller command.
|
||||||
|
|
||||||
## Linux
|
## Linux
|
||||||
- **AppImage**: `installer/build_linux.sh` (needs PyInstaller +
|
- **AppImage**: `installer/build_linux.sh` (needs PyInstaller +
|
||||||
`linuxdeploy` or `appimagetool` + your ffmpeg in PATH).
|
`linuxdeploy` or `appimagetool` + your ffmpeg in PATH).
|
||||||
- **Debian/Ubuntu**: just run the PyInstaller binary, or install ffmpeg via
|
- **Debian/Ubuntu**: just run the PyInstaller binary, or install ffmpeg via
|
||||||
apt and run from source. A `.desktop` entry template is in `installer/`.
|
apt and run from source. A `.desktop` entry template is in `installer/`.
|
||||||
- **Flatpak** is also a fine route if you want it in your store.
|
- **Flatpak** is also a fine route if you want it in your store.
|
||||||
|
- AI mode: `WITH_AI=1 ./installer/build_linux.sh` bundles `demucs`/`torch`
|
||||||
|
into the binary (bigger build).
|
||||||
|
|
||||||
## Notes on voice-isolation quality
|
## Notes on voice-isolation quality
|
||||||
The built-in reducer is classic DSP (spectral gating / STFT noise
|
- **Fast mode** is classic DSP (spectral gating / STFT noise estimation) —
|
||||||
estimation) — great for fan, hum, light room tone. For music or heavy
|
great for fan, hum, light room tone; it can't remove music or a second
|
||||||
background speech, a deep-learning model (e.g. Demucs / UVR) gives
|
speaker.
|
||||||
better separation but costs ~1–2 GB of downloads and runs far slower
|
- **AI mode** (Demucs `htdemucs_ft`) is a source separator — the right tool
|
||||||
without a GPU. For voice *training* samples, the DSP route is usually
|
for speech-over-music or competing voices. It's a product feature behind a
|
||||||
more than enough, and it keeps the installer small and startup instant.
|
radio button, never a hard dependency: the core app stays small and starts
|
||||||
|
in < 2 s without `torch`/`demucs` installed.
|
||||||
|
- The AI path *replaces* the DSP path — one or the other, never stacked
|
||||||
|
(double-processing adds artifacts).
|
||||||
|
|
|
||||||
|
|
@ -0,0 +1,27 @@
|
||||||
|
# Third-Party Components
|
||||||
|
|
||||||
|
## Core (always)
|
||||||
|
|
||||||
|
| Component | License | Notes |
|
||||||
|
|---|---|---|
|
||||||
|
| PySide6 (Qt for Python) | LGPL v3 / GPL | GUI toolkit |
|
||||||
|
| numpy, scipy, soundfile | BSD / BSD-3 / 3-clause BSD | audio math & IO |
|
||||||
|
| noisereduce | Unlicense / MIT (see repo) | DSP voice isolation ("fast" mode) |
|
||||||
|
| ffmpeg / ffprobe | LGPL v2.1+ (GPL builds exist) | media I/O; bundled binary must match license |
|
||||||
|
|
||||||
|
## Optional — AI voice separation (only if the user installs `requirements-ai.txt`)
|
||||||
|
|
||||||
|
| Component | License | Notes |
|
||||||
|
|---|---|---|
|
||||||
|
| demucs 4.1.x (Meta/Facebook) | Code: MIT | source separation engine |
|
||||||
|
| PyTorch (torch 2.x) | BSD-3 | deep-learning runtime |
|
||||||
|
| htdemucs_ft weights (~90 MB) | **MIT** (per the model card of `adefossez/HTDemucs-ft` on Hugging Face) | downloaded on first use from Hugging Face, cached in the user data dir — never bundled |
|
||||||
|
| huggingface_hub | Apache-2.0 | model download plumbing |
|
||||||
|
|
||||||
|
### License gate (project policy)
|
||||||
|
|
||||||
|
We only use checkpoints whose license we can point to. Demucs code = MIT,
|
||||||
|
`htdemucs_ft` weights = MIT (model card). They are only ever *downloaded by*
|
||||||
|
the *user* on first use — never bundled by us. RoFormer / UVR community
|
||||||
|
checkpoints have mixed or research-only licensing and are therefore **not**
|
||||||
|
used.
|
||||||
Binary file not shown.
|
|
@ -3,7 +3,7 @@
|
||||||
; Compile: ISCC.exe installer\VoiceGrab.iss (from project root, after build_win.ps1)
|
; Compile: ISCC.exe installer\VoiceGrab.iss (from project root, after build_win.ps1)
|
||||||
|
|
||||||
#define MyAppName "VoiceGrab"
|
#define MyAppName "VoiceGrab"
|
||||||
#define MyAppVersion "1.0.0"
|
#define MyAppVersion "1.1.0"
|
||||||
#define MyAppExeName "VoiceGrab.exe"
|
#define MyAppExeName "VoiceGrab.exe"
|
||||||
|
|
||||||
[Setup]
|
[Setup]
|
||||||
|
|
|
||||||
|
|
@ -1,5 +1,6 @@
|
||||||
#!/usr/bin/env bash
|
#!/usr/bin/env bash
|
||||||
# VoiceGrab — Linux build (produces dist/VoiceGrab-<arch> binary; AppImage if linuxdeploy available)
|
# VoiceGrab — Linux build (produces dist/VoiceGrab-<arch> binary; AppImage if linuxdeploy available)
|
||||||
|
# WITH_AI=1 ./installer/build_linux.sh → also bundle torch+demucs (bigger build)
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
cd "$(dirname "$0")/.."
|
cd "$(dirname "$0")/.."
|
||||||
|
|
||||||
|
|
@ -8,6 +9,13 @@ PY="${PYTHON:-.venv/bin/python}"
|
||||||
command -v ffmpeg >/dev/null || { echo "ffmpeg not found in PATH (needed to resolve the binary)"; exit 1; }
|
command -v ffmpeg >/dev/null || { echo "ffmpeg not found in PATH (needed to resolve the binary)"; exit 1; }
|
||||||
|
|
||||||
"$PY" -m pip install pyinstaller
|
"$PY" -m pip install pyinstaller
|
||||||
|
if [ "${WITH_AI:-0}" = "1" ]; then
|
||||||
|
echo "Building with AI voice separation (torch+demucs)…"
|
||||||
|
"$PY" -m pip install -r requirements-ai.txt
|
||||||
|
AI_FLAGS=(--collect-all demucs --collect-all torch)
|
||||||
|
else
|
||||||
|
AI_FLAGS=()
|
||||||
|
fi
|
||||||
|
|
||||||
FFMPEG_BIN="$(command -v ffmpeg)"
|
FFMPEG_BIN="$(command -v ffmpeg)"
|
||||||
FFPROBE_BIN="$(command -v ffprobe || true)"
|
FFPROBE_BIN="$(command -v ffprobe || true)"
|
||||||
|
|
@ -20,6 +28,7 @@ fi
|
||||||
"$PY" -m PyInstaller --noconfirm --onefile --windowed --name VoiceGrab \
|
"$PY" -m PyInstaller --noconfirm --onefile --windowed --name VoiceGrab \
|
||||||
--add-binary "$FFMPEG_BIN:." \
|
--add-binary "$FFMPEG_BIN:." \
|
||||||
"${EXTRA[@]}" \
|
"${EXTRA[@]}" \
|
||||||
|
"${AI_FLAGS[@]}" \
|
||||||
--add-data "assets/icon.png:assets" \
|
--add-data "assets/icon.png:assets" \
|
||||||
--icon assets/icon.ico \
|
--icon assets/icon.ico \
|
||||||
voicegrab.py
|
voicegrab.py
|
||||||
|
|
|
||||||
|
|
@ -1,5 +1,10 @@
|
||||||
# VoiceGrab — Windows build script
|
# VoiceGrab — Windows build script
|
||||||
# Usage (from project root): powershell -ExecutionPolicy Bypass -File installer\build_win.ps1
|
# Usage (from project root): powershell -ExecutionPolicy Bypass -File installer\build_win.ps1 [-BuildAI]
|
||||||
|
# -BuildAI also bundle the optional AI voice separation (torch+demucs;
|
||||||
|
# exe grows by ~2 GB). Core users don't need it.
|
||||||
|
param(
|
||||||
|
[switch]$BuildAI
|
||||||
|
)
|
||||||
$ErrorActionPreference = "Stop"
|
$ErrorActionPreference = "Stop"
|
||||||
$root = Split-Path -Parent $PSScriptRoot
|
$root = Split-Path -Parent $PSScriptRoot
|
||||||
Set-Location $root
|
Set-Location $root
|
||||||
|
|
@ -10,6 +15,16 @@ if (-not (Test-Path ".venv\Scripts\python.exe")) {
|
||||||
}
|
}
|
||||||
& ".venv\Scripts\python.exe" -m pip install --upgrade pip
|
& ".venv\Scripts\python.exe" -m pip install --upgrade pip
|
||||||
& ".venv\Scripts\python.exe" -m pip install -r requirements.txt pyinstaller
|
& ".venv\Scripts\python.exe" -m pip install -r requirements.txt pyinstaller
|
||||||
|
if ($BuildAI) {
|
||||||
|
Write-Host "Building with AI voice separation (torch+demucs)…"
|
||||||
|
& ".venv\Scripts\python.exe" -m pip install -r requirements-ai.txt
|
||||||
|
}
|
||||||
|
|
||||||
|
# AI-mode collection flags (no-op when not building the AI pack)
|
||||||
|
$AI_FLAGS = @()
|
||||||
|
if ($BuildAI) {
|
||||||
|
$AI_FLAGS = @("--collect-all", "demucs", "--collect-all", "torch")
|
||||||
|
}
|
||||||
|
|
||||||
# 2. Grab a static ffmpeg build (BtbN build — includes libmp3lame)
|
# 2. Grab a static ffmpeg build (BtbN build — includes libmp3lame)
|
||||||
$ffdir = "third_party\ffmpeg\bin"
|
$ffdir = "third_party\ffmpeg\bin"
|
||||||
|
|
@ -33,6 +48,7 @@ if (-not (Test-Path "$ffdir\ffmpeg.exe")) {
|
||||||
--add-binary "$ffdir\ffprobe.exe;." `
|
--add-binary "$ffdir\ffprobe.exe;." `
|
||||||
--add-data "assets\icon.png;assets" `
|
--add-data "assets\icon.png;assets" `
|
||||||
--icon "assets\icon.ico" `
|
--icon "assets\icon.ico" `
|
||||||
|
@AI_FLAGS `
|
||||||
voicegrab.py
|
voicegrab.py
|
||||||
|
|
||||||
Write-Host "Built dist\VoiceGrab.exe"
|
Write-Host "Built dist\VoiceGrab.exe"
|
||||||
|
|
|
||||||
|
|
@ -0,0 +1,19 @@
|
||||||
|
# VoiceGrab — OPTIONAL AI voice separation (heavy mode)
|
||||||
|
#
|
||||||
|
# The core app does NOT need this file; it only enables the
|
||||||
|
# "AI voice separation (Demucs)" mode.
|
||||||
|
#
|
||||||
|
# Install:
|
||||||
|
# pip install -r requirements-ai.txt
|
||||||
|
#
|
||||||
|
# Notes:
|
||||||
|
# - `torch` from PyPI on Windows/Linux ships CUDA wheels (~2 GB with the
|
||||||
|
# nvidia-* deps). CPU-only machines can install the smaller CPU build:
|
||||||
|
# pip install torch --index-url https://download.pytorch.org/whl/cpu
|
||||||
|
# - The Demucs `htdemucs_ft` model (~90 MB) downloads on first use and is
|
||||||
|
# cached under the app data dir (<appdata>/VoiceGrab/models/).
|
||||||
|
# - Licenses: see THIRD_PARTY.md (code MIT; weights usable per Meta's
|
||||||
|
# model license — fine for personal use).
|
||||||
|
|
||||||
|
torch>=2.4,<3
|
||||||
|
demucs>=4.1,<4.2
|
||||||
|
|
@ -0,0 +1,252 @@
|
||||||
|
#!/usr/bin/env python3
|
||||||
|
"""VoiceGrab separator tests (GPU-upgrade.md §6 automation).
|
||||||
|
|
||||||
|
Run with the project venv:
|
||||||
|
.venv/bin/python tests/test_separators.py
|
||||||
|
or with pytest:
|
||||||
|
.venv/bin/pytest tests/
|
||||||
|
|
||||||
|
The AI (Demucs) end-to-end test needs the model cached or a network
|
||||||
|
connection (~90 MB, first run only) and takes ~15 s on CPU.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import math
|
||||||
|
import os
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
import tempfile
|
||||||
|
import wave
|
||||||
|
|
||||||
|
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||||
|
ROOT = os.path.dirname(HERE)
|
||||||
|
sys.path.insert(0, ROOT)
|
||||||
|
|
||||||
|
import numpy as np # noqa: E402
|
||||||
|
|
||||||
|
import voicegrab as vg # noqa: E402
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# fixtures / helpers
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
def _synthetic_voice(n: int, sr: int = 48000) -> np.ndarray:
|
||||||
|
"""Speech-like signal (130 Hz fundamental + harmonics, syllable AM)."""
|
||||||
|
t = np.arange(n) / sr
|
||||||
|
sig = np.zeros_like(t)
|
||||||
|
for h in range(1, 31):
|
||||||
|
sig += (0.6 / h) * np.sin(2 * np.pi * 130.0 * h * t + h * 0.7)
|
||||||
|
sig *= 0.35 + 0.65 * (0.5 + 0.5 * np.sin(2 * np.pi * 4.5 * t))
|
||||||
|
sig /= np.abs(sig).max()
|
||||||
|
rng = np.random.default_rng(7)
|
||||||
|
noise = 0.5 * np.cumsum(rng.normal(0, 1, len(t)))
|
||||||
|
noise /= np.abs(noise).max()
|
||||||
|
return np.clip(0.7 * sig + 0.5 * noise, -1, 1).astype(np.float32)
|
||||||
|
|
||||||
|
|
||||||
|
def _write_test_mp3(path: str, seconds: int = 5) -> str:
|
||||||
|
"""Create a seconds-long mp3 with ffmpeg (voice-like tone + noise)."""
|
||||||
|
cmd = [
|
||||||
|
"ffmpeg", "-hide_banner", "-v", "error", "-y",
|
||||||
|
"-f", "lavfi", "-i", f"sine=frequency=220:duration={seconds}",
|
||||||
|
"-f", "lavfi", "-i", f"anoisesrc=d={seconds}:c=pink:a=0.3",
|
||||||
|
"-filter_complex",
|
||||||
|
"[0][1]amix=inputs=2:weights=1 0.5,atrim=0:" + str(seconds) + ",asetpts=PTS-STARTPTS[out]",
|
||||||
|
"-map", "[out]", "-c:a", "libmp3lame", path,
|
||||||
|
]
|
||||||
|
subprocess.run(cmd, check=True, capture_output=True)
|
||||||
|
return path
|
||||||
|
|
||||||
|
|
||||||
|
def _mp3_duration(path: str) -> float:
|
||||||
|
out = subprocess.run(
|
||||||
|
["ffprobe", "-v", "error", "-show_entries", "format=duration",
|
||||||
|
"-of", "default=nw=1:nk=1", path],
|
||||||
|
capture_output=True, text=True, check=True).stdout.strip()
|
||||||
|
return float(out)
|
||||||
|
|
||||||
|
|
||||||
|
def _mp3_bitrate(path: str) -> int:
|
||||||
|
out = subprocess.run(
|
||||||
|
["ffprobe", "-v", "error", "-show_entries", "format=bit_rate",
|
||||||
|
"-of", "default=nw=1:nk=1", path],
|
||||||
|
capture_output=True, text=True, check=True).stdout.strip()
|
||||||
|
return int(out)
|
||||||
|
|
||||||
|
|
||||||
|
def _ai_available() -> bool:
|
||||||
|
return vg._ai_deps_present()
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# unit tests
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
def test_modes_constants():
|
||||||
|
assert vg.MODE_NONE == "none"
|
||||||
|
assert vg.MODE_DSP == "dsp"
|
||||||
|
assert vg.MODE_AI == "ai"
|
||||||
|
assert vg.AI_MODEL_NAME == "htdemucs_ft"
|
||||||
|
|
||||||
|
|
||||||
|
def test_dsp_available_and_shape():
|
||||||
|
sep = vg.DSSeparator(intensity=0.8, stationary=True)
|
||||||
|
ok, reason = sep.available()
|
||||||
|
assert ok, f"DSP should be available in this venv: {reason}"
|
||||||
|
y = _synthetic_voice(48000)
|
||||||
|
out, sr = sep.separate(y, 48000, lambda m: None)
|
||||||
|
assert sr == 48000, "DSP must not change the sample rate"
|
||||||
|
assert out.dtype == np.float32
|
||||||
|
assert out.ndim == 1 and len(out) == len(y)
|
||||||
|
assert np.isfinite(out).all(), "no NaN/Inf allowed"
|
||||||
|
assert float(np.sqrt(np.mean(out**2))) > 0, "must not be all-silence"
|
||||||
|
|
||||||
|
|
||||||
|
def test_ai_available_report():
|
||||||
|
sep = vg.DemucsSeparator()
|
||||||
|
ok, reason = sep.available()
|
||||||
|
if _ai_available():
|
||||||
|
assert ok
|
||||||
|
else:
|
||||||
|
assert not ok
|
||||||
|
assert "requirements-ai.txt" in reason
|
||||||
|
|
||||||
|
|
||||||
|
def test_ai_graceful_without_torch():
|
||||||
|
"""DemucsSeparator.available() must degrade gracefully when torch is absent."""
|
||||||
|
code = (
|
||||||
|
"import sys, types\n"
|
||||||
|
"class Blocker:\n"
|
||||||
|
" def find_spec(self, name, path=None, target=None):\n"
|
||||||
|
" if name == 'torch' or name.startswith('torch.'):\n"
|
||||||
|
" raise ModuleNotFoundError('blocked: ' + name)\n"
|
||||||
|
" return None\n"
|
||||||
|
"sys.meta_path.insert(0, Blocker())\n"
|
||||||
|
"sys.path.insert(0, %r)\n"
|
||||||
|
"import voicegrab as vg\n"
|
||||||
|
"ok, reason = vg.DemucsSeparator().available()\n"
|
||||||
|
"assert not ok, 'should report unavailable when torch is blocked'\n"
|
||||||
|
"assert 'requirements-ai.txt' in reason\n"
|
||||||
|
"print('graceful-degradation OK')\n"
|
||||||
|
) % ROOT
|
||||||
|
r = subprocess.run([sys.executable, "-c", code],
|
||||||
|
capture_output=True, text=True, timeout=120)
|
||||||
|
assert r.returncode == 0, r.stderr
|
||||||
|
assert "graceful-degradation OK" in r.stdout
|
||||||
|
|
||||||
|
|
||||||
|
def test_ai_pick_device_order():
|
||||||
|
if not _ai_available():
|
||||||
|
print(" (skipped: torch/demucs not installed)")
|
||||||
|
return
|
||||||
|
sep = vg.DemucsSeparator(device="cuda")
|
||||||
|
assert sep._pick_device() == "cuda", "explicit device must be honored"
|
||||||
|
auto = vg.DemucsSeparator()._pick_device()
|
||||||
|
assert auto in ("cpu", "cuda", "mps")
|
||||||
|
|
||||||
|
|
||||||
|
def test_ai_separate_unit():
|
||||||
|
if not _ai_available():
|
||||||
|
print(" (skipped: torch/demucs not installed)")
|
||||||
|
return
|
||||||
|
sep = vg.DemucsSeparator()
|
||||||
|
y = _synthetic_voice(3 * 48000) # 3 s @ 48 kHz
|
||||||
|
out, sr = sep.separate(y, 48000, lambda m: None)
|
||||||
|
assert sr == 44100, "Demucs runs at 44.1 kHz"
|
||||||
|
assert out.dtype == np.float32 and out.ndim == 1
|
||||||
|
assert np.isfinite(out).all()
|
||||||
|
expected = 3 * 44100
|
||||||
|
assert abs(len(out) - expected) <= 44100 * 0.02, f"duration drift: {len(out)} vs {expected}"
|
||||||
|
assert float(np.sqrt(np.mean(out**2))) > 0, "must not be all-silence"
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# regression: full CLI pipeline (GPU-upgrade.md §6.2)
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
def test_cli_dsp_regression():
|
||||||
|
with tempfile.TemporaryDirectory(prefix="voicegrab-test-") as wd:
|
||||||
|
src = os.path.join(wd, "in.mp3")
|
||||||
|
out = os.path.join(wd, "out.mp3")
|
||||||
|
_write_test_mp3(src, 5)
|
||||||
|
r = subprocess.run(
|
||||||
|
[sys.executable, os.path.join(ROOT, "voicegrab.py"),
|
||||||
|
"--cli", src, "0", "5", out, "--mode", "dsp"],
|
||||||
|
capture_output=True, text=True, timeout=300)
|
||||||
|
assert r.returncode == 0, r.stderr
|
||||||
|
assert os.path.exists(out)
|
||||||
|
dur = _mp3_duration(out)
|
||||||
|
assert abs(dur - 5.0) <= 0.15, f"duration {dur} not within 5.0 ± 0.15 s"
|
||||||
|
br = _mp3_bitrate(out)
|
||||||
|
assert 300_000 <= br <= 350_000, f"expected ~320 kbps, got {br}"
|
||||||
|
|
||||||
|
|
||||||
|
def test_cli_ai_regression():
|
||||||
|
if not _ai_available():
|
||||||
|
print(" (skipped: torch/demucs not installed)")
|
||||||
|
return
|
||||||
|
with tempfile.TemporaryDirectory(prefix="voicegrab-test-") as wd:
|
||||||
|
src = os.path.join(wd, "in.mp3")
|
||||||
|
out = os.path.join(wd, "out.mp3")
|
||||||
|
_write_test_mp3(src, 5)
|
||||||
|
r = subprocess.run(
|
||||||
|
[sys.executable, os.path.join(ROOT, "voicegrab.py"),
|
||||||
|
"--cli", src, "0", "5", out, "--mode", "ai"],
|
||||||
|
capture_output=True, text=True, timeout=600)
|
||||||
|
assert r.returncode == 0, r.stderr
|
||||||
|
assert os.path.exists(out)
|
||||||
|
dur = _mp3_duration(out)
|
||||||
|
assert abs(dur - 5.0) <= 0.15, f"duration {dur} not within 5.0 ± 0.15 s"
|
||||||
|
br = _mp3_bitrate(out)
|
||||||
|
assert 300_000 <= br <= 350_000, f"expected ~320 kbps, got {br}"
|
||||||
|
|
||||||
|
|
||||||
|
def test_cli_none_mode():
|
||||||
|
with tempfile.TemporaryDirectory(prefix="voicegrab-test-") as wd:
|
||||||
|
src = os.path.join(wd, "in.mp3")
|
||||||
|
out = os.path.join(wd, "out.mp3")
|
||||||
|
_write_test_mp3(src, 5)
|
||||||
|
r = subprocess.run(
|
||||||
|
[sys.executable, os.path.join(ROOT, "voicegrab.py"),
|
||||||
|
"--cli", src, "0", "5", out, "--mode", "none"],
|
||||||
|
capture_output=True, text=True, timeout=300)
|
||||||
|
assert r.returncode == 0, r.stderr
|
||||||
|
assert abs(_mp3_duration(out) - 5.0) <= 0.15
|
||||||
|
|
||||||
|
|
||||||
|
def test_cli_rejects_bad_mode():
|
||||||
|
with tempfile.TemporaryDirectory(prefix="voicegrab-test-") as wd:
|
||||||
|
src = os.path.join(wd, "in.mp3")
|
||||||
|
_write_test_mp3(src, 5)
|
||||||
|
r = subprocess.run(
|
||||||
|
[sys.executable, os.path.join(ROOT, "voicegrab.py"),
|
||||||
|
"--cli", src, "0", "5", os.path.join(wd, "o.mp3"), "--mode", "bogus"],
|
||||||
|
capture_output=True, text=True, timeout=120)
|
||||||
|
assert r.returncode == 2
|
||||||
|
assert "none | dsp | ai" in r.stdout
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
tests = [v for k, v in sorted(globals().items()) if k.startswith("test_")]
|
||||||
|
failed = 0
|
||||||
|
for t in tests:
|
||||||
|
name = t.__name__
|
||||||
|
try:
|
||||||
|
t()
|
||||||
|
print(f"PASS {name}")
|
||||||
|
except AssertionError as exc:
|
||||||
|
failed += 1
|
||||||
|
print(f"FAIL {name}: {exc}")
|
||||||
|
except Exception as exc: # noqa: BLE001
|
||||||
|
failed += 1
|
||||||
|
print(f"ERROR {name}: {type(exc).__name__}: {exc}")
|
||||||
|
total = len(tests)
|
||||||
|
print(f"\n{total - failed}/{total} passed")
|
||||||
|
return 1 if failed else 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
383
voicegrab.py
383
voicegrab.py
|
|
@ -3,22 +3,27 @@
|
||||||
|
|
||||||
- Pick an input video (mp4, mkv, webm, avi, mov, ...)
|
- Pick an input video (mp4, mkv, webm, avi, mov, ...)
|
||||||
- Set start / stop timestamps
|
- Set start / stop timestamps
|
||||||
- Optional: isolate voice with noise reduction (no GPU or big models needed)
|
- Optional voice processing (one of, mutually exclusive):
|
||||||
|
* fast noise reduction (CPU DSP via noisereduce — no downloads)
|
||||||
|
* AI voice separation (Demucs htdemucs_ft — GPU if available, optional install)
|
||||||
|
* no processing (loudness-normalize only)
|
||||||
- Export MP3 (320 kbps, loudness-normalized to -16 LUFS — good for AI voice training)
|
- Export MP3 (320 kbps, loudness-normalized to -16 LUFS — good for AI voice training)
|
||||||
- Waveform preview with the selected range highlighted
|
- Waveform preview with the selected range highlighted
|
||||||
"""
|
"""
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import math
|
||||||
import os
|
import os
|
||||||
import platform
|
import platform
|
||||||
import shutil
|
import shutil
|
||||||
import subprocess
|
import subprocess
|
||||||
import sys
|
import sys
|
||||||
import tempfile
|
import tempfile
|
||||||
|
import threading
|
||||||
import wave
|
import wave
|
||||||
|
|
||||||
APP_NAME = "VoiceGrab"
|
APP_NAME = "VoiceGrab"
|
||||||
APP_VERSION = "1.0.0"
|
APP_VERSION = "1.1.0"
|
||||||
|
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
# Resource / executable helpers
|
# Resource / executable helpers
|
||||||
|
|
@ -33,7 +38,7 @@ def app_dir() -> str:
|
||||||
|
|
||||||
|
|
||||||
def appdata_dir() -> str:
|
def appdata_dir() -> str:
|
||||||
"""Per-user directory for saved model noise profiles."""
|
"""Per-user directory for cached AI models (and any saved profiles)."""
|
||||||
if platform.system() == "Windows":
|
if platform.system() == "Windows":
|
||||||
base = os.environ.get("APPDATA", os.path.expanduser("~"))
|
base = os.environ.get("APPDATA", os.path.expanduser("~"))
|
||||||
elif platform.system() == "Darwin":
|
elif platform.system() == "Darwin":
|
||||||
|
|
@ -179,6 +184,203 @@ class VoiceReducer:
|
||||||
return nr.reduce_noise(**kwargs)
|
return nr.reduce_noise(**kwargs)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Separation backends — one mode at a time (DSP and AI are alternatives, not
|
||||||
|
# stacked: running noisereduce after Demucs adds double-processing artifacts)
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
MODE_NONE = "none" # loudness-normalize only
|
||||||
|
MODE_DSP = "dsp" # fast STFT noise reduction (noisereduce)
|
||||||
|
MODE_AI = "ai" # deep-learning separation (Demucs htdemucs_ft)
|
||||||
|
|
||||||
|
AI_MODEL_NAME = "htdemucs_ft" # MIT-licensed weights; we keep its "vocals" stem
|
||||||
|
|
||||||
|
|
||||||
|
class Separator:
|
||||||
|
"""Common interface for the voice-processing modes."""
|
||||||
|
|
||||||
|
name = "base"
|
||||||
|
|
||||||
|
def available(self) -> tuple[bool, str]:
|
||||||
|
"""Cheap check (no heavy imports). Returns (ok, human-readable reason)."""
|
||||||
|
raise NotImplementedError
|
||||||
|
|
||||||
|
def separate(self, y: "np.ndarray", sr: int, log=print) -> tuple["np.ndarray", int]:
|
||||||
|
"""Return (float32 mono audio, sample_rate)."""
|
||||||
|
raise NotImplementedError
|
||||||
|
|
||||||
|
def cancel(self):
|
||||||
|
"""Best-effort cancel of a running separation (no-op for fast modes)."""
|
||||||
|
|
||||||
|
|
||||||
|
def _ai_deps_present() -> bool:
|
||||||
|
"""Fast presence check for torch+demucs without importing them (importing
|
||||||
|
torch takes ~1-2 s and ~400 MB RAM — too expensive for a startup UI check)."""
|
||||||
|
import importlib.util
|
||||||
|
try:
|
||||||
|
return (importlib.util.find_spec("torch") is not None
|
||||||
|
and importlib.util.find_spec("demucs") is not None)
|
||||||
|
except Exception:
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
class DSSeparator(Separator):
|
||||||
|
"""Current noisereduce STFT path — unchanged behavior."""
|
||||||
|
|
||||||
|
name = MODE_DSP
|
||||||
|
|
||||||
|
def __init__(self, intensity: float = 0.8, stationary: bool = False):
|
||||||
|
self.intensity = float(intensity)
|
||||||
|
self.stationary = stationary
|
||||||
|
|
||||||
|
def available(self) -> tuple[bool, str]:
|
||||||
|
if not HAVE_NR:
|
||||||
|
return False, (f"noisereduce is not installed "
|
||||||
|
f"({NR_IMPORT_ERROR or 'unknown import error'})")
|
||||||
|
return True, ""
|
||||||
|
|
||||||
|
def separate(self, y, sr, log=print) -> tuple["np.ndarray", int]:
|
||||||
|
log("Isolating voice — fast noise reduction (CPU)…")
|
||||||
|
red = VoiceReducer(self.intensity, self.stationary)
|
||||||
|
return red.reduce(y, sr), sr
|
||||||
|
|
||||||
|
|
||||||
|
class DemucsSeparator(Separator):
|
||||||
|
"""AI voice separation with Demucs (htdemucs_ft). Lazy imports; GPU-aware.
|
||||||
|
|
||||||
|
Only imported/used when the AI mode is actually selected, so the core app
|
||||||
|
stays a hard-dependency-free ~100 MB tool.
|
||||||
|
"""
|
||||||
|
|
||||||
|
name = MODE_AI
|
||||||
|
|
||||||
|
def __init__(self, model_name: str = AI_MODEL_NAME, device: str | None = None):
|
||||||
|
self.model_name = model_name
|
||||||
|
self.device = device # None → auto-pick: cuda → mps → cpu
|
||||||
|
self._api = None # demucs.api.Separator (lazy, worker thread)
|
||||||
|
self._device_used = None
|
||||||
|
self._cancel = threading.Event()
|
||||||
|
|
||||||
|
# -- cheap checks (safe to call from the GUI thread) -------------------
|
||||||
|
def available(self) -> tuple[bool, str]:
|
||||||
|
if not _ai_deps_present():
|
||||||
|
return False, ("demucs/torch not installed — "
|
||||||
|
"pip install -r requirements-ai.txt "
|
||||||
|
"(pulls torch, ~2 GB with CUDA wheels)")
|
||||||
|
return True, ""
|
||||||
|
|
||||||
|
def cancel(self):
|
||||||
|
self._cancel.set()
|
||||||
|
|
||||||
|
# -- heavy work (worker thread only) ------------------------------------
|
||||||
|
def _pick_device(self) -> str:
|
||||||
|
if self.device:
|
||||||
|
return self.device
|
||||||
|
import torch
|
||||||
|
if torch.cuda.is_available():
|
||||||
|
return "cuda"
|
||||||
|
mps = getattr(torch.backends, "mps", None)
|
||||||
|
if mps is not None and mps.is_available():
|
||||||
|
return "mps"
|
||||||
|
return "cpu"
|
||||||
|
|
||||||
|
def _get_api(self, log):
|
||||||
|
if self._api is None:
|
||||||
|
# Cache downloads under <appdata>/VoiceGrab/models. Must be set
|
||||||
|
# before torch/huggingface_hub are first imported in this thread.
|
||||||
|
root = os.path.join(appdata_dir(), "models")
|
||||||
|
os.makedirs(root, exist_ok=True)
|
||||||
|
os.environ.setdefault("TORCH_HOME", os.path.join(root, "torch"))
|
||||||
|
os.environ.setdefault("HF_HOME", os.path.join(root, "huggingface"))
|
||||||
|
from demucs.api import Separator as _DemucsApi
|
||||||
|
self._device_used = self._pick_device()
|
||||||
|
log(f"Loading AI model '{self.model_name}' "
|
||||||
|
f"(first use downloads ~90 MB, cached afterwards)…")
|
||||||
|
self._api = _DemucsApi(self.model_name, device=self._device_used)
|
||||||
|
log("AI model ready.")
|
||||||
|
return self._api
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def _model_segment(model) -> float:
|
||||||
|
try:
|
||||||
|
for sub in model.models: # BagOfModels
|
||||||
|
seg = getattr(sub, "segment", None)
|
||||||
|
if seg:
|
||||||
|
return float(seg)
|
||||||
|
except AttributeError:
|
||||||
|
pass
|
||||||
|
seg = getattr(model, "segment", None)
|
||||||
|
return float(seg) if seg else 8.0
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def _estimate_total_chunks(d: dict, api) -> int:
|
||||||
|
"""Expected 'end' callback count = submodels × shifts × ceil(len/stride)."""
|
||||||
|
try:
|
||||||
|
length = int(d["audio_length"])
|
||||||
|
seg = DemucsSeparator._model_segment(api.model)
|
||||||
|
n_subs = int(d.get("models", 1) or 1)
|
||||||
|
stride = int(0.75 * seg * api.samplerate) # overlap=0.25
|
||||||
|
if stride <= 0:
|
||||||
|
return 0
|
||||||
|
return max(1, math.ceil(length / stride)) * n_subs
|
||||||
|
except Exception:
|
||||||
|
return 0 # unknown → log raw chunk counts instead
|
||||||
|
|
||||||
|
def separate(self, y, sr, log=print) -> tuple["np.ndarray", int]:
|
||||||
|
import numpy as np
|
||||||
|
import torch
|
||||||
|
|
||||||
|
api = self._get_api(log)
|
||||||
|
device = self._device_used
|
||||||
|
if device == "cuda":
|
||||||
|
log("AI voice separation on GPU (CUDA)…")
|
||||||
|
elif device == "mps":
|
||||||
|
log("AI voice separation on GPU (Apple MPS)…")
|
||||||
|
else:
|
||||||
|
log("AI voice separation on CPU (no GPU detected — expect a wait)…")
|
||||||
|
|
||||||
|
# (C, T) float32 at the original rate; demucs resamples to 44.1 kHz
|
||||||
|
# and duplicates mono→stereo for us (see demucs.audio.convert_audio).
|
||||||
|
wav = torch.from_numpy(np.ascontiguousarray(y, dtype=np.float32))
|
||||||
|
if wav.ndim == 1:
|
||||||
|
wav = wav[None, :]
|
||||||
|
|
||||||
|
state = {"done": 0, "total": None, "last_pct": -1}
|
||||||
|
|
||||||
|
def _progress(d):
|
||||||
|
if self._cancel.is_set():
|
||||||
|
raise KeyboardInterrupt() # demucs' documented way to abort
|
||||||
|
if state["total"] is None:
|
||||||
|
state["total"] = self._estimate_total_chunks(d, api)
|
||||||
|
if d.get("state") == "end":
|
||||||
|
state["done"] += 1
|
||||||
|
total = state["total"]
|
||||||
|
if total:
|
||||||
|
pct = min(99, int(state["done"] * 100 / total))
|
||||||
|
if state["done"] >= total or pct >= state["last_pct"] + 10:
|
||||||
|
state["last_pct"] = pct
|
||||||
|
log(f"AI separation: {state['done']}/{total} chunks ({pct}%)…")
|
||||||
|
else:
|
||||||
|
log(f"AI separation: chunk {state['done']} done…")
|
||||||
|
|
||||||
|
self._cancel.clear()
|
||||||
|
try:
|
||||||
|
api.update_parameter(callback=_progress)
|
||||||
|
_ref, stems = api.separate_tensor(wav, sr)
|
||||||
|
except KeyboardInterrupt:
|
||||||
|
raise FfmpegError("Cancelled by user.")
|
||||||
|
|
||||||
|
key = ("vocals" if "vocals" in stems
|
||||||
|
else "speech" if "speech" in stems
|
||||||
|
else next(iter(stems)))
|
||||||
|
v = stems[key]
|
||||||
|
while v.dim() > 1:
|
||||||
|
v = v[0] # first channel → mono
|
||||||
|
out = v.detach().cpu().numpy().astype(np.float32)
|
||||||
|
out = np.where(np.isfinite(out), out, 0.0) # defensive: never emit NaN/Inf
|
||||||
|
return out, int(api.samplerate)
|
||||||
|
|
||||||
|
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
# Audio IO helpers (numpy/soundfile)
|
# Audio IO helpers (numpy/soundfile)
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
|
|
@ -199,21 +401,24 @@ def write_wav_f32(path: str, y, sr: int):
|
||||||
# GUI
|
# GUI
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
def _import_pyside():
|
def _import_pyside():
|
||||||
from PySide6 import QtCore, QtGui, QtWidgets # noqa: F401
|
from PySide6 import QtCore, QtGui, QtWidgets # noqa: F401
|
||||||
return QtCore, QtGui, QtWidgets
|
return QtCore, QtGui, QtWidgets
|
||||||
|
|
||||||
|
|
||||||
|
def _peak_normalize(y):
|
||||||
|
"""Peak-normalize (99.5th percentile) to ~ -1 dBFS reference."""
|
||||||
|
import numpy as _np
|
||||||
|
p = _np.percentile(_np.abs(y), 99.5)
|
||||||
|
if p < 1e-6:
|
||||||
|
return y
|
||||||
|
return (y / p * 0.891) # ~ -1 dBFS reference peak
|
||||||
|
|
||||||
|
|
||||||
def run_gui() -> int:
|
def run_gui() -> int:
|
||||||
QtCore, QtGui, QtWidgets = _import_pyside()
|
QtCore, QtGui, QtWidgets = _import_pyside()
|
||||||
|
|
||||||
def _normalize(y):
|
|
||||||
import numpy as _np
|
|
||||||
p = _np.percentile(_np.abs(y), 99.5)
|
|
||||||
if p < 1e-6:
|
|
||||||
return y
|
|
||||||
return (y / p * 0.891) # ~ -1 dBFS reference peak
|
|
||||||
|
|
||||||
class Worker(QtCore.QThread):
|
class Worker(QtCore.QThread):
|
||||||
log = QtCore.Signal(str)
|
log = QtCore.Signal(str)
|
||||||
done = QtCore.Signal(object, str) # (success, message)
|
done = QtCore.Signal(object, str) # (success, message)
|
||||||
|
|
@ -221,6 +426,11 @@ def run_gui() -> int:
|
||||||
def __init__(self, job):
|
def __init__(self, job):
|
||||||
super().__init__()
|
super().__init__()
|
||||||
self.job = job
|
self.job = job
|
||||||
|
self._sep: Separator | None = None
|
||||||
|
|
||||||
|
def cancel(self):
|
||||||
|
if self._sep is not None:
|
||||||
|
self._sep.cancel()
|
||||||
|
|
||||||
def run(self):
|
def run(self):
|
||||||
import numpy as _np
|
import numpy as _np
|
||||||
|
|
@ -230,15 +440,26 @@ def run_gui() -> int:
|
||||||
wav = extract_wav(self.job["input"], self.job["start"], self.job["stop"], wd)
|
wav = extract_wav(self.job["input"], self.job["start"], self.job["stop"], wd)
|
||||||
y, sr = read_wav_f32(wav)
|
y, sr = read_wav_f32(wav)
|
||||||
|
|
||||||
if self.job["isolate"]:
|
mode = self.job.get("mode", MODE_DSP)
|
||||||
if not HAVE_NR:
|
if mode == MODE_AI:
|
||||||
raise FfmpegError(f"Voice isolation unavailable: {NR_IMPORT_ERROR}")
|
self._sep = DemucsSeparator()
|
||||||
self.log.emit("Isolating voice (noise reduction)\u2026")
|
elif mode == MODE_DSP:
|
||||||
red = VoiceReducer(self.job["intensity"], self.job["stationary"])
|
self._sep = DSSeparator(self.job["intensity"], self.job["stationary"])
|
||||||
y = red.reduce(y, sr)
|
else:
|
||||||
|
self._sep = None
|
||||||
|
if self._sep is not None:
|
||||||
|
ok, reason = self._sep.available()
|
||||||
|
if not ok:
|
||||||
|
raise FfmpegError(f"Voice processing unavailable: {reason}")
|
||||||
|
y, sr = self._sep.separate(y, sr, self.log.emit)
|
||||||
|
if mode == MODE_AI:
|
||||||
|
# Separated vocals sit a few dB below the mix; bring
|
||||||
|
# peaks up before loudnorm (post-gain step).
|
||||||
|
self.log.emit("Leveling separated vocals…")
|
||||||
|
y = _peak_normalize(y)
|
||||||
else:
|
else:
|
||||||
self.log.emit("Normalizing loudness…")
|
self.log.emit("Normalizing loudness…")
|
||||||
y = _normalize(y)
|
y = _peak_normalize(y)
|
||||||
|
|
||||||
work_wav = os.path.join(wd, "work.wav")
|
work_wav = os.path.join(wd, "work.wav")
|
||||||
write_wav_f32(work_wav, y, sr)
|
write_wav_f32(work_wav, y, sr)
|
||||||
|
|
@ -257,10 +478,10 @@ def run_gui() -> int:
|
||||||
self.done.emit(True, self.job["output"])
|
self.done.emit(True, self.job["output"])
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
self.done.emit(False, str(exc))
|
self.done.emit(False, str(exc))
|
||||||
return _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize)
|
return _main_loop(QtCore, QtGui, QtWidgets, Worker)
|
||||||
|
|
||||||
|
|
||||||
def _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) -> int:
|
def _main_loop(QtCore, QtGui, QtWidgets, Worker) -> int:
|
||||||
app = QtWidgets.QApplication(sys.argv)
|
app = QtWidgets.QApplication(sys.argv)
|
||||||
app.setApplicationName(APP_NAME)
|
app.setApplicationName(APP_NAME)
|
||||||
app.setApplicationVersion(APP_VERSION)
|
app.setApplicationVersion(APP_VERSION)
|
||||||
|
|
@ -321,11 +542,26 @@ def _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) -> int:
|
||||||
row2.addWidget(btn_listen)
|
row2.addWidget(btn_listen)
|
||||||
lay.addLayout(row2)
|
lay.addLayout(row2)
|
||||||
|
|
||||||
# Row 3: isolation + output
|
# Row 3: processing mode (one of: fast DSP / AI / none)
|
||||||
self.chk_isolate = QtWidgets.QCheckBox("Isolate voice (reduce background noise)")
|
proc_box = QtWidgets.QGroupBox("Voice processing (choose one)")
|
||||||
self.chk_isolate.setChecked(True)
|
proc_lay = QtWidgets.QVBoxLayout(proc_box)
|
||||||
self.chk_isolate.toggled.connect(self._isolate_toggled)
|
self.rb_dsp = QtWidgets.QRadioButton(
|
||||||
lay.addWidget(self.chk_isolate)
|
"Fast noise reduction (CPU) — fan / hum / room tone; no downloads")
|
||||||
|
self.rb_ai = QtWidgets.QRadioButton(
|
||||||
|
"AI voice separation (Demucs) — background music & second speakers; "
|
||||||
|
"uses GPU if available")
|
||||||
|
self.rb_none = QtWidgets.QRadioButton(
|
||||||
|
"No processing — just loudness-normalize")
|
||||||
|
self.rb_dsp.setChecked(True)
|
||||||
|
self.rb_ai.setToolTip(
|
||||||
|
"Deep-learning vocal separation. First use downloads ~90 MB of models. "
|
||||||
|
"GPU (NVIDIA CUDA / Apple) is strongly recommended; on CPU a 60 s clip "
|
||||||
|
"takes a few minutes. Long AI runs can be cancelled between chunks.")
|
||||||
|
for rb in (self.rb_dsp, self.rb_ai, self.rb_none):
|
||||||
|
proc_lay.addWidget(rb)
|
||||||
|
lay.addWidget(proc_box)
|
||||||
|
self.rb_dsp.toggled.connect(self._proc_mode_changed)
|
||||||
|
self.rb_ai.toggled.connect(self._proc_mode_changed)
|
||||||
|
|
||||||
iso_row = QtWidgets.QHBoxLayout()
|
iso_row = QtWidgets.QHBoxLayout()
|
||||||
iso_row.addWidget(QtWidgets.QLabel("Noise reduction strength:"))
|
iso_row.addWidget(QtWidgets.QLabel("Noise reduction strength:"))
|
||||||
|
|
@ -339,11 +575,12 @@ def _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) -> int:
|
||||||
"Stationary noise (fan / hum — better if constant)")
|
"Stationary noise (fan / hum — better if constant)")
|
||||||
iso_row.addWidget(self.chk_stationary)
|
iso_row.addWidget(self.chk_stationary)
|
||||||
iso_row.addStretch(1)
|
iso_row.addStretch(1)
|
||||||
self.lbl_iso_status = QtWidgets.QLabel("")
|
|
||||||
self.lbl_iso_status.setStyleSheet("color:#64748b;")
|
|
||||||
iso_row.addWidget(self.lbl_iso_status)
|
|
||||||
lay.addLayout(iso_row)
|
lay.addLayout(iso_row)
|
||||||
self.iso_row_widget = iso_row
|
|
||||||
|
self.lbl_proc_status = QtWidgets.QLabel("")
|
||||||
|
self.lbl_proc_status.setStyleSheet("color:#64748b;")
|
||||||
|
self.lbl_proc_status.setWordWrap(True)
|
||||||
|
lay.addWidget(self.lbl_proc_status)
|
||||||
|
|
||||||
out_row = QtWidgets.QHBoxLayout()
|
out_row = QtWidgets.QHBoxLayout()
|
||||||
out_row.addWidget(QtWidgets.QLabel("Output MP3:"))
|
out_row.addWidget(QtWidgets.QLabel("Output MP3:"))
|
||||||
|
|
@ -363,6 +600,14 @@ def _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) -> int:
|
||||||
|
|
||||||
# Bottom bar
|
# Bottom bar
|
||||||
bar = QtWidgets.QHBoxLayout()
|
bar = QtWidgets.QHBoxLayout()
|
||||||
|
self.btn_cancel = QtWidgets.QPushButton("Cancel")
|
||||||
|
self.btn_cancel.setMinimumHeight(40)
|
||||||
|
self.btn_cancel.setEnabled(False)
|
||||||
|
self.btn_cancel.setToolTip(
|
||||||
|
"Cancels an in-progress AI separation at the next chunk boundary. "
|
||||||
|
"Fast DSP runs finish in seconds and are not cancelable.")
|
||||||
|
self.btn_cancel.clicked.connect(self._cancel)
|
||||||
|
bar.addWidget(self.btn_cancel)
|
||||||
self.btn_export = QtWidgets.QPushButton("Export MP3")
|
self.btn_export = QtWidgets.QPushButton("Export MP3")
|
||||||
self.btn_export.setMinimumHeight(40)
|
self.btn_export.setMinimumHeight(40)
|
||||||
self.btn_export.setStyleSheet(
|
self.btn_export.setStyleSheet(
|
||||||
|
|
@ -375,15 +620,32 @@ def _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) -> int:
|
||||||
bar.addWidget(self.lbl_status, 1)
|
bar.addWidget(self.lbl_status, 1)
|
||||||
lay.addLayout(bar)
|
lay.addLayout(bar)
|
||||||
|
|
||||||
if not HAVE_NR:
|
# ----- availability / default mode (cheap checks only) -----
|
||||||
self.lbl_iso_status.setText(f"⚠ voice isolation unavailable ({NR_IMPORT_ERROR})")
|
dsp_ok, dsp_reason = DSSeparator().available()
|
||||||
self.chk_isolate.setEnabled(False)
|
ai_ok, ai_reason = DemucsSeparator().available()
|
||||||
|
self.spn_intensity.setEnabled(dsp_ok)
|
||||||
|
self.chk_stationary.setEnabled(dsp_ok)
|
||||||
|
if not dsp_ok:
|
||||||
|
self.rb_dsp.setEnabled(False)
|
||||||
|
self.rb_none.setChecked(True)
|
||||||
|
self.rb_ai.setEnabled(ai_ok)
|
||||||
|
self._proc_mode_changed(self.rb_dsp.isChecked())
|
||||||
|
hints = []
|
||||||
|
if not ai_ok:
|
||||||
|
hints.append(f"⚠ AI separation unavailable — {ai_reason}")
|
||||||
|
else:
|
||||||
|
hints.append(
|
||||||
|
"AI mode ready — GPU (NVIDIA/Apple) used if present, otherwise CPU "
|
||||||
|
"(slower); first use downloads ~90 MB of models, cached afterwards.")
|
||||||
|
if not dsp_ok:
|
||||||
|
hints.append(f"⚠ fast noise reduction unavailable — {dsp_reason}")
|
||||||
|
self.lbl_proc_status.setText(" ".join(hints))
|
||||||
|
|
||||||
# ---------- helpers ----------
|
# ---------- helpers ----------
|
||||||
def _isolate_toggled(self, on):
|
def _proc_mode_changed(self, *_):
|
||||||
for w in self.iso_row_widget.items():
|
dsp = self.rb_dsp.isChecked()
|
||||||
if isinstance(w, QtWidgets.QWidget) and w not in (self.lbl_iso_status,):
|
self.spn_intensity.setEnabled(dsp)
|
||||||
w.setEnabled(on)
|
self.chk_stationary.setEnabled(dsp)
|
||||||
|
|
||||||
@staticmethod
|
@staticmethod
|
||||||
def _parse_time(s: str) -> float:
|
def _parse_time(s: str) -> float:
|
||||||
|
|
@ -515,11 +777,18 @@ def _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) -> int:
|
||||||
out = os.path.join(os.path.dirname(os.path.abspath(self.input_path)),
|
out = os.path.join(os.path.dirname(os.path.abspath(self.input_path)),
|
||||||
base + "_voicegrab.mp3")
|
base + "_voicegrab.mp3")
|
||||||
self.btn_export.setEnabled(False)
|
self.btn_export.setEnabled(False)
|
||||||
|
self.btn_cancel.setEnabled(True)
|
||||||
self.status("Working… (see log in status area)")
|
self.status("Working… (see log in status area)")
|
||||||
|
if self.rb_ai.isChecked():
|
||||||
|
mode = MODE_AI
|
||||||
|
elif self.rb_none.isChecked():
|
||||||
|
mode = MODE_NONE
|
||||||
|
else:
|
||||||
|
mode = MODE_DSP
|
||||||
self.worker = Worker({
|
self.worker = Worker({
|
||||||
"input": self.input_path, "start": s, "stop": e,
|
"input": self.input_path, "start": s, "stop": e,
|
||||||
"output": out,
|
"output": out,
|
||||||
"isolate": self.chk_isolate.isChecked(),
|
"mode": mode,
|
||||||
"intensity": self.spn_intensity.value() / 100.0,
|
"intensity": self.spn_intensity.value() / 100.0,
|
||||||
"stationary": self.chk_stationary.isChecked(),
|
"stationary": self.chk_stationary.isChecked(),
|
||||||
"lufts": -16,
|
"lufts": -16,
|
||||||
|
|
@ -528,8 +797,15 @@ def _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) -> int:
|
||||||
self.worker.done.connect(self._export_done)
|
self.worker.done.connect(self._export_done)
|
||||||
self.worker.start()
|
self.worker.start()
|
||||||
|
|
||||||
|
def _cancel(self):
|
||||||
|
if self.worker is None:
|
||||||
|
return
|
||||||
|
self.worker.cancel()
|
||||||
|
self.status("Cancelling… (AI separation stops at the next chunk boundary)")
|
||||||
|
|
||||||
def _export_done(self, ok, msg):
|
def _export_done(self, ok, msg):
|
||||||
self.btn_export.setEnabled(True)
|
self.btn_export.setEnabled(True)
|
||||||
|
self.btn_cancel.setEnabled(False)
|
||||||
self.worker = None
|
self.worker = None
|
||||||
if ok:
|
if ok:
|
||||||
self.status(f"✔ Saved: {msg}")
|
self.status(f"✔ Saved: {msg}")
|
||||||
|
|
@ -580,10 +856,23 @@ def _main_loop(QtCore, QtGui, QtWidgets, Worker, _normalize) -> int:
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
if "--cli" in sys.argv:
|
if "--cli" in sys.argv:
|
||||||
# Simple headless mode: voicegrab --cli <input> <start> <stop> [output.mp3]
|
# Simple headless mode:
|
||||||
|
# voicegrab --cli <input> <start> <stop> [output.mp3] [--mode none|dsp|ai]
|
||||||
sys.argv = [a for a in sys.argv if a != "--cli"]
|
sys.argv = [a for a in sys.argv if a != "--cli"]
|
||||||
|
mode = MODE_DSP
|
||||||
|
if "--mode" in sys.argv:
|
||||||
|
i = sys.argv.index("--mode")
|
||||||
|
try:
|
||||||
|
mode = sys.argv[i + 1]
|
||||||
|
except IndexError:
|
||||||
|
print("usage: --mode needs a value: none | dsp | ai")
|
||||||
|
raise SystemExit(2)
|
||||||
|
del sys.argv[i:i + 2]
|
||||||
|
if mode not in (MODE_NONE, MODE_DSP, MODE_AI):
|
||||||
|
print(f"unknown --mode {mode!r} (use: none | dsp | ai)")
|
||||||
|
raise SystemExit(2)
|
||||||
if len(sys.argv) < 4:
|
if len(sys.argv) < 4:
|
||||||
print("usage: voicegrab --cli <input> <start> <stop> [output.mp3]")
|
print("usage: voicegrab --cli <input> <start> <stop> [output.mp3] [--mode none|dsp|ai]")
|
||||||
raise SystemExit(2)
|
raise SystemExit(2)
|
||||||
inp, s, e = sys.argv[1], float(sys.argv[2]), float(sys.argv[3])
|
inp, s, e = sys.argv[1], float(sys.argv[2]), float(sys.argv[3])
|
||||||
out = sys.argv[4] if len(sys.argv) > 4 else os.path.join(
|
out = sys.argv[4] if len(sys.argv) > 4 else os.path.join(
|
||||||
|
|
@ -593,9 +882,21 @@ if __name__ == "__main__":
|
||||||
with tempfile.TemporaryDirectory(prefix="voicegrab-") as wd:
|
with tempfile.TemporaryDirectory(prefix="voicegrab-") as wd:
|
||||||
wav = extract_wav(inp, s, e, wd)
|
wav = extract_wav(inp, s, e, wd)
|
||||||
y, sr = read_wav_f32(wav)
|
y, sr = read_wav_f32(wav)
|
||||||
if HAVE_NR:
|
if mode == MODE_AI:
|
||||||
red = VoiceReducer(0.8, stationary=True)
|
sep = DemucsSeparator()
|
||||||
y = red.reduce(y, sr)
|
ok, reason = sep.available()
|
||||||
|
if not ok:
|
||||||
|
print(f"AI separation unavailable: {reason}")
|
||||||
|
raise SystemExit(3)
|
||||||
|
y, sr = sep.separate(y, sr, print)
|
||||||
|
y = _peak_normalize(y) # post-gain before loudnorm
|
||||||
|
elif mode == MODE_DSP:
|
||||||
|
sep = DSSeparator(intensity=0.8, stationary=True)
|
||||||
|
ok, reason = sep.available()
|
||||||
|
if not ok:
|
||||||
|
print(f"Fast noise reduction unavailable: {reason}")
|
||||||
|
raise SystemExit(3)
|
||||||
|
y, sr = sep.separate(y, sr, print)
|
||||||
red_wav = os.path.join(wd, "reduced.wav")
|
red_wav = os.path.join(wd, "reduced.wav")
|
||||||
write_wav_f32(red_wav, y, sr)
|
write_wav_f32(red_wav, y, sr)
|
||||||
subprocess.run([
|
subprocess.run([
|
||||||
|
|
|
||||||
Loading…
Reference in New Issue