Online Vocal & Accompaniment Extractor — AI-Powered Stem Separation (Spleeter & Demucs)
AI-powered vocal & accompaniment extraction. Two modes: Quick mode (Spleeter, 2 stems: vocals + accompaniment) and Pro mode (Demucs, 4 stems: vocals + drums + bass + other). Fully browser-based, no upload required.
Try it Now →Why Choose Formavide?
- Privacy First: Most tools run locally in your browser. GPU features use encrypted cloud processing with auto-delete.
- No Sign-Up: Open your browser and start. No account needed.
- Free to Use: Basic features are always free.
- Cross-Platform: Chrome, Edge, Safari, and mobile.
How to Use
- Choose a mode: Quick (Spleeter, 2 stems, works on most devices) or Pro (Demucs, 4 stems, requires 8GB+ RAM).
- Upload an audio file — supports MP3, WAV, FLAC, M4A and more.
- Wait for the AI model to download on first use (Quick: ~30MB, Pro: ~173MB). It caches for future use.
- Click "Start Separation". Processing runs entirely in your browser — files never leave your device.
- Download the separated stems (vocals, accompaniment, drums, bass, etc.), each available as MP3.
FAQs
How many separation modes are available?
Two modes: Quick mode (Spleeter, 2 stems: vocals + accompaniment, ~30MB model) and Pro mode (Demucs, 4 stems: vocals + drums + bass + other, ~173MB model).
Which mode should I choose?
Quick mode works on most devices and is faster. Pro mode offers more detailed separation but requires 8GB+ RAM and desktop Chrome.
What audio formats are supported?
MP3, WAV, FLAC, M4A, OGG and common audio formats.
Why does the first use require downloading a model?
The AI model runs locally in your browser. Quick mode model is ~30MB, Pro mode is ~173MB. It caches after first download.
How AI Stem Separation Works
AI separation models don't "listen" to music — they convert audio into a spectrogram: an image with time on one axis, frequency on the other, brightness as energy. The model learns what vocals versus drums look like in that image, labels every frequency bin, then reconstructs vocal energy into a standalone track.
Formavide ships two models: Quick mode (Spleeter — lightweight and fast) and Pro mode (Demucs — Meta's music source separation model, markedly better but heavier). Both run entirely in your browser; your music never uploads.
Quick Mode vs Pro Mode
Workflow tip: run Quick first; re-run the same song in Pro only if needed. Quick is fine for karaoke tracks and casual covers — go straight to Pro for remixing and sampling.
| Quick Mode | Pro Mode | |
|---|---|---|
| Model | Spleeter | Demucs |
| Output stems | 2 (vocals + accompaniment) | 4 (vocals + drums + bass + other) |
| Model size | ~30MB | ~173MB |
| Speed | Fast | Slower, better quality |
| Requirements | Most devices | 8GB+ RAM and desktop Chrome recommended |
Common Use Cases
- Karaoke instrumentals for parties, practice, and cover videos
- Acapellas for remixes, mashups, and vocal samples
- Ear training: isolate bass or drums to study arrangements
- Live DJ toolkits: layer acapellas over instrumentals
- Podcast post: extract voice, then denoise and compress it specifically
Tips for Cleaner Separation
- Feed the best source: WAV/FLAC over MP3, 320kbps over 128kbps
- Avoid multiply-compressed audio (e.g. re-encoded platform rips)
- Heavily reverbed or distorted vocals (metal, dense mixes) are inherently harder for every AI tool
- Run the separated vocal through a denoiser for a second-pass cleanup