Separate vocals and instrumental tracks from any song with AI (demucs htdemucs_ft). Download a ZIP with the vocal and karaoke stems.
Free AI vocal remover and karaoke maker. Split any song into a clean vocal stem and an instrumental (karaoke) stem in your browser. Built on Meta's open-source demucs model (htdemucs_ft), the same AI used by audio engineers for music production. No account, no watermark — the audio is processed on our servers and the two stems arrive in a single ZIP.
Choose an audio file or drop it here
Drop an MP3/WAV/FLAC/M4A into the box to get ready
Note: This tool uploads the audio to the server for AI processing. Do not upload sensitive files.
Processing can take a while for long songs (CPU-based AI model).
About Vocal Separation
What is vocal separation?
Vocal separation (also called vocal removal or stem separation) is the
process of splitting a finished song into two parts: the vocals — the
singing, rap or speech — and the instrumental (everything else: drums,
bass, guitars, keyboards, synths). It is the technology that makes karaoke
possible without hiring a live band, and it lets producers sample, remix
or cover songs that were never released as separate tracks.
Until the late 2010s this required either access to the original studio
multi-tracks (rare) or phase-cancellation tricks that only worked on
mono recordings. Modern AI models such as demucs take a fully mixed song
as input and produce two clean stems, even on tracks with heavy reverb,
side-chain compression or layered backing vocals.
How to use this tool
Prepare your audio. Any common format works — MP3, WAV, FLAC,
OGG, M4A, AAC, OPUS, WMA, AIFF or M4B. Sample rate is normalised
automatically; bit depth is preserved.
Drop the file into the upload box above, or click to browse. The
selected file name will appear below the box.
Pick an output format. Choose MP3 for a small file you can play
on any device, or WAV for an uncompressed master suitable for
further editing in a DAW.
Click Separate. The browser uploads the file to our servers,
where the AI model (htdemucs_ft) processes it. The progress bar shows
a "loading the AI model" step the very first time you use the tool,
then "separating vocals and instrumental" while the song is being
analysed. A 3-minute song typically takes 60–180 seconds.
Download the ZIP. When the job finishes, click Download ZIP
(vocals + instrumental) to save a single archive containing the two
stems. Both stems share the source file's tempo and pitch — they are
ready to drop straight into a project.
Tips for the best results
Use the highest-quality source you have. Lossy formats like 128
kbps MP3 introduce artefacts that the AI cannot undo. A 320 kbps MP3,
AAC, FLAC or WAV source will always sound noticeably cleaner.
Shorter clips separate faster and better. If you only need a
chorus or a verse, cut the relevant region in your DAW first; the
model analyses the whole file.
Mono is fine. Most pop, rock and hip-hop releases are mixed in
stereo, but the AI handles mono sources just as well.
Heavy mastering is your enemy. Loudness-war tracks that have been
brick-wall-limited strip the spatial cues the model relies on. A song
that has been professionally mixed but not mastered usually separates
much better than the same song after -8 LUFS limiting.
What "instrumental" actually means here
demucs htdemucs_ft is a hybrid transformer that splits a song into
four internal sources — vocals, drums, bass and other (guitars,
keys, synths, anything that is not vocals, drums or bass). The
--two-stems vocals mode this tool uses returns two stems:
Vocals — the model's vocal prediction.
Instrumental — the original song minus the vocal prediction. This
is computed by subtracting the vocals from the input, which is what
makes the instrumental sound "hollow" in some frequencies. That is
normal, and is the same trade-off that every AI vocal remover makes.
If you need finer-grained stems (drums only, bass only, etc.) the full
demucs model exposes them as a separate option; the two-stem mode is
what most karaoke and remix workflows actually need.
Common uses
Karaoke backing tracks — drop the instrumental into your karaoke
app or DJ software.
Acapella extraction for mashups, remixes and bootlegs.
Sampling — pull a clean vocal or instrumental loop out of a
finished record without paying for an official remix clearance.
Practice and teaching — sing along with the instrumental, then
compare your cover to the original acapella.
Audio restoration — isolate vocals for cleanup, noise reduction
or pitch correction without touching the backing track.
Technical notes
The model runs CPU-only on our server. The first request after a fresh
container start spends about 20–30 seconds loading the 340 MB model
weights into memory; subsequent requests start processing immediately.
Audio larger than the server's per-job limit (currently 50 MB) will be
rejected — split long files first if needed. There is no time
Frequently asked questions
Is this vocal remover free?
Yes — no account, no payment, no watermark. The page is funded by unobtrusive ads and the server cost is absorbed by us.
Is my audio uploaded or shared?
The file is uploaded to our server for the duration of the separation job and deleted right after. We do not store, transcode, share or analyse your audio beyond what is needed to produce the stems.
What audio formats are supported?
MP3, WAV, FLAC, OGG, M4A, AAC, OPUS, WMA, AIFF and M4B. Sample rate is normalised automatically; bit depth is preserved.
Can it separate individual instruments like drums or bass?
Not in this tool — the underlying model is configured for the most common use case (vocals versus everything else). The full four-stem mode (vocals, drums, bass, other) is available in the upstream demucs project if you need finer control.
Why does the first request take so long?
The first vocal-separator request after a fresh server start spends about 20–30 seconds loading the 340 MB AI model into memory. Subsequent requests start immediately. The progress bar shows "loading the AI model" during that step so you know it has not stalled.
Should I pick MP3 or WAV output?
MP3 is much smaller and is fine for karaoke or casual listening. WAV is uncompressed and is what you want if you intend to edit the stems further in a DAW, because lossy compression adds artefacts that survive further processing.
How long does the separation take?
A 3-minute song typically takes 60–180 seconds on our CPU server. Very long files (over 10 minutes) may exceed the per-job limit; split the file first if you need longer segments.
What if my song already has vocals removed?
If the input is purely instrumental, the "vocals" stem will be silent or near-silent hiss and the "instrumental" stem will be identical to the input.