Skip to main content
Free tool · Nothing uploaded

Vocal remover that runs in your browser

Most vocal removers upload your file, put it in a queue and email you a link. This one downloads the separation model to your browser instead and runs it there, so the audio never leaves the tab. You get an instrumental and an acapella, both as WAV.

Runs in your browser · No account · Up to 10 minutes per track

Hear the split

Stem splitter, instrumental maker, acapella extractor, karaoke maker

The same 15 seconds of one song: the mix that went in, and the two stems this page made from it.

  • Original mix

    Heartbreak Is a Highway · 15 s

  • Vocals only (acapella)

    Heartbreak Is a Highway · 15 s

  • Instrumental (karaoke)

    Heartbreak Is a Highway · 15 s

The separator

One file in, two stems out

The first run downloads about 64 MB of model, which your browser then caches. Separation is real work — expect it to take a few minutes on a normal laptop, and longer than the song itself unless your browser supports WebGPU.

What is different here

The file stays on your machine

That is not a privacy slogan bolted onto the same architecture as everyone else — it is a different architecture, and it has consequences worth knowing before you choose a tool.

Nothing is uploaded

There is no upload endpoint in this page at all. Your unreleased mix, your client’s stems and your voice memo are read by the tab and stay there. Closing it deletes everything.

The stems sum back to the original

The model estimates the instrumental and the vocal is what is left over, so adding the two back together reconstructs your file exactly. That is what makes the instrumental usable as a karaoke bed rather than an approximation of one.

Your machine does the work

Which is the trade: no queue and no cost, but no server farm either. WebGPU browsers finish in well under the length of the song; without it you are on single-threaded WebAssembly and it takes noticeably longer.

How it works

Spectrograms, not magic

Worth knowing what the tool is doing, partly because it explains the wait and partly because it tells you when the result will be good.

  1. 01

    Your file is decoded and resampled

    To 44.1 kHz stereo, because that is what the model was trained on. A mono file is played through both sides; anything the browser can decode is accepted.

  2. 02

    The audio becomes a picture of itself

    Six-second chunks are turned into spectrograms — a 6144-point transform every 1024 samples, which is what the model expects to be handed.

  3. 03

    The model estimates the instrumental

    Chunk by chunk, with the edges cross-faded so you do not hear a seam every six seconds. This is the part that takes the time.

  4. 04

    The vocal is what is left

    Subtracting the estimated instrumental from your original gives the vocal, and guarantees the two halves add back up to what you started with.

Two honest limits. Separation is imperfect on dense mixes — heavy reverb, doubled vocals and loud cymbals are where you will hear artefacts, in both stems. And a long track holds the mixture and both stems in memory at once, which is why the ceiling is 10 minutes rather than an album side.

Credit

The model is not ours

Separation uses UVR-MDX-NET-Inst_HQ_1, trained and released by the Ultimate Vocal Remover project under the MIT licence, which asks that developers who use its models credit it. The architecture is MDX-Net, from KUIELab. We built the browser side; they built the part that does the actual work.

Questions

The ones worth answering

Is it free, and is there a limit?

Free, with no account and no per-track cap: your computer does the work, so there is nothing to meter. The only limits are 10 minutes per track and your machine’s memory.

Is my audio uploaded anywhere?

No. The page has no upload path: your browser downloads the model and runs it locally, so the audio never leaves the tab. The only sizeable request is the model itself.

What do I get?

Two stems, the instrumental and the vocal, as 44100 Hz stereo WAV. They add back up to your original exactly, because the vocal is what remains once the instrumental is estimated.

Can I make a karaoke track or an acapella?

Both, from one run: the instrumental is your karaoke or backing track, the vocal is your acapella. Dense mixes leave some artefacts, so listen to both before you use one.

Which files, and how long?

Anything your browser can decode: MP3, WAV, M4A/AAC, FLAC or OGG, up to 10 minutes. The first run downloads about 64 MB of model once; a split takes a few minutes on a normal laptop.

Can I use the stems commercially?

That depends on the song, not the tool: separating a recording gives you no rights to it. Your own recordings are yours to use. For music you can license, generate your own.