Audio Production

How to Isolate Vocals and Create Studio Karaoke Tracks Online

A browser can strip a centred vocal out of a stereo mix using arithmetic rather than a machine learning model. It works far better on some records than others, and it helps to know which before you start.

How centre channel cancellation works

Stereo is two channels. Anything panned dead centre in the mix, meaning it was placed identically in the left and the right, appears as the same signal in both. Subtract one channel from the other and identical content cancels to zero, leaving only what differed between them. That difference is the sides of the mix: the wide guitars, the stereo keyboards, the reverb tails.

That is the entire technique. Left minus right gives you the instrumental, left plus right gives you the centre. Analogue hardware has been doing it for decades, it costs almost nothing to compute, and it needs no model, no training data and no server. It also depends completely on the vocal genuinely sitting in the centre.

Why it works well on some records and badly on others

It is at its best on older recordings and simply mixed ones: a lead vocal placed straight up the middle, dry or close to it, with the band spread around it. Feed it a sixties soul record or a straightforward acoustic session and the voice can drop out almost completely.

Modern pop is much harder. Producers routinely widen the vocal bus, double the lead and pan the copies apart, or sit the voice in a wide stereo reverb and delay. None of that is identical in both channels, so none of it cancels. You get a ghost of the vocal rather than silence, and the reverb usually survives even when the dry voice mostly does not.

What else disappears along with the vocal

This is the part nobody warns you about. The technique has no idea what a voice is. It removes whatever is centred, and in a conventional mix the centre is a crowded place. The kick drum is centred. The snare is centred. The bass is centred almost without exception, because low frequencies are kept in the middle to keep vinyl cutting and large PA systems happy. Lead instruments and solos often sit there too.

So a plain centre-cancelled track frequently arrives with the bottom end scooped out and the drums weakened. This tool protects the lowest part of that: everything below about 120 Hz, where the kick drum and the bass fundamentals live, is left out of the cancellation. The snare, the upper body of the bass and any centred lead instrument still go with the vocal, and on some records the result is simply unusable, because the problem is in the mix rather than in the tool.

What the two files actually are

Worth being blunt about, because the names oversell them. The instrumental is each original channel with the centre, above that 120 Hz line, subtracted from it. At full intensity a centred vocal cancels. Lower the intensity slider and part of the centre stays in, which leaves some of the vocal but makes the track sound less hollow. Mono playback is still a problem: a phone speaker or a mono PA adds the two channels together, the sides cancel, and at full intensity you hear little more than the bass and kick. If you plan to use it in public, check it in mono first.

The second file is not an isolated vocal. It is the centre of the mix above 120 Hz: the voice, and with it anything else centred, such as the snare or a lead guitar. It is a useful reference for picking out a lyric or following a melody. It is not an acapella and it will not hold up if you treat it as one.

Both files come back as WAV, so they are large, and the separation mode lets you ask for just one of them. The source has to be stereo: a mono file has no sides to work with, and the tool says so rather than pretending it did something.

What it is actually good for

Rough karaoke of older material. Working out a bassline or a vocal melody by ear. Checking what a producer put in the centre of a mix, which is a genuinely instructive thing to listen to. A quick backing track to practise over at home.

What it is not for: anything released, sold or published. The artifacts are audible on a decent pair of speakers, the balance is wrong, and separating a recording gives you no rights to it that you did not already have.

When you need real separation

Source separation models such as Demucs and Spleeter do a genuinely different job. They are trained on large quantities of music and estimate each instrument independently, so they can pull a vocal out of a wide, reverberant modern mix and hand you separate drum and bass stems as well. The results are in a different league, and for anything close to professional work they are the only sensible answer.

The cost is compute and convenience. These models are heavy, so most web services offering them are uploading your audio to a server to run it, which is a completely different privacy proposition from what happens here. Running Demucs on your own machine means a Python environment and some patience. If you need clean stems for real work, that is the road to take. If you want a karaoke version of a record from 1971 in ten seconds without the file leaving your laptop, arithmetic is enough.

Nothing is uploaded

The decode, the subtraction and the WAV writing all happen in the page you have open. Your music never reaches a server, which matters if the track is your own unreleased recording rather than something off a streaming service. Open the Network tab in your browser's developer tools and run a separation: no request carries the audio, because none is made.

Run it

Vocal remover separates the vocals from the backing track in this tab, without uploading anything.

Open Vocal remover