Schema der Stem-Trennung: Links eine einzelne Wellenform des fertigen Songs, die in ein Spektrogramm umgewandelt wird. In der Mitte ein neuronales Netz, das vier Schablonen erzeugt. Rechts vier getrennte Ausgabespuren, beschriftet mit Gesang, Schlagzeug, Bass und Restinstrumente.

Stem Separation

Stem separation refers to software that breaks a finished, mixed piece of music back down into its individual components: vocals, drums, bass, and remaining instruments. Modern programs achieve this so cleanly using learning systems that the results end up in karaoke apps, DJ software, and remixes.

A song you stream is a single, finished audio track. Vocals, drums, bass, and guitar are all fused together into one sound within it. Stem separation reverses this step. Software breaks the finished recording back down into individual tracks, so-called stems. Typically there are four: vocals, drums, bass, and everything else. The result is not a true reconstruction of the original recording, but a very good estimate.

Why karaoke suddenly works for any song

In the past, a karaoke version required the original files from the recording studio. These files are held by the record label and are rarely released. Anyone who only had the published version could practically not remove the vocals. There were tricks, but they sounded muffled and destroyed the rest of the music along with it.

Stem separation lifts this restriction. Any song available as an audio file can be broken down in minutes. This applies to more than just karaoke. DJs mix only the drums of one track live with the melody of another. Music students isolate the bass track to play along with it. Film studios separate dialogue from background noise to redub old films.

Legally, this is tricky. The separation itself is permitted, but redistributing the stems generally is not. An isolated vocal track from a copyrighted song remains protected material. Yet this is exactly where new conflicts arise: such vocal tracks sometimes serve as training material for voice impersonations.

How the software sorts out the sonic mush

First, the recording is converted into an image, a so-called spectrogram. Time runs from left to right on it, pitch from bottom to top. Bright spots mean loud tones. A drum hit looks different on it than a sustained vocal note. The former is short and smeared across many pitches, the latter appears as a narrow, horizontal line.

A neural network learns to recognize these patterns. This is a learning program that is fed many examples rather than given rules. During training, it receives songs whose individual tracks are known. So it gets the mix and the correct solution at the same time. From millions of such comparisons, it learns which part of the image belongs to which instrument.

In operation, the network generates a kind of template for each track. It specifies which portions of the spectrogram belong to the vocals and which do not. An audible file is then calculated back from each template. A common misconception is that the result is lossless. Where two instruments share exactly the same pitch, the software has to guess. That’s why you often hear a watery reverberation at such points, which experts call artifacts.

From Spotify to the remix scene

Well-known tools include Demucs, a freely available model from Meta's research, and commercial services like LALAL.AI or Moises. DJ programs such as Serato and Rekordbox have also built the function in directly. There, it computes in real time while the track is playing. On a normal laptop, separating a song usually takes under a minute.

In the news, the term mainly comes up in two contexts. First, in music restoration: the 2023 Beatles single “Now and Then” became possible because software extracted John Lennon’s voice from a noisy cassette recording. Second, in the debate over AI music, because separated vocal tracks provide the raw material for voice clones.

Stem separation should be distinguished from transcription. Transcription converts music into notation, delivering symbols rather than sound. Stem separation, by contrast, delivers actual audio files that can be played. Both methods are often combined when an app is meant to display the guitar part for a song.

Related Products

Latest News

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.