Vocal Split takes a song apart with AI — the voice on one side, the music on the other, or all four parts — then turns it into a karaoke video with every word lighting up as it is sung, an MP3+G for karaoke players, a mix of its parts with your own recordings, or a whole karaoke party on your television. Everything happens on your own computer. Nothing is uploaded, ever.
Vocal Split is one app for the whole journey. Every step keeps its result, so you never do the slow part twice.
Open an MP3, WAV, FLAC, M4A or OGG file. The AI pulls the vocal away from the music — or splits drums, bass and the rest as well.
Paste the lyrics. Vocal Split finds when every line — and, with the aligner, every word — is sung. Drag anything that needs correcting.
An MP4 video for any TV, in 720p, 1080p, 4K or phone shape, or an MP3+G for karaoke players and hosts' software.
Queue your singers, put the words on the television, change the key for each person, and let the next song start by itself.
The Splitter page is where every song starts. It uses Demucs, the open AI model that musicians and researchers use for source separation, running entirely on your computer and sped up with Intel OpenVINO.

Below the player sits an Output panel modelled on classic studio hardware: a left and right VU needle, each with a green lamp that lights while the needle is in the normal range and a red one that lights when it reaches the red, and, between them, a 15-column spectrum analyser with falling peaks. It shows what actually leaves the app — after the equalizer and effects — and it tells you what is playing (“Output · Vocals”).

Every stem arrives as its own card in its own colour — purple for vocals, green for the instrumental, rose for drums, amber for bass, sky blue for the rest and orange for the trimmed original — so you can tell them apart at a glance. The header says how long the separation took.

Every AI separation leaves a little of the music in the vocal and a little of the voice in the music. Vocal Split adds its own cleanup pass after the model: it compares the two stems moment by moment and takes out the bleed and the hiss, while leaving everything the model was sure about alone.
Fifteen bands from 25 Hz to 16 kHz, two thirds of an octave apart, each from −12 to +12 dB. Mud at 250 Hz, boxiness at 400 Hz, presence at 2.5 kHz and air at 10 kHz each get a fader of their own. The equalizer applies to everything the app plays — and, if you like, to everything it saves.

The Effects rack runs after the equalizer, in the order shown by the lit signal path at the top: Equalizer → Bass → Gain → Width → Reverb → Fade. Stages that are doing nothing show “bypass”, so you always know what is active.

A laptop, a phone or a television's own speakers play almost nothing below 150–200 Hz — and that is where nearly all of a bass line lives. Separate the bass on a laptop and it can seem to be silent, even though the meters show it moving. Bass on small speakers fixes what you hear rather than what is recorded:
Move anything you listen to up to six semitones up or down and play it at 80–120% of its speed, with a high-quality time-stretch that keeps voices natural. The Original button puts both back. This changes listening only; the karaoke page has its own key and tempo for the video (see below). A track takes a moment to move the first time, then every key you have tried is kept ready.
The Karaoke page puts the words over the instrumental as a video for your television — every word lighting up as it is sung, with a countdown before each verse, the next line waiting underneath and a progress bar along the bottom. The result is a standard MP4 (H.264) that plays from a USB stick on practically any TV.

[Anna], [Mark], or [Both] for lines sung together. Headings like Anna: and [Anna & Mark] are understood too, while [Chorus] or [x2] are simply ignored. Each singer gets a colour of their own in the video.
The right-hand column of the Karaoke page holds everything about how the video comes out.

Choose a ready-made look, or design your own — the typeface, the colours, the weight, where the lines sit, how big they are, and your logo in a corner.
Every picture below was written by the app itself. Songs, names and lyrics are invented for this page.











The phone format is a different picture, not a smaller one. Vocal Split writes it with its own layout: the words take nearly twice the share of the width, wrap sooner, and sit in the middle of the tall frame — ready for TikTok, Instagram Reels and YouTube Shorts.
MP3+G (CD+G) is the format of karaoke discs and of the software karaoke hosts use every night — software that can change the key live. Vocal Split writes a pair of files with the same name: Song.mp3 and Song.cdg.

The Party page turns your laptop into a karaoke machine. Build a queue of singers and songs, put the words on the TV, and let Vocal Split move from one singer to the next.





| Key | On the party screen |
|---|---|
| Space | Start the next song · pause · resume |
| → or N | Skip to the next singer |
| ← or R | Restart the song |
| ↑ ↓ | Key up or down a semitone |
| Esc | Pause and close the screen |
The Batch page works through a queue of tracks one after another — separating each, and making the karaoke video and MP3+G for every song that has its words beside it. Start it in the evening and find a finished library in the morning.

Vocal Split batch\
Artist - Song\
Artist - Song - Vocals.mp3
Artist - Song - Instrumental.mp3
Artist - Song (karaoke).mp4
Artist - Song (karaoke).mp3
Artist - Song (karaoke).cdg
Every separated song also appears in History, and its timing is kept — so the Karaoke page opens it ready to adjust.
The Mixer page puts every part of a separation on its own line, with a level, a place between the speakers, a mute and a solo. Add a track of your own — a take you sang, a backing track from another source, a guide — and Vocal Split lines it up with the song by listening to it: where it starts, its speed and its key. Everything is heard as you set it, and the result is saved as one file.


Add a track… takes MP3, WAV, FLAC, M4A, AAC, OGG, Opus, WMA and AIFF files, at any sample rate — several at once if you like. Each is compared with the song and with every one of its parts, and placed against the one it agrees with most: a sung take against the vocal, a backing track against the instrumental. It works in the background, and the track shows Lining it up with the song… until it is done.
The History page keeps your recent separations, so you can play and save their stems again without waiting for the AI a second time.



Your settings, presets and the equalizer and effects you left come back exactly as they were the next time you open Vocal Split. The interface is in English.
Try the whole app for free on a minute of any song. The Home Edition is the complete karaoke maker for your house; the Pro Edition adds what you need to make karaoke for a living. Pricing will be announced at launch.
| Feature | Free Edition | Home Edition | Pro Edition |
|---|---|---|---|
| Vocal & instrumental separation | First 60 s · 3 songs a day | Whole songs, unlimited | Whole songs, unlimited |
| Vocals only / instrumental only | First 60 s | ✓ | ✓ |
| All four parts (drums, bass, other) | — | — | ✓ |
| Cleanup pass · three AI models · MP3 / WAV / FLAC | ✓ | ✓ | ✓ |
| 15-band equalizer, effects, VU meters | ✓ | ✓ | ✓ |
| Key & tempo (listening and videos) | — | ✓ | ✓ |
| Karaoke videos — automatic & AI word timing, manual correction | ✓ | ✓ | ✓ |
| Duets in colour · Latin letters under Cyrillic | ✓ | ✓ | ✓ |
| Ready-made looks, backgrounds, pictures, count-in, guide vocal | ✓ | ✓ | ✓ |
| Video sizes | 720p · 1080p | 720p · 1080p · Vertical 9:16 | 720p · 1080p · Vertical · 4K |
| Custom typeface, colours, placement, size · logo | — | — | ✓ |
| MP3+G export | — | — | ✓ |
| Karaoke party on the TV | ✓ | ✓ | ✓ |
| History · cutting a run | History | ✓ | ✓ |
| Mixer — levels, pan, mute & solo for every part, your own tracks lined up automatically | — | — | ✓ |
| Batch processing · “Artist - Song” file names | — | — | ✓ |
| Use | Personal use | Personal use | Commercial — bars, hosts, channels |
| Item | Requirement |
|---|---|
| System | Windows 10 or Windows 11, 64-bit |
| Processor | Any modern 64-bit Intel or AMD processor. Separation runs on the processor: on a modest two-core laptop a 4-minute song takes around 14 minutes; faster processors take proportionally less. The 60-second free preview takes about a minute. |
| Memory | 8 GB recommended |
| Disk | About 800 MB for the app, plus room for your separations (a 4-minute song is about 20 MB of MP3 stems). The optional word aligner is a one-time 1.2 GB download. |
| Video | Any graphics card that encodes H.264 (NVIDIA, Intel, AMD) — or Windows' own encoder; a software encoder is included as a fallback |
| Party | A TV or projector connected as a second screen (recommended), or a single screen with the keyboard |
| Internet | Not needed to separate, make videos or run a party. Only for the optional extra AI models and the word aligner, downloaded once. |
Vocal Split never uploads your audio, your lyrics or anything about what you do. Separations, settings and presets are kept in your own Windows profile, and Settings has a button that deletes all of it. The only time the app reaches the internet is to download an optional model you asked for, or when you click a link to this website.
Yes. Separation with the default model, karaoke videos, MP3+G, the party and everything else work offline. Only the two optional AI models and the word aligner are downloaded — once, when you first choose them.
No. Everything happens on your own computer. There is no account, no cloud processing and no tracking of what you separate.
MP3, WAV, FLAC, M4A and OGG for separation. The Party page also plays karaoke videos (MP4, MKV, MOV, AVI, WEBM, M4V) and MP3+G files, and the Mixer (Pro) also takes AAC, Opus, WMA and AIFF tracks.
Yes, in the Pro Edition's Mixer. Add a take you sang, a backing track or a guide, and it is lined up with the song automatically — its start, its speed and its key — then mixed with the parts you choose and saved as one file. If it does not sound enough like the song to be placed with confidence, it is left at the start for you to move by hand in 10 or 100 ms steps.
Because nearly all of a bass line is below 150–200 Hz, where laptop, phone and TV speakers play almost nothing — the meters show it moving, but the speakers cannot reproduce it. Listen on headphones, or turn on Bass on small speakers on the Effects page (or pick the “Small speakers” preset): it adds the bass's harmonics, which small speakers can play, so you hear the bass line.
Vocal Split uses Demucs, one of the best open separation models available, followed by its own cleanup pass that removes most of the bleed between voice and music. Results depend on the recording — heavily reverberant or very dense mixes are always harder — but for karaoke the instrumental is usually clean enough that nobody notices anything is missing.
You paste them — or import a .txt or .lrc file. Vocal Split deliberately does not download lyrics from the internet: you stay in control of the words and of the rights to use them.
Drag the dividers on the vocal's waveform to move where a line starts, right-click to change the singer, and watch the preview update instantly. Turn on “Stop after the timing” to check everything before the video is written. Your corrections are kept with the song.
Videos are standard MP4 files with H.264 video and AAC sound — the combination practically every TV, media player and USB stick setup accepts. Choose 720p for very old sets, 1080p for most, and 4K for big screens and projectors.
Yes — any language and any alphabet. For Bulgarian, Serbian and Russian, Vocal Split can also add the lyrics in Latin letters under the Cyrillic line.
Write who sings on a line of its own before their lines — for example [Anna], [Mark] or [Both]. Each singer's words light up in their own colour, and lines sung together in a mix of both. You can also right-click any line on the timeline to change who sings it.
MP3+G is the classic karaoke format used by karaoke discs, players and hosts' software (which can often change the key live). If you only play karaoke at home from a USB stick, the MP4 video is all you need. If you run karaoke nights with dedicated software, MP3+G (in the Pro Edition) fits straight into your library.
Commercial use needs the Pro Edition licence. Please note that the licence covers the software; the rights to the music you process are your own responsibility.
The words fill the whole screen and the keyboard runs the party: Space starts and pauses, → skips, ← restarts, ↑ ↓ change the key, and Esc brings you back to the app. With a TV connected as a second screen, the words go to the TV while you manage the queue on the laptop.
It depends on your processor. On a modest two-core laptop a 4-minute song takes around 14 minutes; faster processors take proportionally less. You can mark just the part you need, and a batch can run overnight.
The whole app on the first 60 seconds of any song, up to three songs a day — including karaoke videos and the party — so you can judge the quality for yourself before buying.
Not at the moment — Vocal Split is made for Windows 10 and 11.
The interface is in English. The installer is available in fourteen languages.
Vocal Split is in internal testing, and the installer is waiting for its certified code signature. Then a small group of people who live and breathe karaoke and audio will test it before release. If you host karaoke nights, run a karaoke channel or work with vocals and stems, apply for the beta.
Coming soon · Join the beta programme