Placing A Turkish Vocal So It Does Not Sound Pasted On
A vocal from another session carries another room, another mic and another clock. Four reasons it sits on top of your beat, and a fix for each.
Separation is a guess, not a cut. Where the watery modulation, the whistles and the ghosts in the silence come from, and how to work around them.
A stem splitter does not find the vocal inside a mix and lift it out. It estimates what the vocal probably was, cell by cell, across a grid of time and frequency, and everything that sounds wrong in the result is the shape of that estimate. This is a different problem from taking a room off a voice. De-reverb works on one source and one decay, while separation has to decide who owns energy two sources were sharing. That is why a split acapella can sound watery and metallic even when the mix it came from sounded fine.
Four families show up and they want different answers. Watery modulation means the model is changing its decision from frame to frame. Whistling means isolated bins are switching on and off. A soft breath in front of a hard consonant is pre-echo. A voice still audible under a section you muted is residue. Guessing which one you have costs you the session.
Most of the damage lives above the fundamental range of the voice, where hats, cymbals and sibilance overlap it and none of that material has clean harmonic structure to attribute. Low-pass the acapella hard for a moment and listen to what is left. If the bottom two thirds hold up, you have a top end problem, and a top end can be replaced, resynthesised or hidden.
Residue lives where the vocal is not. Slice the acapella into phrases, delete everything between them, and the ghost of the instrumental leaves with the gaps. Short fades at the edges cost nothing. This single edit removes more perceived artifact than any spectral repair, because the ear notices a beat leaking under silence far more than it notices roughness inside a word.
Every stage changes what the next one hears. Separation leaves a smeared, partly modulated signal, and the offline de-reverb in RAW deals with that better than it deals with a pitched copy of the same thing. Pitch last, once the source is as clean as it is going to get, so the artifacts do not get stretched into something the ear reads as a deliberate effect.
Separation happens on a grid. The audio is cut into short overlapping windows and each window is split into frequency bins, so the model answers one question thousands of times a second: how much of this cell belongs to the voice. With one source in a cell the answer is easy and the output sounds clean. When a hat, a snare tail and a consonant land in the same cell, the model has to divide energy that was never separable.
Errors in neighbouring cells do not average out, they modulate. A bin credited to the vocal in one frame and to the instrumental in the next turns steady energy into flutter, which is the watery sound. An isolated bin surviving without its harmonic family becomes a short tone with nothing around it, which is the whistle. Transients fail differently: a hit is shorter than the analysis window, so energy belonging after the strike spreads across the window and arrives early as pre-echo.
A split stem is not a recording. It is a rendering of a model's opinion about a recording, and the moment you treat it as a source you start defending its weakest parts. Producers who get usable results from separation do the opposite. They decide what they actually need, which is usually one phrase, one ad lib or one chord, and throw the rest of the estimate away along with most of its artifacts.
From there the work is layering, not repair. Band the stem so it covers only what it is good at, put something underneath it, and let the noise floor of the beat do the masking. Cleaning still helps, but it is a different job from taking a room off a single vocal: here you clean after a guess rather than after a space. The metallic edge that survives responds to the same narrow cuts you would use to tame harshness on any bright source.

RAW pulls the room off a vocal on your own machine, so the pass after a split is fixing space instead of stacking a second estimate on the first.
Guides
You may also like
A vocal from another session carries another room, another mic and another clock. Four reasons it sits on top of your beat, and a fix for each.
Reflections are a distance problem, not a gear problem. What to change at the microphone in an untreated room, and which damage no later pass can undo.