In the gaps, delay sounds great: it carries the last note, fills the space and keeps the mood alive. Then the next line starts, and you have to make out the words through the previous one. Turn the effect down and the clarity returns, but that lovely tail disappears.
You do not always have to choose one or the other. First, work out whether the timing and number of repeats are getting in the way, or whether the voice simply has to share the same frequency range with the echo.
Do not make the echo repeat everything
Listen to where the first repeat lands. If it keeps colliding with the start of a new phrase, try a different delay time. A quarter note and a dotted eighth produce different patterns; the part itself determines the right setting, not just the project tempo.
Next, check Feedback, the amount of signal sent around for another pass. Sometimes the first response fits perfectly, while the third and fourth get in the way. In that case, shortening the tail is more useful than turning down every repeat.
Filters help too. Taking out a little low end and softening the top of the echo is a simple way to separate it from the lead voice. But if the delay has become dull and the words are still getting lost, there is little point in cutting more frequencies at random.
Ducking or unmasking: what should you turn down?
Conventional ducking lowers the whole echo while the vocal is playing and brings it back in the gaps. A compressor on the return receives its control signal from the vocal track through a sidechain. This works well when you want the repeats to step forward noticeably between lines. Its limitation is that the rest of the tail gets quieter along with the frequencies that are in the way.
Unmasking is more selective: it clears individual frequency regions for the voice. You can do this with a dynamic EQ or a spectral processor with a suitable external sidechain. The vocal then controls which parts of the echo are attenuated, without necessarily pulling down the entire space.
Some delay plugins have this analysis built in. In GammaDelay by PhotonDSP [64-bit VST3], for example, UNMASK listens to the input signal and dynamically attenuates conflicting regions in the repeats. The analysis uses 32 ERB-like detectors per channel, with dynamic filters handling the attenuation. In practice, this lets more of the tail remain audible while the vocal takes the space it needs.
UNMASK in GammaDelay [64-bit VST3] lets you adjust the strength and selectivity of the response. Start gently: find the point where the words become clearer without the echo losing its character. The reference is the plugin's own input, so the vocal part needs to feed that input for this kind of vocal processing.
Ducking still has its place. It suits a pronounced rise of the echo in the gaps; unmasking suits situations where you want a more continuous sense of space.
A wide echo, a focused voice
There is another approach: add width only to the repeats. The vocal itself stays close and focused, with a more spacious backdrop around it.
You can set this up with separate processing on the return, but some delays already separate these jobs. In GammaDelay [64-bit VST3], Detune and Haas affect only the echo, leaving the dry voice unchanged. After widening, listen in mono: attractive width should not turn the tail into something thin or almost inaudible.
Finally, return to the passage where the words were getting lost and play the whole arrangement. If the next line is easy to follow and the gap still carries the tail you wanted, the setting is doing its job. There is no need to reduce the delay further just to make it sound cleaner in Solo.
