A Study on Improving Multi-class Audio Source Separation
Via Decoupled CLAP Query Optimization and an Automated Data Engine
Abstract
Language-queried audio source separation (LASS) enables extracting any sound source using natural language. However, adapting LASS models to application-specific sound classes is challenging due to noisy training data and limited semantic coverage of the CLAP-based control signals. We propose a framework comprised of an automated data engine for training-data curation and a two-stage optimization process for class-specific CLAP control signals. Our objective evaluations across seven sound classes show that data refinement and control signal optimization consistently improve source separation performance. Subjective evaluation with 17 participants further demonstrates perceptual improvements of the model trained with optimized control signals over baseline and similar commercial models.
Audio Source Separation Samples:
| Class | Mixture | Google Pixel 8 | Samsung S25 | Baseline | Ours | Ground Truth |
|---|---|---|---|---|---|---|
| Ambient Noise |
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
| Ambient Noise |
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
| Animal |
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
| Animal |
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
| Crowd Noise |
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
| Crowd Noise |
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
| Music |
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
| Music |
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
| Nature |
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
| Nature |
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
| Speech |
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
| Speech |
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
| Wind Noise |
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
| Wind Noise |
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|
0:00 / 0:00
|