Preview

Journal of Instrument Engineering

Advanced search
Open Access Open Access  Restricted Access Subscription Access

Scenarios for classifying multicomponent audio signals using neural networks

https://doi.org/10.17586/0021-3454-2026-69-4-295-302

Abstract

Methods for classifying audio signals received from several simultaneously active sources and with partially overlapping features are being investigated. Real audio recordings often contain sounds from multiple sources, which significantly complicates the task of automatic recognition and reduces the accuracy of standard models trained on single-component signals. The purpose of the study is to evaluate the effectiveness of various classification scenarios for multicomponent audio signals. The ResNet18, ResNet34, and ResNet50 architectures are used in the experiments. Models trained on single-component audio signals and tested on multicomponent ones using classical spectral filters and the Demucs neural network separator, as well as models trained directly on multicomponent signals, are considered.

Multicomponent signal training provides the highest classification accuracy, up to 88.5% on test sets. The use of filtering and neural network separation increases the accuracy of models trained on single-component audio signals, but it is not possible to fully compensate for the difference between data distributions. The model trained on multicomponent audio signals demonstrates extremely low accuracy when tested on single-component ones, which reveal the limitations of its direct application to the tasks of identifying individual sources. The results obtained emphasize the importance of forming realistic training sets and the need to develop hybrid approaches combining models for single-component and multicomponent signals in order to increase the versatility and stability of classifiers.

About the Author

N. I. Miroshnichenko
St. Petersburg State University of Aerospace Instrumentation
Russian Federation

Nikita I. Miroshnichenko - Department of Applied Informatics; Assistant

St. Petersburg 



References

1. Silk S., Biswas S., Solanki S. Audio Signal Separation and Classification: A Review Paper // Intern. J. Innov. Res. Comput. Commun. Eng. 2014. Vol. 2, N 11 [Электронный ресурс]: <https://www.rroij.com/open-access/audio-signalseparation-and-classificationa-review-paper.pdf>.

2. Grais E. M., Roma G., Simpson A. J. R., Plumbley M. D. Discriminative Enhancement for Single Channel Audio Source Separation using Deep Neural Networks // arXiv. 2016. arXiv:1609.01678 [Электронный ресурс]: <https://arxiv.org/abs/1609.01678>.

3. Takahashi N., Mitsufuji Y. Multi-scale Multi-band DenseNets for Audio Source Separation //arXiv. 2017. arXiv:1706.09588 [Электронный ресурс]: <https://arxiv.org/abs/1706.09588>.

4. Takahashi N., Goswami N., Mitsufuji Y. MMDenseLSTM: An Efficient Combination of Convolutional and Recurrent Neural Networks for Audio Source Separation // arXiv. 2018. arXiv:1805.02410 [Электронный ресурс]: <https://arxiv.org/abs/1805.02410>.

5. Grais E. M., Ward D., Plumbley M. D. Raw Multi-Channel Audio Source Separation using Multi-Resolution Convolutional Auto-Encoders // arXiv. 2018. arXiv:1803.00702 [Электронный ресурс]: <https://arxiv.org/abs/1803.00702>.

6. Wang C., Jia M., Zhang X. Deep Encoder/Decoder Dual-Path Neural Network for Speech Separation in Noisy Reverberation Environments // EURASIP J. Audio, Speech, Music Processing. 2023. Vol. 2023. Art. nо. 41. DOI: 10.1186/s13636-023-00307-5.

7. Ochieng P., Li Y., Smith J. Deep Neural Network Techniques for Monaural Speech Enhancement and Separation: State-of-the-Art Analysis // Artif. Intell. Rev. 2023. Vol. 56. P. 3651–3703. DOI: 10.1007/s10462-023-10612-2.

8. Wang Z., Li Z. Speech Separation Using Advanced Deep Neural Network Methods: A Recent Survey // Algorithms (MDPI). 2025. Vol. 9, N 11. Art. nо. 289. DOI: 10.3390/a9110289.

9. Chen Y.-S., Lin Z.-J., Bai M. R. A Multichannel Learning-Based Approach for Sound Source Separation in Reverberant Environments // EURASIP J. Audio, Speech, Music Processing. 2021. Vol. 2021. Art. nо. 38 [Электронный ресурс]: <https://asmp-eurasipjournals.springeropen.com/articles/10.1186/s13636-021-00227-2>.

10. Teng Zhang & Ji Wu. Learning Long-Term Filter Banks for Audio Source Separation and Audio Scene Classification // EURASIP J. Audio, Speech, Music Processing. 2018. Vol. 2018. Art. nо. 4.

11. Deleforge F.-E., Serizel R., Alameda-Pineda X. ESC-50: Dataset for Environmental Sound Classification. 2015 [Электронный ресурс]: <https://github.com/karoldvl/ESC-50>.

12. Sharma J., Granmo O.-C., Goodwin M. Environmental Sound Classification using Multiple Feature Channels and Attention based Deep Convolutional Neural Network. arXiv preprint, 2019 [Электронный ресурс]: <https://arxiv.org/abs/1908.11219>.

13. Binandeh Dehaghani P., Pena D., Aguiar A. P. Investigation of Feature Selection and Pooling Methods for Environmental Sound Classification. ResearchGate, 2025 [Электронный ресурс]: <https://www.researchgate.net/publication/397596087_Investigation_of_Feature_Selection_and_Pooling_Methods_for_Environmental_Sound_Classification>.

14. Miroshnichenko N.I. (Applied Artificial Intelligence: Prospects and Risks), Proc. of the II Intern. Scientific Conf., October 21, 2025, St. Petersburg, 2025, рр. 145–150, https://guap.ru/content/aai/sbornik2025.pdf. (in Russ.)

15. Every N., Szymanski B. Separation of Synchronous Pitched Notes by Spectral Filtering of Harmonics. ResearchGate, 2006 [Электронный ресурс]: <https://www.researchgate.net/publication/3457636_Separation_of_synchronous_pitched_notes_by_spectral_filtering_of_harmonics>.

16. Dannenberg R. B., Hu N. A Spectral-Filtering Approach to Music Signal Separation. DAFx, 2004 [Электронный ресурс]: <https://www.dafx.de/paper-archive/2004/P_197.PDF>.

17. Défossez A., Usunier N., Bottou L., Bach F. Music Source Separation in the Waveform Domain. arXiv preprint, 2019 [Электронный ресурс]: <https://arxiv.org/abs/1911.13254>.


Review

For citations:


Miroshnichenko N.I. Scenarios for classifying multicomponent audio signals using neural networks. Journal of Instrument Engineering. 2026;69(4):295-302. (In Russ.) https://doi.org/10.17586/0021-3454-2026-69-4-295-302

Views: 278

JATS XML

ISSN 0021-3454 (Print)
ISSN 2500-0381 (Online)