<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">pribor</journal-id><journal-title-group><journal-title xml:lang="ru">Известия высших учебных заведений. Приборостроение</journal-title><trans-title-group xml:lang="en"><trans-title>Journal of Instrument Engineering</trans-title></trans-title-group></journal-title-group><issn pub-type="ppub">0021-3454</issn><issn pub-type="epub">2500-0381</issn><publisher><publisher-name>Национальный исследовательский университет ИТМО</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.17586/0021-3454-2026-69-4-295-302</article-id><article-id custom-type="elpub" pub-id-type="custom">pribor-501</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>СИСТЕМНЫЙ АНАЛИЗ, УПРАВЛЕНИЕ И ОБРАБОТКА ИНФОРМАЦИИ</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="en"><subject>SYSTEM ANALYSIS, MANAGEMENT AND INFORMATION PROCESSING</subject></subj-group></article-categories><title-group><article-title>Сценарии классификации многокомпонентных аудиосигналов с использованием нейронных сетей</article-title><trans-title-group xml:lang="en"><trans-title>Scenarios for classifying multicomponent audio signals using neural networks</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Мирошниченко</surname><given-names>Н. И.</given-names></name><name name-style="western" xml:lang="en"><surname>Miroshnichenko</surname><given-names>N. I.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Никита Игоревич Мирошниченко — кафедра прикладной информатики; ассистент</p><p>Санкт-Петербург</p></bio><bio xml:lang="en"><p>Nikita I. Miroshnichenko - Department of Applied Informatics; Assistant</p><p>St. Petersburg </p></bio><email xlink:type="simple">miroshnichenko.nikita.97@yandex.ru</email><xref ref-type="aff" rid="aff-1"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>Санкт-Петербургский государственный университет аэрокосмического приборостроения</institution><country>Россия</country></aff><aff xml:lang="en"><institution>St. Petersburg State University of Aerospace Instrumentation</institution><country>Russian Federation</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2026</year></pub-date><pub-date pub-type="epub"><day>06</day><month>05</month><year>2026</year></pub-date><volume>69</volume><issue>4</issue><fpage>295</fpage><lpage>302</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Национальный исследовательский университет ИТМО, 2026</copyright-statement><copyright-year>2026</copyright-year><copyright-holder xml:lang="ru">Национальный исследовательский университет ИТМО</copyright-holder><copyright-holder xml:lang="en">Национальный исследовательский университет ИТМО</copyright-holder><license xlink:href="https://pribor.ifmo.ru/jour/about/submissions#copyrightNotice" xlink:type="simple"><license-p>https://pribor.ifmo.ru/jour/about/submissions#copyrightNotice</license-p></license></permissions><self-uri xlink:href="https://pribor.ifmo.ru/jour/article/view/501">https://pribor.ifmo.ru/jour/article/view/501</self-uri><abstract><p>Исследуются методы классификации аудиосигналов, полученных из нескольких одновременно активных источников и с частично пересекающимися признаками. В реальных аудиозаписях часто содержатся звуки нескольких источников, что значительно усложняет задачу автоматического распознавания и снижает точность стандартных моделей, обученных на однокомпонентных сигналах. Цель исследования — оценить эффективность различных сценариев классификации многокомпонентных аудиосигналов. В экспериментах использованы архитектуры ResNet18, ResNet34 и ResNet50. Рассмотрены модели, обученные на однокомпонентных аудиосигналах и протестированные на многокомпонентных с применением классических спектральных фильтров и нейросетевого разделителя Demucs, а также модели, обученные непосредственно на многокомпонентных сигналах. Обучение на многокомпонентных сигналах обеспечило наивысшую точность классификации — до 88,5 % на тестовых наборах. Использование фильтрации и нейросетевого разделения повышает точность моделей, обученных на однокомпонентных аудиосигналах, но полностью компенсировать различие между распределениями данных не удается. Модель, обученная на многокомпонентных аудиосигналах, при тестировании на однокомпонентных показала крайне низкую точность, что выявило ограничения ее прямого применения к задачам идентификации отдельных источников. Полученные результаты подчеркивают важность формирования реалистичных тренировочных наборов и необходимость разработки гибридных подходов, объединяющих модели для однокомпонентных и многокомпонентных сигналов, чтобы повысить универсальность и устойчивость классификаторов.</p></abstract><trans-abstract xml:lang="en"><p>Methods for classifying audio signals received from several simultaneously active sources and with partially overlapping features are being investigated. Real audio recordings often contain sounds from multiple sources, which significantly complicates the task of automatic recognition and reduces the accuracy of standard models trained on single-component signals. The purpose of the study is to evaluate the effectiveness of various classification scenarios for multicomponent audio signals. The ResNet18, ResNet34, and ResNet50 architectures are used in the experiments. Models trained on single-component audio signals and tested on multicomponent ones using classical spectral filters and the Demucs neural network separator, as well as models trained directly on multicomponent signals, are considered.</p><p>Multicomponent signal training provides the highest classification accuracy, up to 88.5% on test sets. The use of filtering and neural network separation increases the accuracy of models trained on single-component audio signals, but it is not possible to fully compensate for the difference between data distributions. The model trained on multicomponent audio signals demonstrates extremely low accuracy when tested on single-component ones, which reveal the limitations of its direct application to the tasks of identifying individual sources. The results obtained emphasize the importance of forming realistic training sets and the need to develop hybrid approaches combining models for single-component and multicomponent signals in order to increase the versatility and stability of classifiers.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>классификация аудиосигналов</kwd><kwd>многокомпонентные сигналы</kwd><kwd>акустические события</kwd><kwd>нейронные сети</kwd><kwd>ResNet</kwd><kwd>обработка звука</kwd><kwd>разделение источников</kwd></kwd-group><kwd-group xml:lang="en"><kwd>classification of audio signals</kwd><kwd>multicomponent signals</kwd><kwd>acoustic events</kwd><kwd>neural networks</kwd><kwd>ResNet</kwd><kwd>sound processing</kwd><kwd>source separation</kwd></kwd-group></article-meta></front><back><ref-list><title>References</title><ref id="cit1"><label>1</label><citation-alternatives><mixed-citation xml:lang="ru">Silk S., Biswas S., Solanki S. Audio Signal Separation and Classification: A Review Paper // Intern. J. Innov. Res. Comput. Commun. Eng. 2014. Vol. 2, N 11 [Электронный ресурс]: &lt;https://www.rroij.com/open-access/audio-signalseparation-and-classificationa-review-paper.pdf&gt;.</mixed-citation><mixed-citation xml:lang="en">Silk S., Biswas S., Solanki S. Audio Signal Separation and Classification: A Review Paper // Intern. J. Innov. Res. Comput. Commun. Eng. 2014. Vol. 2, N 11 [Электронный ресурс]: &lt;https://www.rroij.com/open-access/audio-signalseparation-and-classificationa-review-paper.pdf&gt;.</mixed-citation></citation-alternatives></ref><ref id="cit2"><label>2</label><citation-alternatives><mixed-citation xml:lang="ru">Grais E. M., Roma G., Simpson A. J. R., Plumbley M. D. Discriminative Enhancement for Single Channel Audio Source Separation using Deep Neural Networks // arXiv. 2016. arXiv:1609.01678 [Электронный ресурс]: &lt;https://arxiv.org/abs/1609.01678&gt;.</mixed-citation><mixed-citation xml:lang="en">Grais E. M., Roma G., Simpson A. J. R., Plumbley M. D. Discriminative Enhancement for Single Channel Audio Source Separation using Deep Neural Networks // arXiv. 2016. arXiv:1609.01678 [Электронный ресурс]: &lt;https://arxiv.org/abs/1609.01678&gt;.</mixed-citation></citation-alternatives></ref><ref id="cit3"><label>3</label><citation-alternatives><mixed-citation xml:lang="ru">Takahashi N., Mitsufuji Y. Multi-scale Multi-band DenseNets for Audio Source Separation //arXiv. 2017. arXiv:1706.09588 [Электронный ресурс]: &lt;https://arxiv.org/abs/1706.09588&gt;.</mixed-citation><mixed-citation xml:lang="en">Takahashi N., Mitsufuji Y. Multi-scale Multi-band DenseNets for Audio Source Separation //arXiv. 2017. arXiv:1706.09588 [Электронный ресурс]: &lt;https://arxiv.org/abs/1706.09588&gt;.</mixed-citation></citation-alternatives></ref><ref id="cit4"><label>4</label><citation-alternatives><mixed-citation xml:lang="ru">Takahashi N., Goswami N., Mitsufuji Y. MMDenseLSTM: An Efficient Combination of Convolutional and Recurrent Neural Networks for Audio Source Separation // arXiv. 2018. arXiv:1805.02410 [Электронный ресурс]: &lt;https://arxiv.org/abs/1805.02410&gt;.</mixed-citation><mixed-citation xml:lang="en">Takahashi N., Goswami N., Mitsufuji Y. MMDenseLSTM: An Efficient Combination of Convolutional and Recurrent Neural Networks for Audio Source Separation // arXiv. 2018. arXiv:1805.02410 [Электронный ресурс]: &lt;https://arxiv.org/abs/1805.02410&gt;.</mixed-citation></citation-alternatives></ref><ref id="cit5"><label>5</label><citation-alternatives><mixed-citation xml:lang="ru">Grais E. M., Ward D., Plumbley M. D. Raw Multi-Channel Audio Source Separation using Multi-Resolution Convolutional Auto-Encoders // arXiv. 2018. arXiv:1803.00702 [Электронный ресурс]: &lt;https://arxiv.org/abs/1803.00702&gt;.</mixed-citation><mixed-citation xml:lang="en">Grais E. M., Ward D., Plumbley M. D. Raw Multi-Channel Audio Source Separation using Multi-Resolution Convolutional Auto-Encoders // arXiv. 2018. arXiv:1803.00702 [Электронный ресурс]: &lt;https://arxiv.org/abs/1803.00702&gt;.</mixed-citation></citation-alternatives></ref><ref id="cit6"><label>6</label><citation-alternatives><mixed-citation xml:lang="ru">Wang C., Jia M., Zhang X. Deep Encoder/Decoder Dual-Path Neural Network for Speech Separation in Noisy Reverberation Environments // EURASIP J. Audio, Speech, Music Processing. 2023. Vol. 2023. Art. nо. 41. DOI: 10.1186/s13636-023-00307-5.</mixed-citation><mixed-citation xml:lang="en">Wang C., Jia M., Zhang X. Deep Encoder/Decoder Dual-Path Neural Network for Speech Separation in Noisy Reverberation Environments // EURASIP J. Audio, Speech, Music Processing. 2023. Vol. 2023. Art. nо. 41. DOI: 10.1186/s13636-023-00307-5.</mixed-citation></citation-alternatives></ref><ref id="cit7"><label>7</label><citation-alternatives><mixed-citation xml:lang="ru">Ochieng P., Li Y., Smith J. Deep Neural Network Techniques for Monaural Speech Enhancement and Separation: State-of-the-Art Analysis // Artif. Intell. Rev. 2023. Vol. 56. P. 3651–3703. DOI: 10.1007/s10462-023-10612-2.</mixed-citation><mixed-citation xml:lang="en">Ochieng P., Li Y., Smith J. Deep Neural Network Techniques for Monaural Speech Enhancement and Separation: State-of-the-Art Analysis // Artif. Intell. Rev. 2023. Vol. 56. P. 3651–3703. DOI: 10.1007/s10462-023-10612-2.</mixed-citation></citation-alternatives></ref><ref id="cit8"><label>8</label><citation-alternatives><mixed-citation xml:lang="ru">Wang Z., Li Z. Speech Separation Using Advanced Deep Neural Network Methods: A Recent Survey // Algorithms (MDPI). 2025. Vol. 9, N 11. Art. nо. 289. DOI: 10.3390/a9110289.</mixed-citation><mixed-citation xml:lang="en">Wang Z., Li Z. Speech Separation Using Advanced Deep Neural Network Methods: A Recent Survey // Algorithms (MDPI). 2025. Vol. 9, N 11. Art. nо. 289. DOI: 10.3390/a9110289.</mixed-citation></citation-alternatives></ref><ref id="cit9"><label>9</label><citation-alternatives><mixed-citation xml:lang="ru">Chen Y.-S., Lin Z.-J., Bai M. R. A Multichannel Learning-Based Approach for Sound Source Separation in Reverberant Environments // EURASIP J. Audio, Speech, Music Processing. 2021. Vol. 2021. Art. nо. 38 [Электронный ресурс]: &lt;https://asmp-eurasipjournals.springeropen.com/articles/10.1186/s13636-021-00227-2&gt;.</mixed-citation><mixed-citation xml:lang="en">Chen Y.-S., Lin Z.-J., Bai M. R. A Multichannel Learning-Based Approach for Sound Source Separation in Reverberant Environments // EURASIP J. Audio, Speech, Music Processing. 2021. Vol. 2021. Art. nо. 38 [Электронный ресурс]: &lt;https://asmp-eurasipjournals.springeropen.com/articles/10.1186/s13636-021-00227-2&gt;.</mixed-citation></citation-alternatives></ref><ref id="cit10"><label>10</label><citation-alternatives><mixed-citation xml:lang="ru">Teng Zhang &amp; Ji Wu. Learning Long-Term Filter Banks for Audio Source Separation and Audio Scene Classification // EURASIP J. Audio, Speech, Music Processing. 2018. Vol. 2018. Art. nо. 4.</mixed-citation><mixed-citation xml:lang="en">Teng Zhang &amp;  Ji Wu. Learning Long-Term Filter Banks for Audio Source Separation and Audio Scene Classification // EURASIP J. Audio, Speech, Music Processing. 2018. Vol. 2018. Art. nо. 4.</mixed-citation></citation-alternatives></ref><ref id="cit11"><label>11</label><citation-alternatives><mixed-citation xml:lang="ru">Deleforge F.-E., Serizel R., Alameda-Pineda X. ESC-50: Dataset for Environmental Sound Classification. 2015 [Электронный ресурс]: &lt;https://github.com/karoldvl/ESC-50&gt;.</mixed-citation><mixed-citation xml:lang="en">Deleforge F.-E., Serizel R., Alameda-Pineda X. ESC-50: Dataset for Environmental Sound Classification. 2015 [Электронный ресурс]: &lt;https://github.com/karoldvl/ESC-50&gt;.</mixed-citation></citation-alternatives></ref><ref id="cit12"><label>12</label><citation-alternatives><mixed-citation xml:lang="ru">Sharma J., Granmo O.-C., Goodwin M. Environmental Sound Classification using Multiple Feature Channels and Attention based Deep Convolutional Neural Network. arXiv preprint, 2019 [Электронный ресурс]: &lt;https://arxiv.org/abs/1908.11219&gt;.</mixed-citation><mixed-citation xml:lang="en">Sharma J., Granmo O.-C., Goodwin M. Environmental Sound Classification using Multiple Feature Channels and Attention based Deep Convolutional Neural Network. arXiv preprint, 2019 [Электронный ресурс]: &lt;https://arxiv.org/abs/1908.11219&gt;.</mixed-citation></citation-alternatives></ref><ref id="cit13"><label>13</label><citation-alternatives><mixed-citation xml:lang="ru">Binandeh Dehaghani P., Pena D., Aguiar A. P. Investigation of Feature Selection and Pooling Methods for Environmental Sound Classification. ResearchGate, 2025 [Электронный ресурс]: &lt;https://www.researchgate.net/publication/397596087_Investigation_of_Feature_Selection_and_Pooling_Methods_for_Environmental_Sound_Classification&gt;.</mixed-citation><mixed-citation xml:lang="en">Binandeh Dehaghani P., Pena D., Aguiar A. P. Investigation of Feature Selection and Pooling Methods for Environmental Sound Classification. ResearchGate, 2025 [Электронный ресурс]: &lt;https://www.researchgate.net/publication/397596087_Investigation_of_Feature_Selection_and_Pooling_Methods_for_Environmental_Sound_Classification&gt;.</mixed-citation></citation-alternatives></ref><ref id="cit14"><label>14</label><citation-alternatives><mixed-citation xml:lang="ru">Мирошниченко Н. И. Методы глубокого обучения в задаче аудиоклассификации // Прикладной искусственный интеллект: перспективы и риски. II Междунар. науч. конф., 21 окт. 2025 г. Сб. докл. СПбГУАП, 2025. С. 145–150 [Электронный ресурс]: &lt;https://guap.ru/content/aai/sbornik2025.pdf&gt;. (дата обращения: 26.11.2025)</mixed-citation><mixed-citation xml:lang="en">Miroshnichenko N.I.  (Applied Artificial Intelligence: Prospects and Risks), Proc. of the II Intern. Scientific Conf., October 21, 2025, St. Petersburg, 2025, рр. 145–150, https://guap.ru/content/aai/sbornik2025.pdf. (in Russ.)</mixed-citation></citation-alternatives></ref><ref id="cit15"><label>15</label><citation-alternatives><mixed-citation xml:lang="ru">Every N., Szymanski B. Separation of Synchronous Pitched Notes by Spectral Filtering of Harmonics. ResearchGate, 2006 [Электронный ресурс]: &lt;https://www.researchgate.net/publication/3457636_Separation_of_synchronous_pitched_notes_by_spectral_filtering_of_harmonics&gt;.</mixed-citation><mixed-citation xml:lang="en">Every N., Szymanski B. Separation of Synchronous Pitched Notes by Spectral Filtering of Harmonics. ResearchGate, 2006 [Электронный ресурс]: &lt;https://www.researchgate.net/publication/3457636_Separation_of_synchronous_pitched_notes_by_spectral_filtering_of_harmonics&gt;.</mixed-citation></citation-alternatives></ref><ref id="cit16"><label>16</label><citation-alternatives><mixed-citation xml:lang="ru">Dannenberg R. B., Hu N. A Spectral-Filtering Approach to Music Signal Separation. DAFx, 2004 [Электронный ресурс]: &lt;https://www.dafx.de/paper-archive/2004/P_197.PDF&gt;.</mixed-citation><mixed-citation xml:lang="en">Dannenberg R. B., Hu N. A Spectral-Filtering Approach to Music Signal Separation. DAFx, 2004 [Электронный ресурс]: &lt;https://www.dafx.de/paper-archive/2004/P_197.PDF&gt;.</mixed-citation></citation-alternatives></ref><ref id="cit17"><label>17</label><citation-alternatives><mixed-citation xml:lang="ru">Défossez A., Usunier N., Bottou L., Bach F. Music Source Separation in the Waveform Domain. arXiv preprint, 2019 [Электронный ресурс]: &lt;https://arxiv.org/abs/1911.13254&gt;.</mixed-citation><mixed-citation xml:lang="en">Défossez A., Usunier N., Bottou L., Bach F. Music Source Separation in the Waveform Domain. arXiv preprint, 2019 [Электронный ресурс]: &lt;https://arxiv.org/abs/1911.13254&gt;.</mixed-citation></citation-alternatives></ref></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
