<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">pribor</journal-id><journal-title-group><journal-title xml:lang="ru">Известия высших учебных заведений. Приборостроение</journal-title><trans-title-group xml:lang="en"><trans-title>Journal of Instrument Engineering</trans-title></trans-title-group></journal-title-group><issn pub-type="ppub">0021-3454</issn><issn pub-type="epub">2500-0381</issn><publisher><publisher-name>Национальный исследовательский университет ИТМО</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.17586/0021-3454-2026-69-6-514-522</article-id><article-id custom-type="elpub" pub-id-type="custom">pribor-548</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>ВЫЧИСЛИТЕЛЬНЫЕ СИСТЕМЫ И ИХ ЭЛЕМЕНТЫ</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="en"><subject>COMPUTING SYSTEMS AND THEIR ELEMENTS</subject></subj-group></article-categories><title-group><article-title>Гибкая настройка производительности нейросетевого процессора путем параметризации архитектуры</article-title><trans-title-group xml:lang="en"><trans-title>Flexible tuning of neural network processor performance by parametrizing the architecture</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Табунщик</surname><given-names>С. М.</given-names></name><name name-style="western" xml:lang="en"><surname>Tabunshchik</surname><given-names>S. M.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Сергей Михайлович Табунщик — аспирант, факультет программной инженерии и компьютерной техники</p><p>Санкт-Петербург</p></bio><bio xml:lang="en"><p>Sergei M. Tabunshchik — Post-Graduate Student, Faculty of Software Engineering and Computer Technology</p><p>St. Petersburg</p></bio><email xlink:type="simple">sergei_tabunshchik@itmo.ru</email><xref ref-type="aff" rid="aff-1"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>Университет ИТМО</institution><country>Россия</country></aff><aff xml:lang="en"><institution>ITMO University</institution><country>Russian Federation</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2026</year></pub-date><pub-date pub-type="epub"><day>18</day><month>07</month><year>2026</year></pub-date><volume>69</volume><issue>6</issue><fpage>514</fpage><lpage>522</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Национальный исследовательский университет ИТМО, 2026</copyright-statement><copyright-year>2026</copyright-year><copyright-holder xml:lang="ru">Национальный исследовательский университет ИТМО</copyright-holder><copyright-holder xml:lang="en">Национальный исследовательский университет ИТМО</copyright-holder><license xlink:href="https://pribor.ifmo.ru/jour/about/submissions#copyrightNotice" xlink:type="simple"><license-p>https://pribor.ifmo.ru/jour/about/submissions#copyrightNotice</license-p></license></permissions><self-uri xlink:href="https://pribor.ifmo.ru/jour/article/view/548">https://pribor.ifmo.ru/jour/article/view/548</self-uri><abstract><p>Рассматривается проблема проектирования аппаратных ускорителей для систем краевого искусственного интеллекта с учетом ограничений на используемые ресурсы. Разработана модульная параметризуемая архитектура нейросетевого процессора, которая позволяет гибко настраивать характеристики ускорителя, включая размеры вычислительных элементов и размер внутренней буферной памяти, под требования целевой системы. Эффективность полученных результатов подтверждена путем моделирования нейросетевого ускорителя на уровне регистровых передач с использованием параметров сверточного слоя нейронной сети YOLOv5s. Исследована зависимость площади кристалла от параметров архитектуры с использованием блока программируемой логики (ПЛИС) системы на кристалле ZYNQ-7000. Показано, что увеличение размеров систолического массива повышает производительность, но эффект прироста снижается при больших значениях. Масштабирование систолического массива также увеличивает количество логических элементов (LUT, Look-Up Table) и однобитных регистров (FF, Flip-Flop) ПЛИС. Увеличение объема внутренней памяти существенно влияет на потребление элементов блочной памяти (BRAM) ПЛИС, но не оказывает значительного воздействия на общую производительность. Результаты экспериментов подтвердили преимущество предложенного подхода в части адаптивности и гибкости настройки под различные сценарии применения. Перспективными областями применения являются системы, требующие высокой производительности при ограниченных ресурсах. Полученные результаты могут служить основой для разработки новых поколений нейросетевых процессоров с возможностью масштабирования производительности и учетом ограничений на ресурсы.</p></abstract><trans-abstract xml:lang="en"><p>The problem of designing hardware accelerators for edge artificial intelligence systems is considered, taking into account the limitations on the resources used. A modular parameterizable architecture of the neural network processor is been developed, which allows flexibly adjusting the characteristics of the accelerator, including the size of computing elements and the size of the internal buffer memory, to the requirements of the target system. The effectiveness of the obtained results is confirmed by modeling a neural network accelerator at the level of register transfers using the parameters of the convolutional layer of the YOLOv5s neural network. The dependence of the crystal area on the architecture parameters using the programmable logic unit (Field Programmable Gate Array, FPGA) of the ZYNQ-7000 system on a chip is investigated. It has been shown that increasing the size of the systolic array increases productivity, but the effect of the increase decreases with large values. Scaling the systolic array also increases the number of logic gates and single-bit registers of the FPGA. Increasing the amount of internal memory significantly affects the consumption of FPGA block memory elements, but does not significantly affect overall performance. The experimental results confirmed the advantage of the proposed approach in terms of adaptability and flexibility of customization for various application scenarios. Promising areas of application are systems that require high performance with limited resources. The results obtained can serve as a basis for the development of new generations of neural network processors with the ability to scale performance and taking into account resource constraints.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>нейросетевой процессор</kwd><kwd>модульная архитектура</kwd><kwd>параметризуемая микроархитектура</kwd><kwd>производительность</kwd><kwd>площадь кристалла</kwd><kwd>аппаратный ускоритель</kwd><kwd>ускорение нейронных сетей</kwd><kwd>систолический массив</kwd></kwd-group><kwd-group xml:lang="en"><kwd>neural network processor</kwd><kwd>modular architecture</kwd><kwd>parameterizable microarchitecture</kwd><kwd>performance</kwd><kwd>crystal area</kwd><kwd>hardware accelerator</kwd><kwd>neural network acceleration</kwd><kwd>systolic array</kwd></kwd-group></article-meta></front><back><ref-list><title>References</title><ref id="cit1"><label>1</label><citation-alternatives><mixed-citation xml:lang="ru">Nakhle F. Shrinking the giants: Paving the way for TinyAI // Device. 2024. Vol. 2, N 8. Р. 100411. DOI:10.1016/j.device.2024.100411.</mixed-citation><mixed-citation xml:lang="en">Nakhle F. Device, 2024, no. 8(2), pp. 100411, DOI:10.1016/j.device.2024.100411.</mixed-citation></citation-alternatives></ref><ref id="cit2"><label>2</label><citation-alternatives><mixed-citation xml:lang="ru">Gill S. S. et al. Edge AI: A taxonomy, systematic review and future directions // Cluster Computing. 2025. Vol. 28, N 1. Р. 1–53.</mixed-citation><mixed-citation xml:lang="en">Gill S.S. et al. Cluster Computing, 2025, no. 1(28), pp. 1–53.</mixed-citation></citation-alternatives></ref><ref id="cit3"><label>3</label><citation-alternatives><mixed-citation xml:lang="ru">Chen Y. H. et al. Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks // IEEE Journal of Solid-State Circuits. 2016. Vol. 52, N 1. Р. 127–138.</mixed-citation><mixed-citation xml:lang="en">Chen Y.H. et al. IEEE Journal of Solid-State Circuits, 2016, no. 1(52), pp. 127–138.</mixed-citation></citation-alternatives></ref><ref id="cit4"><label>4</label><citation-alternatives><mixed-citation xml:lang="ru">Chen Y. H. et al. Eyeriss v2: A flexible accelerator for emerging deep neural networks on mobile devices // IEEE Journal on Emerging and Selected Topics in Circuits and Systems. 2019. Vol. 9, N 2. Р. 292–308.</mixed-citation><mixed-citation xml:lang="en">Chen Y.H. et al. IEEE Journal on Emerging and Selected Topics in Circuits and Systems, 2019, no. 2(9), pp. 292–308.</mixed-citation></citation-alternatives></ref><ref id="cit5"><label>5</label><citation-alternatives><mixed-citation xml:lang="ru">Moons B. et al. 14.5 envision: A 0.26-to-10tops/w subword-parallel dynamic-voltage-accuracy-frequency-scalable convolutional neural network processor in 28nm fdsoi // 2017 IEEE Intern. Solid-State Circuits Conf. (ISSCC). IEEE, 2017. Р. 246–247.</mixed-citation><mixed-citation xml:lang="en">Moons B. et al. 2017 IEEE International Solid-State Circuits Conference (ISSCC), IEEE, 2017, рр. 246–247.</mixed-citation></citation-alternatives></ref><ref id="cit6"><label>6</label><citation-alternatives><mixed-citation xml:lang="ru">Yin S. et al. A high energy efficient reconfigurable hybrid neural network processor for deep learning applications // IEEE Journal of Solid-State Circuits. 2017. Vol. 53, N 4. Р. 968–982.</mixed-citation><mixed-citation xml:lang="en">Yin S. et al. IEEE Journal of Solid-State Circuits, 2017, no. 4(53), pp. 968–982.</mixed-citation></citation-alternatives></ref><ref id="cit7"><label>7</label><citation-alternatives><mixed-citation xml:lang="ru">Jouppi N. P. et al. In-datacenter performance analysis of a tensor processing unit // Proc. of the 44th Ann. Intern. Symp. on Computer Architecture. 2017. Р. 1–12.</mixed-citation><mixed-citation xml:lang="en">Jouppi N.P. et al. Proceedings of the 44th Annual International Symposium on Computer Architecture, 2017, рр. 1–12.</mixed-citation></citation-alternatives></ref><ref id="cit8"><label>8</label><citation-alternatives><mixed-citation xml:lang="ru">Jouppi N. P. et al. A domain-specific supercomputer for training deep neural networks // Communications of the ACM. 2020. Vol. 63, N 7. Р. 67–78.</mixed-citation><mixed-citation xml:lang="en">Jouppi N.P. et al. Communications of the ACM, 2020, no. 7(63), pp. 67–78.</mixed-citation></citation-alternatives></ref><ref id="cit9"><label>9</label><citation-alternatives><mixed-citation xml:lang="ru">Jouppi N. et al. Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings // Proc. of the 50th Ann. Intern. Symp. on Computer Architecture. 2023. Р. 1–14.</mixed-citation><mixed-citation xml:lang="en">Jouppi N. et al. Proceedings of the 50th Annual International Symposium on Computer Architecture, 2023, рр. 1–14.</mixed-citation></citation-alternatives></ref><ref id="cit10"><label>10</label><citation-alternatives><mixed-citation xml:lang="ru">Chatha K. Qualcomm® cloud Al 100: 12TOPS/W scalable, high performance and low latency deep learning inference accelerator // 2021 IEEE Hot Chips 33 Symposium (HCS). IEEE, 2021. Р. 1–19.</mixed-citation><mixed-citation xml:lang="en">Chatha K. 2021 IEEE Hot Chips 33 Symposium (HCS), IEEE, 2021, рр. 1–19.</mixed-citation></citation-alternatives></ref><ref id="cit11"><label>11</label><citation-alternatives><mixed-citation xml:lang="ru">Li Y. et al. Ascend: A scalable and energy-efficient deep neural network accelerator with photonic interconnects // IEEE Transactions on Circuits and Systems I: Regular Papers. 2022. Vol. 69, N 7. Р. 2730–2741.</mixed-citation><mixed-citation xml:lang="en">Li Y. et al. IEEE Transactions on Circuits and Systems I: Regular Papers, 2022, no. 7(69), pp. 2730–2741.</mixed-citation></citation-alternatives></ref><ref id="cit12"><label>12</label><citation-alternatives><mixed-citation xml:lang="ru">Tabunshchik S., Bykovskii S., Kustarev P. Neural processor microarchitecture design for CNN processing acceleration // Optoelectronic Imaging and Multimedia Technology XI. SPIE, 2024. Vol. 13239. Р. 141–149.</mixed-citation><mixed-citation xml:lang="en">Tabunshchik S., Bykovskii S., Kustarev P. Optoelectronic Imaging and Multimedia Technology XI, SPIE, 2024, vol. 13239, рр. 141–149.</mixed-citation></citation-alternatives></ref><ref id="cit13"><label>13</label><citation-alternatives><mixed-citation xml:lang="ru">Peddawad C. et al. Matrix-matrix multiplication using systolic array architecture in bluespec // CS6230: CAD for VLSI. 2015. Vol. 1, N 8.</mixed-citation><mixed-citation xml:lang="en">Peddawad C. et al. CS6230: CAD for VLSI, 2015, no. 8(1).</mixed-citation></citation-alternatives></ref><ref id="cit14"><label>14</label><citation-alternatives><mixed-citation xml:lang="ru">Yu K. et al. MobileNet-YOLO v5s: An improved lightweight method for real-time detection of sugarcane stem nodes in complex natural environments // IEEE Access. 2023. Vol. 11. Р. 104070–104083.</mixed-citation><mixed-citation xml:lang="en">Yu K. et al. IEEE Access, 2023, vol. 11, рр. 104070–104083.</mixed-citation></citation-alternatives></ref><ref id="cit15"><label>15</label><citation-alternatives><mixed-citation xml:lang="ru">Khan A. et al. A survey of the recent architectures of deep convolutional neural networks // Artificial Intelligence Review. 2020. Vol. 53, N 8. Р. 5455–5516.</mixed-citation><mixed-citation xml:lang="en">Khan A. et al. Artificial Intelligence Review, 2020, no. 8(53), pp. 5455–5516.</mixed-citation></citation-alternatives></ref></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
