Preview

Journal of Instrument Engineering

Advanced search
Open Access Open Access  Restricted Access Subscription Access

Cluster reliability during recovery with container migration

https://doi.org/10.17586/0021-3454-2026-69-5-417-427

Abstract

When designing high-availability, fault-tolerant distributed computing systems that must ensure continuous operation with minimal request-processing delays, the consolidation of redundant computing resources through their integration into clusters with container virtualization is becoming increasingly important. To justify architectural choices and enable structural-parametric optimization of such fault-tolerant clusters, it is essential to develop reliability models that reflect the operational specifics of virtual containers as well as the processes of cluster recovery and reconfiguration during container migration. The aim of this article is to construct an analytical reliability model for a multi-node cluster employing container virtualization, in which computing nodes (servers) host multiple containers. A Markov model is proposed that captures the two-stage recovery of cluster nodes: physical restoration of servers followed by container migration with their sequential loading once the server has been physically recovered from a failure. An analysis is carried out to evaluate how the cluster’s availability coefficient depends on its architectural parameters, including the number of servers, hardware failure rates, and the number of deployed containers loaded during migration. The proposed model provides a foundation for substantiating design decisions in the development of fault-tolerant container-based clusters and modern cloud platforms.

About the Authors

M. K. Do
ITMO University
Russian Federation

Do Manh Kiem — Post-Graduate Student; Faculty of Software Engineering and Computer Science

St. Petersburg 



V. A. Bogatyrev
ITMO University
Russian Federation

Vladimir A. Bogatyrev — Dr. Sci, Professor; Faculty of Software Engineering and Computer Science

St. Petersburg 



V.Q. Phung
Le Quy Don Technical University
Viet Nam

Phung Van Quy — Researcher

Ha Noi 



References

1. Aysan H. Fault-Tolerance Strategies and Probabilistic Guarantees for Real-Time Systems, Vasteras, Sweden, Malardalen University, 2012, 190 p.

2. Kopetz H. Real-Time Systems: Design Principles for Distributed Embedded Applications, Springer, 2011, 396 p., DOI: 10.1007/978-1-4419-8237-7.

3. Sorin D.J. Fault Tolerant Computer Architecture, Morgan & Claypool, 2009, 103 p.

4. Koren I., Krishna C.M. Fault Tolerant Systems, San Francisco, Morgan Kaufmann Publishers, 2009, 378 p.

5. Goyal P., Deora S.S. Indian Journal of Cryptography and Network Security, 2022, no. 1(2), pp. 1–5, DOI: 10.54105/ijcns.C1417.051322.

6. Khersonsky N.S., Bolshedvorskaya L.G. Crede Experto: transport, society, education, language, 2024, no. 1, pp. 6–23, DOI: 10.51955/2312-1327_2024_1_6. (in Russ.)

7. Srivastava A., Kumar N. International Journal of Advanced Computer Science and Applications, 2023, no. 1(14), pp. 465–472, DOI: 10.14569/IJACSA-.2023.0140150.

8. Gurjanov A.V., Korobeynikov A.G., Zharinov I.O., Zharinov O.O. III International Workshop on Modeling, Information Processing and Computing, 2021, pp. 103–108.

9. Shukur H.M., Zeebaree S.R.M., Zebari R.R., Zeebaree D.Q., Ahmed O.M., Salih A.A. Journal of Applied Science and Technology Trends, 2020, no. 2(1), pp. 98–105, DOI: 10.38094/jastt1331.

10. Kushchazli A., Safargalieva A., Kochetkova I., Gorshenin A. Mathematics, 2024, no. 3(12), pp. 468, DOI:10.3390/math12030468.

11. Tatarnikova T.M., Arkhiptsev E.D. Proc. of the 27th Intern. Conf. on Soft Computing and Measurements (SCM), 2024, pp. 348–351, DOI: 10.1109/SCM62608.2024.10554143.

12. Bogatyrev V.A., Bogatyrev S.V. Journal of Instrument Engineering, 2016, no. 9(59), pp. 735–740. (in Russ.)

13. Sovetov B.Ya., Tatarnikova T.M., Poymanova E.D. Information and Control Systems, 2020, no. 5(108), pp. 43–49, DOI: 10.31799/1684-8853-2020-5-43-49.

14. Krylov D.R., Poimanova E.D., Tyurlikov A.M. Information Management Systems, 2024, no. 3, pp. 11–23, DOI:10.31799/1684-8853-2024-3-11-23. (in Russ.)

15. Bogatyrev V.A., Bogatyrev A.V. Information Technologies, 2015, no. 7(21), pp. 495–502. (in Russ.)

16. Bogatyrev V.A., Bogatyrev A.V., Bogatyrev S.V. Journal of Instrument Engineering, 2014, no. 4(57), pp. 46–48. (in Russ.)

17. Arustamov S.A., Bogatyrev V.A., Polyakov V.I. Advances in Intelligent Systems and Computing, 2016, vol. 451, pp. 103–109.

18. Bogatyrev V.A. Instruments and Systems: Monitoring, Control, and Diagnostics, 2006, no. 10, pp. 18–21. (in Russ.)

19. Polovko A.M., Gurov S.V. Fundamentals of Reliability Theory, St. Petersburg, 2006, 702 p. (in Russ.)

20. Bogatyrev V., Derkach A. Computers, 2020, no 2(9), pp. 42.

21. Bogatyriev V.A., Bogatyriev S.V., Bogatyriev A.V. Scientific and Technical Journal of Information Technologies, Mechanics and Optics, 2023, no. 3(23), pp. 608–617, DOI: 10.17586/2226-1494-2023-23-3-608-617. (in Russ.)

22. Bogatyrev V.A., Phung V.Q. Scientific and Technical Journal of Information Technologies, Mechanics and Optics, 2025, no. 5(25), pp. 988–995, DOI: 10.17586/2226-1494-2025-25-5-988-995. (in Russ.)

23. Phung Van Quy, Bogatyrev V.A., Karmanovskiy N.S., Le V. Scientific and Technical Journal of Information Technologies, Mechanics and Optics, 2024, no. 2(24), pp. 249–255, DOI: 10.17586/2226-1494-2024-24-2-249-255. (in Russ.)

24. Compastié M., Badonnel R., Festor O., He R. Computers & Security, 2020, vol. 97, pp. 101905, DOI: 10.1016/j.cose.2020.101905.

25. Choudhary A., Govil M.C., Singh G., Awasthi L.K., Pilli E.S., Kapil D. Journal of Cloud Computing, 2017, vol. 6, pp. 23, DOI: 10.1186/s13677-017-0092-1.


Review

For citations:


Do M.K., Bogatyrev V.A., Phung V. Cluster reliability during recovery with container migration. Journal of Instrument Engineering. 2026;69(5):417-427. (In Russ.) https://doi.org/10.17586/0021-3454-2026-69-5-417-427

Views: 192

JATS XML

ISSN 0021-3454 (Print)
ISSN 2500-0381 (Online)