Preview

Scientific and Technical Journal of Information Technologies, Mechanics and Optics

Advanced search

The journal "Scientific and Technical Journal of Information Technologies, Mechanics and Optics" is one of the oldest scientific periodicals in Russia based on a technical educational institution. The first issue dates back to 1936. For many years, the journal was published under the name "Proceedings of the Leningrad Institute of Precision Mechanics and Optics". The publication was resumed in 2001 as a scientific research and educational periodical.

Due to the change in the status of the University, the journal changed its name several times: "Scientific and Technical Bulletin of the St. Petersburg State Institute of Precision Mechanics and Optics (Technical University)" - until 2003; "Scientific and Technical Bulletin of the St. Petersburg State University of Information Technologies, Mechanics and Optics" - from 2004 to 2011.

In December 2011, the journal was published under its current name "Scientific and Technical Journal of Information Technologies, Mechanics and Optics".

The journal is included in the Scopus database

The journal is included in the Russian Science Citation Index (RSCI)

The journal is included in the Directory of Open Access Journals (DOAJ).

The journal is catalogued in Ulrich's Periodicals Directory.

All articles are posted on the platform of the Scientific Electronic Library http://elibrary.ru/

Editorial address: office 2136, ITMO University, Lomonosov st., 9, Saint Petersburg, Russian Federation
Correspondence address: ITMO University, Kronverksky pr., 49, litera A, Saint Petersburg, Russian Federation, 197101
Phone: +7 (812) 480-02-75
E-mail: ntvitmo@itmo.ru

Current issue

Vol 26, No 4 (2026)
View or download the full issue PDF (Russian)

REVIEW PAPERS

673-682 2
Abstract

A systematic review of modern neural network methods for Speech Enhancement is presented, aimed at improving speech intelligibility and quality under acoustic distortions. The reviewed methods can be applied in voice control systems, telecommunications, hearing aids, and human–machine interaction interfaces. Key architectural approaches are considered, including classical recurrent and convolutional networks as well as modern hybrid architectures with attention mechanisms (Transformer, Conformer), state-space models (Mamba), and advanced recurrent blocks (xLSTM). The advantages and disadvantages of different architectures are shown in terms of speech restoration quality and computational efficiency. Specific problems of existing methods are highlighted, including high computational cost and insufficient generalization capability under non-stationary noise conditions. The need for further research in the development of lightweight models for mobile devices, multi-distortion suppression methods, and the integration of neural network noise suppression with generative models to achieve a new level of speech signal restoration quality is demonstrated.

683-694 7
Abstract

This review article presents a generalized, integrated classification of scoring functions for molecular docking of protein interaction models. It describes various types of scoring functions and examines relevant issues in biopharmaceuticals, medical cybernetics, computational biology and biophysics, applied mathematics. This article may be of interest to a wide range of readers, including scientists from various fields, software engineers developing automated bioinformation processing systems, lecturers in relevant topics, graduate and postgraduate students in bioinformatics, systems analysis, medical software systems research.

OPTICAL ENGINEERING

695-704 4
Abstract

The paper presents the results of an investigation that employed vibrational spectroscopy as a monitoring technique for greenhouse gases at the “Rossyanka” Carbon Measurement Supersite, in Kaliningrad Oblast. The application of these methodologies is indicative of their efficacy in the quantitative detection of greenhouse gases, with a particular focus on carbon dioxide (CO2). The capacity of infrared (IR) and Raman spectroscopy (RS) for the monitoring of O2 in field conditions is illustrated, with a discrepancy of approximately 100 ppm observed between the calculated and experimental values of concentrations in field samples. Methods of IR and RS using a Virsa (Renishaw) and Shimadzu IRPrestige 21 spectrometer have been applied. The survey was carried out by recording the spectra of gases sealed in a glass vial for medicines with a volume of 10 ml. The obtained samples of the spectra of the gas mixture were compared with the spectra of a reference sample with a known concentration. The results demonstrate the feasibility of using IR and RS wile determine greenhouse gas concentrations at carbon supersites throughout the country, including the “Rossyanka” Carbon Measurement Supersite where CO2 concentrations can be determined with a registration threshold of 300500 ppm and the isolation of the main characteristic vibrational modes. For the samples obtained at the “Rossyanka” Carbon Measurement Supersite, the applied vibrational spectroscopy methods enabled the quantification of only the concentration of CO2. The presence of CH4 was not detected in the samples due to the low concentrations of the gas which were found to be as low as 10 ppm according to gas chromatography. It is expected that further studies using signal amplification methods, such as using multi-pass cuvettes, will solve this problem when registering low greenhouse gas concentrations. It is expected that further studies using signal amplification methods, such as using multi-pass cuvettes and RS, will solve this problem when registering low greenhouse gas concentrations. 

705-716 7
Abstract

Laser-based material processing on inclined or non-uniform surfaces is often affected by beam defocus and alignment errors which can reduce process accuracy and repeatability. Adaptive laser mechanisms provide a promising approach to mitigating these effects; however, their practical implementation requires validated system architectures that support reliable integration of sensing, actuation, and control functions. This paper presents the design and experimental evaluation of a prototype adaptive laser mechanism developed as a dedicated test platform for laser beam autofocusing and laser beam detection subsystems. The prototype is intentionally constrained to a single rotational degree of freedom to limit mechanical and computational complexity while retaining the essential characteristics of adaptive beam control. The experimental setup integrates a continuous-wave visible laser source operating at a wavelength of 532 nm, a camerabased sensing unit with polarization filtering, an actuator system, an MPU6050 inclination measurement subsystem, and a rotating platform to introduce controlled surface inclinations. Experimental evaluation conducted at surface inclination angles up to 0.262 rad demonstrated stable and repeatable laser beam projection under varying angular conditions. Quantitative assessment of the inclination measurement and positioning subsystems showed a maximum measurement error of 0.001 rad, settling times between 2.52 s and 5.18 s, and steady-state positioning errors below 0.001 rad. Qualitative observations further confirmed the capability of the prototype to support systematic investigation of laser beam behavior in response to surface inclination, although visual differences in beam spot size, shape, and intensity distribution remained subtle. The presented prototype contributes a practical and scalable foundation for the development of adaptive laser mechanisms by enabling controlled evaluation of beam projection performance in a simplified experimental environment. The results indicate that quantitative characterization of laser beam behavior cannot rely on visual inspection alone, thereby motivating the development of dedicated laser beam detection methods. Future work will focus on extending the mechanism to multiple degrees of freedom and on the quantitative analysis of laser beam spot characteristics using advanced image-based detection algorithms.

717-723 11
Abstract

Baltic amber possesses exceptional optical properties and transparency, making it considered one of the most valuable materials. Analysis of its structural features allows one to assess the purity and authenticity of amber. This study aims to understand the fundamental processes within the amber polymer matrix. The development of new, more reliable criteria for amber authentication and classification, as well as a fundamental understanding of the relationship between molecular dynamics and the macroscopic properties of natural polymers, is relevant. This paper presents the results of a study of molecular dynamics in amber and its interaction with water. Currently, there are no published values for the T1 and T2 nuclear magnetic resonance (NMR) relaxation times in amber which is a complex amorphous natural polymer with variable composition and structure. The aim of this study was to systematize existing scientific papers on the application of nuclear magnetic resonance to amber samples, identify gaps in current knowledge, and substantiate the relevance of new, more sophisticated techniques, such as 2D T1-T2 NMR relaxometry, for studying the complex structure of the polymer. An analysis of the relationships between relaxation characteristics, porosity, and water saturation of amber was conducted. The use of 2D T1-T2 relaxometry method enabled a transition from a phenomenological description to a mechanistic understanding of the structure and properties of amber at the molecular level. It has been established that there are two types of pores in amber, the sizes of which differ by two orders of magnitude. It is shown that water treatment leads to the appearance of new components with dynamics different from the original matrix, which is interpreted as the filling of micropores with water and a change in the flexibility of polymer chains. Information about the pores is of interest from the point of view of amber processing technology, since the size of the pores determines its fragility and cracking ability. The work established a correlation between relaxation times and the degree of saturation of Baltic amber with water. Fused amber has shorter T2 times, which is probably due to a decrease in molecular mobility due to the formation of additional cross-links in the polymer matrix. Conversely, the T1 time may exhibit more complex behavior related to the efficiency of cross-relaxation. The method enables selective identification of components: separating signals from a rigid polymer network (long T1, short T2 — high T1/T2 ratio), mobile molecular fragments (moderate T1 and T2), and possible liquid inclusions (short T1, long T2 — low T1/T2 ratio). The implementation of 2D T1-T2 relaxometry method is a relevant and promising direction in studying the structure and properties of complex polymers.

724-731 7
Abstract

 CsPb(BrхCl1–х)3 perovskite nanocrystals were synthesized in fluorophosphate glass by melt quenching followed by subsequent heat treatment. The variation of the chlorine (Cl) and bromine (Br) ion ratio in the glass affected the composition and, consequently, the band gap ΔЕg of the perovskite nanocrystals, enabling tunable photoluminescence from green to blue emission. The spectral and luminescent properties of CsPb(BrхCl1–х)3 nanocrystals formed in fluorophosphate glass were investigated for different Br/Cl ratios. Fluorophosphate glasses containing CsPb(BrхCl1–х)3 nanocrystals were prepared by high-temperature synthesis from batch chemicals followed by additional heat treatment above the glass transition temperature. The heat-treatment temperature was determined from differential scanning calorimetry data obtained using a STA 449F1 Jupiter Nietzsche analyzer. Absorption spectra were measured with a PerkinElmer Lambda 650 double-beam spectrophotometer. Photoluminescence spectra were recorded using a PerkinElmer LS50B spectrofluorometer. The absolute photoluminescence quantum yield was measured with a Hamamatsu absolute quantum yield measurement system equipped with an integrating sphere. CsPb(BrхCl1–х)3 perovskite nanocrystals with different Br/Cl ratios were successfully formed in fluorophosphate glass. The growth of nanocrystals in the glass matrix was controlled by isothermal treatment at temperatures above Tg through adjustment of the treatment temperature and duration. Optical measurements confirmed the formation of CsPb(BrхCl1–х)3 nanocrystals. The photoluminescence peak position shifted within the range of 430–512 nm. A decrease in the photoluminescence quantum yield of CsPb(BrхCl1–х)3 was observed with increasing Cl concentration in the nanocrystals. This effect can be attributed to the enhancement of nonradiative recombination pathways in the glass matrix, including the formation of deep trap states and increased surface recombination. It is concluded that fluorophosphate glasses containing CsPb(BrхCl1–х)3 nanocrystals can be used as blue and green phosphors in the 450–500 nm spectral range.

MATERIAL SCIENCE AND NANOTECHNOLOGIES

732-738 3
Abstract

In the context of global climate change, the imperative for decarbonization and improvement in efficiency of processes and equipment within the energy sector has become increasingly urgent. One promising approach to improving energy efficiency involves the recovery of low-grade waste heat through the implementation of thermoelectric generators (TEG). Silicide-based materials are considered promising for such devices. However, their utilization is constrained by insufficient thermal stability at elevated temperatures. Specifically, the efficient material Mg2Si0.4Sn0.6 undergoes degradation above 400 °C, whereas the more thermally stable Mg2Si has lower thermoelectric performance. To address this limitation, a segmented design of the n-type leg in TEG is proposed, integrating both materials. The objective of this study is to model and optimize a TEG with a segmented n-type leg, thereby extending the operational temperature range while maintaining high energy conversion efficiency. The simulation was performed with COMSOL Multiphysics software. A three-dimensional stationary model of a TEG consisting of 127 leg pairs has been developed. The p-type leg was composed of the higher manganese silicide MnSi1.75. The segmented n-type leg consisted of Mg2Si on the hot side and Mg2Si0.4Sn0.6 on the cold side. The model employed modules to describe the processes of thermal transport and electrical transport, taking into account the thermoelectric effect. Thermal contact between different parts of the generator was taken into account. To analyze the influence of geometry, the height ratio of the segments varied (with Mg2Si0.4Sn0.6 ranging from 30 % to 80 % of the total leg height). According to the simulation results, the efficiency of the module with a segmented n-type leg increases to 5.3 % compared to 2.9 % for a TEG with n-type legs consisting only of Mg2Si, at a hot-side temperature of 500 °C. With an optimized segmented-leg configuration, this temperature regime becomes acceptable, and the simulated performance is comparable to commercially available generators. The simulation results demonstrate that the proposed segmented silicide-based TEG effectively addresses the thermal stability problem of Mg2Si0.4Sn0.6. This TEG configuration substantially extends the generator operating temperature range while maintaining an acceptable level of efficiency.

739-752 3
Abstract

Advances in computing systems make it possible to simulate relatively complex physical processes for the subsequent application of simulation results in scientific research and technological processes. The paper presents the results of the development of a model, simulation methods, and a full-scale experiment, as well as the application of these developments to the study of certain linear and nonlinear dynamic photoelastic properties of transparent piezocrystals, using lithium niobate as an example. A mathematical model and an experimental setup were developed to study the linear and nonlinear dynamic photoelastic and electro-optical properties of piezocrystals. Full-scale experiments were carried out. An experimental method based on simultaneous measurement and subsequent processing of current, voltage, and intensity of polarization-phase modulation due to dynamic photoelasticity was proposed, which is especially valuable for studying nonlinear phenomena. By taking into account the spatial dispersion and dissipation of wave energy, the developed model makes it possible to simulate the optical effect with high accuracy — dynamic rotation of polarization, i.e., dynamic photoelasticity. A comparison of the simulation results with experimental data, based on analysis, allows us to conclude that the simulation results are consistent with the experimental data and that the mathematical model can be used to solve systemic problems related to the study of the optical properties of perturbed piezocrystals. The developed model makes it possible to determine the distribution of elastic deformations of a crystalline object under piezoelectric disturbance; to identify resonant frequencies, the type and order of wave modes; to determine the rotation angles of polarization for dynamic photoelasticity for arbitrary optical radiation propagation trajectories. The experimental results will make it possible to map the distribution inside the crystal (electric field, strain such as stretching/compression and shear, refractive coefficients) and calculate the conversion of wave energy of optical modes. The practical significance of the research lies in the developed mathematical model and experimental procedure for recording linear and nonlinear effects of dynamic photoelasticity of piezo crystals. The results of the work can be useful in scientific and practical tasks, for example, in the research and design of acousto-optical modulators, the development of devices and systems for wavefront reversal on nonlinear crystals, the construction of various sensors, etc.

AUTOMATIC CONTROL AND ROBOTICS

753-762 4
Abstract

This paper considers the problem of carrier frequency estimation from measurements of an amplitude-modulated signal in the presence of unknown bounded disturbances and a time-varying envelope. Existing identification methods based on regression transformations usually require a small-disturbance condition, which limits their practical applicability. The scientific novelty of this work consists in the development of a robust carrier-frequency estimation algorithm that preserves the regression structure of delay-based signal parameterization and does not require the assumption of small disturbance magnitude. The proposed method is based on the parameterization of the measured signal using delayed measurements, which makes it possible to reduce the problem to a linear regression model. To estimate the unknown parameter, an adaptive algorithm with a saturation nonlinearity is used, limiting the influence of large regression errors. After estimating the regression parameter, the carrier frequency is reconstructed. The effectiveness of the proposed method is verified through numerical experiments for amplitude-modulated signals under noise and impulsive disturbances. A comparison with a normalized gradient algorithm and a power-transformed regression method is performed. The simulation results demonstrate a faster reduction of the estimation error and improved robustness of the proposed algorithm. The obtained results show the advantage of the proposed method over existing frequencyestimation algorithms under strong disturbances. The method can be applied to radio-signal processing, communication systems, and modulation-recognition algorithms. Future research directions include the development of discrete-time implementations of the algorithm and its extension to multi-component signals.

COMPUTER SCIENCE

763-770 10
Abstract

Often, the initial data sets for machine learning contain an excessive number of features, among which there may be highly informative and insignificant, correlating or even noise variables. Solving the feature selection problem not only optimizes machine learning processes, but also opens up new opportunities for implementing machine learning in practice-oriented areas that require high accuracy and trust in algorithms. This phenomenon leads to a number of critical problems: distorted training of the model based on irrelevant patterns, a decrease in its generalizing ability, a sharp increase in computational costs for training and, as a result, difficulty in interpreting the results obtained. The paper proposes a method for optimizing the feature space which uses a combined approach that includes maximizing the measure of model effectiveness, selecting the highest-quality features based on information content and correlation analysis. In addition to optimizing the feature space, the model offers the best classifier for the data set used. As a solution to the objective function, a genetic algorithm with elitism is used to find the optimal value. The proposed model is demonstrated on a dataset in the field of medical diagnostics in comparison with the well-known method of recursive feature selection (RFE), correlation analysis, as well as with the results obtained using the full dataset. The data set consists of measurements of mammary glands using the method of microwave radiothermometry and labels characterizing the severity of temperature anomalies. The dataset contains 62 features and 6 labels and 9,310 measurement records. The results demonstrate the high efficiency of the proposed model either close to or exceeding the value of RFE. As a result of optimizing the feature space for the data set, it was possible to reduce from 62 features to 15, while not only not reducing the accuracy of the model, but even slightly increasing it, the best model of the classifier turned out to be the logistic regression model. Thus, the accuracy indicators were 0.79 for the complete data set, 0.7879 for 15 optimized features based on the RFE method, 0.7175 for 29 features obtained as a result of correlation analysis, and 0.7911 for 15 features obtained using the proposed model. The proposed method not only effectively reduces the feature space, but also increases the accuracy of the classification model, despite a significant decrease in the number of features.

771-782 6
Abstract

Cooperative perception based on Vehicle-to-Everything (V2X) technologies is effective methodology that allows the use of various data sources, such as sensors, cameras, and sensors, enabling unmanned vehicles to exchange data to overcome the limitations of individual perception. However, the integration of multi-agent data systems creates additional vector attack in cooperative perception and information integrity protection. A quantitative metric of information integrity I𝒜(tk), is introduced, based on pairwise consistency of agent observations using Mahalanobis distance. A threat model is developed that classifies types of attacks in V2X systems. A model has been developed that includes the state of agent xi(t) = (pi(t), vi(t), θi(t)), local perception 𝒪i(t), dynamic connectivity graph G(t), and the proposed metric I𝒜(tk), has been described, which quantitatively assesses the integrity of information exchange. The proposed method surpasses analogues in key metrics: area Under the ROC curve (Receiver Operating Characteristic), (Area Under the Curve, AUC) = 0.94 (95 % confidence interval [0.92; 0.96]) versus AUC = 0.87 for the best competitor (MATE, Multi-Agent Trust Estimator; DeLong criterion, p less than 0.05), where p is the achieved level of significance of the statistical criterion. Further prospects and areas of application of the model will focus on conducting experimental validation of the proposed structure based on the V2X-Seq dataset. Carrying out integration with formal verification through the Hamilton–Jacobi equation, performing reachability analysis to ensure security guarantees, introducing trust assessments with adaptive component weights to identify specific Byzantine agents.

783-792 3
Abstract

Given the steady increase in the number of attacks on web applications, the task of automatically detecting such attacks requests remains one of the key challenges in information security. The aim of this work is to develop a new hybrid approach to detecting attacks on web applications in which the feature extraction stage is shifted to the level of pretrained transformers. To achieve this goal, the paper proposes a hybrid method for detecting attacks on web applications based on the combined use of specialized transformer models and neural networks of various architectures. The suggested approach offers that the request URL is processed by a URLBERT model trained on the structural features of web addresses, while the request content is analyzed by a SecRoBERT model tailored for cybersecurity tasks. The resulting contextual representations are combined and fed into parallel blocks of a convolutional neural network and a multilayer perceptron, which extract local and global features, respectively. The proposed approach combines the capabilities of transformer data processing models with the efficient identification of statistical and local dependencies based on a convolutional neural network and a multilayer perceptron, which improves the accuracy of malicious query classification. Experimental evaluation is conducted on three datasets, two of which are used for binary classification and one for multi-class classification. Testing was performed using k-fold cross-validation and by measuring the model output latency to demonstrate the method applicability in real-time attack detection systems. The results of the experimental evaluation of the proposed model demonstrate high accuracy in detecting attacks on web applications, but also reveal a significant drawback related to the considerable time required to process requests. Analysis of the test results and comparison with relevant studies confirm the feasibility of using this model as the second stage of request analysis in web application attack detection systems.

793-805 7
Abstract

Azimuth-resolved optical scattering signals from cell nuclei carry rich information about internal refractive index variations. These two-dimensional patterns provide valuable insights into chromatin organization which plays a critical role in early cancer detection. Our goal was to investigate whether two-dimensional scattering signals could serve as inputs to an inverse problem framework for extracting the spatial correlation length lc and fluctuation magnitude δn of subnuclear refractive index variations using state-of-the-art deep learning approaches. Deriving closed-form expressions connecting azimuth-resolved signals to lc and δn proves intractable, so we turned to a purely data-driven methodology. We implemented a Vision Transformer-based regression pipeline and evaluated it on 198 numerically simulated scattering patterns generated from nuclear models spanning a range of internal structural parameters. We observed strong agreement between ground truth and predicted values for both parameters, achieving mean absolute percent errors of 6.2 % for lc and 9.8 % for δn. Notably, these figures represent substantial improvements over prior Convolutional Neural Network-based methods and fall below the minimum relative spacing between adjacent parameter values in our dataset. Importantly, we demonstrate that the Vision Transformer achieves stable performance on this small dataset through careful regularization including weight decay (1·10–4), dropout (0.1), and early stopping (patience is 15), with learning curves showing no evidence of overfitting. These findings suggest that Vision Transformer architectures offer a promising avenue for mining the information encoded in two-dimensional optical scattering data, delivering markedly better accuracy than conventional Convolutional Neural Network approaches for quantitative chromatin characterization.

806-815 4
Abstract

Modern distributed robotic systems based o n Robot Operating System 2 (ROS 2) critically depend on the quality of the network environment. However, when wireless communication channels, such as Wi-Fi, LoRa, and 4G/5G, are used, various forms of degradation occur, including packet loss, increased jitter, and micro-outages. Standard Data Distribution Service (DDS) mechanisms provide Quality of Service (QoS) metrics, but they do not define application-level adaptation logic. This work addresses the problem of ensuring continuous robot control under unstable communication conditions by developing an intelligent fault-tolerance layer. The authors propose a DDS-aware architecture that integrates QoS Monitor and Policy Engine components. The methodology is based on online monitoring of DDS events, such as DeadlineMissed and LivelinessChanged, as well as statistical analysis of a message delivery window. A finite-state algorithm for classifying the communication state was developed, along with mitigation mechanisms, including dualpath command delivery using primary and backup channels with deduplication based on serial numbers, and dynamic traffic prioritization. Experimental validation was carried out on a Linux-based testbed using tc netem for precise emulation of network degradation scenarios, including 10 % packet loss, delay with jitter, and micro-outages, at a communication frequency of 50 Hz. The comparative analysis demonstrated the superiority of the proposed method over the standard baseline approach. In the scenario with 10 % packet loss, the command delivery ratio, denoted as DR_c, increased from 99.81 % to 100 %, while the 95th percentile latency decreased from 115.52 ms to 3.51 ms. Under micro-outage conditions, the proposed solution completely eliminated control gaps by reducing the maximum blocking time, T_(blk, max), from 1.052 s to 0 s. The system demonstrated a stable recovery time T_rec of approximately 5 s which was determined by the hold-down policy used to prevent oscillations. The results confirm that application-level traffic redundancy and DDS event-based analysis can compensate for physical-layer communication limitations without modifying the ROS 2 core. Unlike standard approaches, the proposed method separates traffic classes, ensuring the delivery of critical control commands even under extreme telemetry degradation. A limitation of the method is the need for preliminary QoS profile configuration; however, this is compensated for by the high portability of the solution across different ROS Middleware implementations, including Fast DDS and Cyclone DDS. This work lays the foundation for the development of self-healing communication environments for mission-critical robotic applications. Robotic systems operating in dynamic and especially wireless environments are subject to communication degradation, including packet loss, increased latency, jitter, and short-term disconnections. These factors reduce the quality of telemetry and may lead to temporary failures in the delivery of control commands. In ROS 2, data exchange is implemented over DDS, which provides QoS policies and events; however, it does not define an application-level fault-tolerance policy. This article proposes a DDS-aware fault-tolerant communication layer for ROS 2 which includes a QoS Monitor and a Policy Engine. The Policy Engine classifies the communication state and triggers mitigation actions, including command prioritization, telemetry stream adaptation, and switching between preconfigured QoS profiles. A reproducible evaluation methodology is proposed for testing the system under artificially introduced communication channel degradations.

816-825 6
Abstract

This study presents a transformer-based Generative Adversarial Network framework designed to model and replicate complex behavioral patterns associated with account manipulation in cybersecurity environments. The primary objective is to generate realistic synthetic user behaviors that emulate malicious activities while ensuring data privacy, thereby providing an ethical and controlled foundation for developing Artificial Intelligence (AI) driven defense mechanisms. The proposed architecture integrates progressive adversarial training, feature matching, and hybrid normalization to stabilize learning and enhance fidelity. A Transformer encoder captures long-range temporal dependencies in behavioral data, while a multi-scale attention discriminator identifies manipulative patterns across varying time scales. A noise-adaptive pre-processing pipeline and composite authenticity evaluation protocol enhance the alignment of synthetic sequences with authentic behavioral statistics. The model surpasses recurrent neural baselines in experimental evaluation, attaining a Fréchet Inception Distance of 18.3, precision of 0.82, and recall of 0.76, alongside enhanced training stability (37 %) and improved behavioral coverage (29 %). Using t-distributed Stochastic Neighbor Embedding and Mahalanobis distance metrics to visualize the data shows that 89 % of the behaviors that were real and those that were generated were the same. This means that the mimicry was very accurate and the timing was very consistent. This study establishes a resilient AI framework for behavioral simulation, facilitating the advancement of adversarially trained, adaptive cybersecurity systems. It stresses even more the need for ethical governance, responsible AI use, and behavioral biometrics to reduce the dual-use risks that come with generative technologies in modern cyber defense.

826-834 8
Abstract

This paper focuses on the development of an automatic speech recognition system for the Livvi-Karelian variety of the Karelian language, as it is spoken under conditions of code-switching between Karelian and Russian. The study of bilingual speech recognition methods is carried out. In order to improve the quality of speech recognition, a methodology for training text data augmentation via partial translation and intra-word code-switched wordforms artificial synthesis was developed. Acoustic modeling was performed by fine-tuning a pre-trained multilingual Wav2Vec2-BERT 2.0 model with the use of the data from two previously collected corpora containing 7.5 hours of speech. Fine-tuning was performed using the Transformers framework. When developing the language model, in order to address the problem of limited code-switching data, an augmentation method was applied based on partial automatic translation of Karelian texts into Russian, followed by the generation of word-forms with intra-word code-switching based on special linguistic rules. On the base of formulated rules, a list of words with intra-word code-switching was generated for a language model. The experiments showed that using a full vocabulary that includes generated hybrid word forms yields a consistent improvement in results. A further reduction in word error rate to 25.82 % on the development set and 29 % on the test portion of the corpus was achieved through linear interpolation of the Karelian language model with the Russian language model (interpolation weight 0.7). The conducted experiments confirm the effectiveness of the developed methodology for developing a bilingual speech recognition system. In particular, it is recommended to combine finetuning of multilingual acoustic models, text augmentation with morphological rules, and language model interpolation. The proposed approach can be applied to developing speech recognition systems for other low-resource languages of Russia spoken in an unbalanced bilingual environment.

835-843 6
Abstract

Filtering methods for three-dimensional object contours in reservoir geological modeling problems are investigated to improve the efficiency of resource base analysis. The problem is solved by applying the mathematical methods of the contourlet transform, diffusion, and contour analysis to filtering the contours of three-dimensional images. The proposed method combines the advantages of multi-scale and multi-directionality, taking into account the local orientation of structures. The problem of three-dimensional images filtering is solved by combining mathematical methods for analyzing three-dimensional images. The essence of the proposed approach consists of using a non-subsampled contourlet transform, anisotropic diffusion controlled by a structural tensor, and subsequent localization of single-voxel contours by suppressing non-maxima and double threshold filtering based on hysteresis. The filtering result is a filtered three-dimensional image with a contour thickness of one voxel. The proposed method showed satisfactory results on test data. The quality of contour filtering allows using the proposed method for subsequent contour identification in three-dimensional objects. The proposed method was compared with three-dimensional gradient filters, the Fourier transform, and the discrete wavelet transform. The method is applicable to problems where the redundancy created by the non-subsampled contourlet transform does not significantly affect the final goal. Further development of the method is seen as taking into account the specifics of the various types of research on the basis of which a three-dimensional model is created.

844-850 6
Abstract

Autoprompting is the process of automatically selecting optimized prompts for language models which is gaining popularity due to the rapid development of prompt engineering driven by extensive research in the field of Large Language Models. This paper presents DistillPrompt — a novel autoprompting method based on Large Language Models that employs a multi-stage integration of task-specific information into prompts using training data. DistillPrompt utilizes distillation, compression, and aggregation operations to explore the prompt space more thoroughly. The method was tested on different datasets for text classification and generation tasks using t-lite-instruct-2.1, gpt-3.5-turbo and gpt-4o-mini language models. The results demonstrate a significant average improvement in key metrics over existing methods in the field, establishing DistillPrompt as one of the most effective non-gradient approaches in autoprompting.

851-859 4
Abstract

Real-time communication platforms generate continuous streams of short, latency-sensitive messages, creating a demanding environment for automated toxicity detection. Transformer-based language models offer strong contextual accuracy, but their inference cost makes them difficult to deploy efficiently in decentralized messaging systems such as the Extensible Messaging and Presence Protocol, where requests arrive asynchronously from multiple servers. Traditional batching strategies struggle in this setting because traffic patterns fluctuate, message lengths vary, and strict latency budgets prevent the accumulation of large batches. This work introduces a full-stack moderation system that combines an OpenFire server plugin, a Graphics Processing Unit (GPU) accelerated microservice for toxicity classification, and an adaptive batch processing method suitable for use in multi-server systems. The plugin intercepts messages at the packet-processing layer and forwards them to a lightweight external inference service that can run on either a Central Processing Unit or a GPU. We introduce Adaptive Cross-Domain Batching (ACDB) as a method to dynamically adjust batch sizes based on the state of the queues, characteristics of the messages being processed, and real-time feedback from GPU usage. Experiments demonstrate that GPU inference provides significant improvements in both throughput and latency, with compute accelerations of up to 28× and total end-to-end accelerations of up to 9×. Comparative evaluation against static batching baselines across multiple traffic scenarios shows that ACDB dynamically adapts batch sizes from 3 to 25 depending on load conditions, reducing latency by 52–59 % compared to large-batch static configurations while maintaining 85–94 % of their throughput. Overall, this system enables scalable, real-time toxic content moderation capabilities for federated chat systems using adaptive batching for Transformer-based inference on dynamic workloads, while preserving strict message delivery constraints under realistic operating conditions.

860-867 5
Abstract

 Analysis of document corpora with complex internal and inter-document links requires not only retrieval of relevant texts but also construction of a compact, verifiable and traceable set of fragments sufficient for downstream reasoning and citation. The purpose of this work is to propose a fragment retrieval method for document corpora in which the meaning of a fragment is determined by its local content, position in the document hierarchy and links to other fragments. The method is based on a joint representation of the corpus as a tree of structural units and a directed link graph as well as on hybrid ranking that combines lexical search, vector similarity, and a link signal. The monotonicity and submodularity of the objective function are shown, which makes it possible to use greedy algorithms with a known approximation guarantee and to perform budgeted context selection for a Retrieval-Augmented Generation (RAG) system. In addition, an evaluation protocol is introduced that separates retrieval quality at the document, fragment, and citation levels. The method is formally validated on tests of lexical mismatch robustness and budgeted selection efficiency compared with simple strategies. Examples from the legal domain are used to illustrate the method. The method can be used as a retrieval layer for RAG systems in question answering, evidence retrieval, regulatory compliance, and analysis of large structurally connected corpora.

MODELING AND SIMULATION

868-876 3
Abstract

Current trends in the development of enterprises are directly related to the active introduction of Information Technology (IT) in all areas of their operation. IT today is not only a tool for collecting, processing and analyzing data, but also a leading factor in ensuring competitive advantages. Most enterprises understand the need and expediency of forming, maintaining, and modernizing an IT infrastructure that includes hardware, software, and a network connecting them into a single system of interaction. Currently, a significant number of large and medium-sized Russian companies need to modernize their IT infrastructure. This is due to the processes of functioning of companies, changes in the volume of work, sales, supplies, the number of employees, interactions with related parties, etc. At the same time, the processes of modernization of the IT infrastructure significantly lag behind the dynamics of changes in the fortunes of companies, which, as a rule, leads to inevitable losses. Due to the wear and tear of the elements of the IT infrastructure, the costs of maintaining them in working order, monitoring their technical condition, maintenance, and scheduled and unplanned repairs increase significantly. In this case, the task arises of determining the timing of the modernization of the IT infrastructure in terms of replacing its elements. Traditionally, this task is solved with known parameters of IT infrastructure modernization such as the uptime allocation function of its elements. The paper proposes a model for determining the timing of modernization of the IT infrastructure of enterprises in conditions when the initial information on the operating time for failure is represented by small samples. The scientific novelty of the work consists in obtaining a general solution to the problem of determining the timing of IT infrastructure modernization in conditions of incomplete data. The optimal time for upgrading the IT infrastructure is determined from the condition of maximizing the mathematical expectation of profit from its operation. The mathematical expectation of profit from the functioning of the IT infrastructure is based on the formula of full mathematical expectation for a complete group of events: an event identical to the fact that an element failure occurred (will occur) after the time of element replacement and an event identical to the fact that an element failure occurred (will occur) before the replacement time. The ratios for the mathematical expectation of profit are obtained with an exponential distribution of the uptime of an element of the IT infrastructure, the optimal value of the element replacement period is found, and it is shown that this value of the replacement period provides, precisely, the maximum value of the average profit. A ratio is obtained for the value of the lower estimate of the mathematical expectation of profit on a set of distribution functions with specified moments equal to the moments determined based on the initial sample of uptime. The problem of determining the optimal replacement time for an element of the IT infrastructure is formulated as the task of finding the maximum minimum of the mathematical expectation of profit, which is reduced to the problem of nonlinear one-dimensional programming. The model performance is demonstrated by the example of determining the timing of disk upgrades in a disk array using the HP EVA P6500 disk array as an example, TOPAZ Ethernet Switches. The results obtained can be used by specialists in assessing and optimizing the time frame for the modernization of the IT infrastructure.

877-886 3
Abstract

Theoretical determination of planetary gear efficiency at design stage for given gear ratio allows for comparison of various designs and selection of most efficient one, providing required kinematic characteristics and minimal power losses. This article examines theoretical determination of planetary gear efficiency for electromechanical systems with distributed parameters for given kinematic design for installed driving and driven links as well as known engagement efficiency of reversing mechanism. Comprehensive analysis of algorithms developed for theoretically estimating planetary gear efficiency is presented. These systems are considered within framework of specific kinematic designs where both driving and driven links are clearly defined, and known engagement efficiency of reversing mechanism is taken into account. Main objective is to provide tools for more accurate and universal determination of planetary gear efficiency. Algorithms for theoretical determination of planetary gear efficiency are developed. Two universal relationships are proposed that allow for estimating energy efficiency of various planetary gear kinematic designs, taking into account functions of driving and driven links. Method enables theoretical evaluation of efficiency at design stage. It is shown that in planetary gearboxes with leading carrier, when selecting kinematic design and gear tooth count to ensure given gear ratio, it is necessary to strive for positive gear ratio. This, all other things being equal, ensures higher efficiency and lower load on gearbox housing. For cycloidal-pinion gearboxes with two-crown satellite with cycloidal tooth profile, central gear, secured in housing, must have greater number of pinions than central gear mounted on driven shaft. Method is applicable to planetary gearboxes and cycloidal-pinion gearboxes operating in reduction and multiplier modes. Obtained results contribute to optimization of gearbox designs and contribute to increased energy savings and improved mechanical efficiency in various applications.

BRIEF PAPERS

887-889 7
Abstract

This study investigates the relationship between a driver emotional state and vehicle maneuvers. Angular velocity was estimated from GPS trajectories using Euler angles, and maneuver boundaries were identified dynamically. Driver emotions were recognized from in-cabin video recordings. The analysis was conducted on data collected from 42 drivers, comprising more than 300 hours of driving. A three-phase representation of maneuvers is proposed in which emotional responses are evaluated before, during, and after each maneuver. Cluster analysis revealed several characteristic patterns of emotional responses, including increased anger during U-turns and a decrease in positive emotions while performing driving maneuvers.



Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.