Homepage Program Systems: Theory and Applications Русская версия
ISSN 2079-3316 Bilingual online scientific Online scientific journal of the Ailamazyan Program System Institute of the Ailamazyan PSI of PSI of Russian Academy of Science of RAS 12+ 
Volume 17 (2026) . Issue 3 (72) . Paper No. 9 (519)

Artificial intelligence and machine learning

Research Article

The effect of initial learning rate on properties of neural networks in automatic speech recognition

Ildus Rustemovich SadrtdinovCorrespondent author

HSE University, Moscow, Russia
Ildus Rustemovich Sadrtdinov — Correspondent author isadrtdinov@hse.ru

Abstract. The initial learning rate of stochastic gradient descent plays a decisive role in the quality of the resulting solution and in the structure of features learned by a neural network. Recent works on image classification have established a taxonomy of three regimes of the initial learning rate and identified a narrow range of high values that yields optimal quality after subsequent fine-tuning with a low learning rate or weight averaging.
This work extends this methodology to speech recognition, where the regime taxonomy is reproduced in its main features, and the identified range of high values retains both its superior final quality after fine-tuning and weight averaging and the linear mode connectivity of the resulting solutions in weight space. The picture of learned features is reproduced in a weaker form: as the learning rate grows, the models shift their reliance toward the lower cepstral coefficients and the low-frequency part of the mel spectrogram, yet within the high-value region this shift no longer singles out the range of best quality. In addition, truncation of the mel-frequency cepstrum is proposed as a portable tool for analyzing learned features, applicable to any model operating on spectrograms.
The obtained results show that the identified regime is a general property of stochastic optimization of neural networks rather than an artifact of the image classification task or of training on a sphere. Therefore, the identified regime is not an artifact of the image classification task and successfully transfers to speech recognition. (In Russian).

Keywords: learning rate, stochastic gradient descent, speech recognition, linear mode connectivity, feature learning

MSC-20202020 Mathematics Subject Classification 68T07; 68T10, 68T20MSC-2020 68-XX: Computer science
MSC-2020 68Txx: Artificial intelligence
MSC-2020 68T07: Artificial neural networks and deep learning
MSC-2020 68T10: Pattern recognition, speech recognition
MSC-2020 68T20: Problem solving in the context of artificial intelligence (heuristics, search strategies, etc.)

For citation: Ildus R. Sadrtdinov. The effect of initial learning rate on properties of neural networks in automatic speech recognition. Program Systems: Theory and Applications, 2026, 17:3, pp. 315–358. (In Russ.). https://psta.psiras.ru/2026/3_315-358.

Full text of article (PDF): https://psta.psiras.ru/read/psta2026_3_315-358.pdf.

The article was submitted 22.07.2026; approved after reviewing 26.08.2026; accepted for publication 12.09.2026; published online 20.09.2026.

© Sadrtdinov I. R.
2026
Editorial address: Ailamazyan Program Systems Institute of the Russian Academy of Sciences, Peter the First Street 4«a», Veskovo village, Pereslavl area, Yaroslavl region, 152021 Russia;   Website:  http://psta.psiras.ru Phone: +7(4852) 695-228;   E-mail: ;   License: CC-BY-4.0License text on the Creative Commons site
© Ailamazyan Program System Institute of Russian Academy of Science (site design) 2010–2026