Artificial intelligence and machine learning
Research Article
The effect of initial learning rate on properties of neural networks in automatic speech recognition
Ildus Rustemovich Sadrtdinov
| HSE University, Moscow, Russia | |
|
|
Abstract.
The initial learning rate of stochastic gradient descent plays a decisive role in the quality of the resulting solution and in the structure of features learned by a neural network.
Recent works on image classification have established a taxonomy of three regimes of the initial learning rate and identified a narrow range of high values that yields optimal quality after subsequent fine-tuning with a low learning rate or weight averaging.
This work extends this methodology to speech recognition, where the regime taxonomy is reproduced in its main features, and the identified range of high values retains both its superior final quality after fine-tuning and weight averaging and the linear mode connectivity of the resulting solutions in weight space.
The picture of learned features is reproduced in a weaker form: as the learning rate grows, the models shift their reliance toward the lower cepstral coefficients and the low-frequency part of the mel spectrogram, yet within the high-value region this shift no longer singles out the range of best quality.
In addition, truncation of the mel-frequency cepstrum is proposed as a portable tool for analyzing learned features, applicable to any model operating on spectrograms.
The obtained results show that the identified regime is a general property of stochastic optimization of neural networks rather than an artifact of the image classification task or of training on a sphere.
Therefore, the identified regime is not an artifact of the image classification task and successfully transfers to speech recognition. (In Russian).
Keywords: learning rate, stochastic gradient descent, speech recognition, linear mode connectivity, feature learning
MSC-2020
68T07; 68T10, 68T20For citation: Ildus R. Sadrtdinov. The effect of initial learning rate on properties of neural networks in automatic speech recognition. Program Systems: Theory and Applications, 2026, 17:3, pp. 315–358. (In Russ.). https://psta.psiras.ru/2026/3_315-358.
Full text of article (PDF): https://psta.psiras.ru/read/psta2026_3_315-358.pdf.
The article was submitted 22.07.2026; approved after reviewing 26.08.2026; accepted for publication 12.09.2026; published online 20.09.2026.