Kalman filtering offers a principled way to incorporate uncertainty and second-order information into neural network optimization, but directly maintaining the full covariance matrix is computationally prohibitive. In this talk, I will present the progression from KOALA to KOALA++ and KOALA+s, which progressively improve covariance modeling while retaining linear memory complexity. I will also discuss a recent theoretical connection showing that, under a particular covariance setting, KOALA++ admits an exact invariant that reduces its update to a scaled SGD step. This provides a new perspective on the relationship between Kalman-style optimization and classical gradient methods.
Bio
Zixuan Xia received his B.Eng. in Software Engineering from Xi’an Jiaotong University in 2024 and his M.Sc. in Computer Science from the University of Bern in 2026. During his master’s studies, he worked on neural network optimization in the Computer Vision Group with Prof. Paolo Favaro and Dr. Aram Davtyan, focusing on Kalman-filter-based optimization for deep learning. In October 2026, he will join Prof. Lydia Y. Chen’s Distributed Machine Learning Systems Lab at the University of Neuchâtel and TU Delft as a PhD student. His research interests include deep learning optimization, multimodal representation learning, and trustworthy generative AI.