PhD Thesis defense by A. Alcalde

Next Monday October 26, 2026 our team member Albert Alcalde will present his PhD Thesis on:
“A Dynamical Perspective on Transformers, with Applications to Data-Driven Modeling”
Advisors:
• Prof. Giovanni Fantuzzi, FAU – Friedrich Alexander-Universität Erlangen-Nürnberg (Germany)
• Prof. Enrique Zuazua, FAU – Friedrich Alexander-Universität Erlangen-Nürnberg (Germany)

Abstract. Transformers form a central class of architectures in modern deep learning, where interpolating maps are built primarily through the alternation of simpler parametric functions known as self-attention and feed-forward layers. This thesis advances the study of these models from a dynamical perspective: inputs are interpreted as collections of particles, layers as time steps, self-attention as a nonlinear interaction rule, and feed-forward layers as pointwise forcing terms. The goal is to characterize rigorous asymptotic regimes of the forward dynamics of Transformers in terms of their parameters and to connect them with interpretable methods for data-driven modeling.

The first part of the thesis analyzes the key components of Transformer models and characterizes their dynamics in a range of parameter regimes. In a pure-attention model where softmax self-attention layers are replaced by the hardmax rule, the dynamics admit a clear geometric interpretation: particles or tokens converge to cluster points determined by special tokens called leaders. The same hardmax limit connects self-attention with the well-known Frank–Wolfe optimization scheme, in which a functional is minimized over a convex set, given here by the convex hull of the tokens. This connection yields convergence rates and relates the clustering results to the concentration phenomena observed empirically in softmax Transformers.

The thesis uses the resulting insight into the geometry of attention to study the expressive power of Transformers through exact sequence interpolation, where finite collections of input tokens must be driven exactly to prescribed target collections of output tokens. Specifically, it proves that both hardmax and softmax Transformers can exactly interpolate such datasets, with explicit bounds on the number of layers and the total parameter count. These results quantitatively characterize when sufficiently deep models can perfectly memorize a training dataset.

Motivated by the long input sequences processed by modern large language models, the second part of the thesis then studies the infinite-token limit of Transformers through a mean-field formulation. By considering the resulting continuity equation, which describes the evolution of token distributions across layers, it obtains quantitative concentration rates as a function of the model parameters. For Gaussian input distributions and affine feed-forward layers, the continuity equation reduces to a system of differential equations for the mean and covariance. The reachability and asymptotic dynamics of this reduced model are studied through their connection with Riccati-type dynamics. This analysis identifies parameter regimes in which the forward dynamics are well-posed in terms of the signs of the parameter matrices, a property relevant to the stability of deep models.

The third part of the thesis is devoted to applications in data-driven modeling, where the goal is to build evolution models from trajectory observations. Leveraging the theoretical study of Transformers, the thesis first proposes a minimal attention-based architecture for learning dynamics. Its minimality allows it to be interpreted as a nonlinear extension of the linear baseline of time-delayed dynamic mode decomposition. The proposed method shows improved performance on chaotic and weakly chaotic datasets, where its nonlinear structure provides additional predictive power compared with the linear baseline.

Finally, the thesis considers the more challenging problem of learning a model with an additional certificate of boundedness of the trajectories. This task is addressed by jointly learning polynomial models and a Lyapunov function that determines a trapping region. This results in a system identification method that provably discovers bounded polynomial ODEs and performs robustly on benchmark problems in ODE discovery.

WHEN
Mon. October 26, 2026 at 13:00H (Berlin time)

WHERE
On-site. FAU

Board of examiners

|| You might like: Upcoming events