My research focuses on the working principles of Generative AI: how can we infer and, at the same time, sample from, the distribution governing real data? Concretely, I study how Diffusion Models and Normalizing Flows learn from finite data resources. I also have a background in statistical physics and have worked on problems in network science and the spread of disease. A full list of my publications can be found here. Below is a selection of recent work.
Recent work
- Local Coverage Governs Memorization in Diffusion Models, Claudia Merger & Sebastian Goldt. Even diffusion models which produce novel examples can still memorize training data. Can we predict which data they memorize? We argue that memorization is strongly shaped by local coverage - i.e. whether a training datapoint lies in a high-density region or in a sparse, isolated one. Using the analogy between diffusion models and kernel density estimation, we show that datapoints from low coverage regions are much more likely to be memorized. For example: training a diffusion model on CIFAR-10, we find that cars are memorized much more often than ships, because in this dataset, car images are much more diverse and locally sparse, making interpolation harder.
In conclusion, memorization is more complicated than a global property of the model. Rather, it is a local phenomenon, controlled by the geometry of the data distribution.
- Generalization Dynamics of Linear Diffusion models, Claudia Merger & Sebastian Goldt. We investigate how the covariance statistics typically found in image data affect learning in diffusion models: Images data usually have a power-law structure in their covariance matrices, with directions of high variance accounting for backround color/ shadows and lower variance directions for finer details. We study when the diffusion model has enough data to accurately fit both high-variance and low variance directions and where and how it overfits to the data when not enough training data is available.
- A theory of learning data statistics in diffusion models, from easy to hard,
Lorenzo Bardone, Claudia Merger, Sebastian Goldt. We demonstrate that diffusion models sequentially learn first simpler, then more complicated statistics and predict at which sample complexity statistics of different orders are learned. We show that the number of data needed to learn higher order statistics with SGD increases polynomially with the dimension, unless these higher order statistics are correlated with relevant lower order statistics (for example, a high variance direction in the data covariance).
- Learning Interacting Theories from Data, Claudia Merger, Alexandre René, et al., Phys. Rev. X 13, 041033. We developed a new method for statistical inference using a trained generative AI model (a Normalizing Flow): We show that there is an exact mapping between the parameters of the Normalizing FLOW and a physical theory of interactions that coordinate degrees of freedom in a system. For images, this constitutes understanding how pixels interact to give rise to images. We also empirically demonstrated that in the case of insufficient model capacity, Normalizing Flows learn effective statistics of a dataset based on lower order cumulants.
c lastname [AT] sissa [dot] it