Andrés Guzmán-Cordero

Machine learning researcher working on quantum many-body problems.

I am a PhD student in Computer Science at Mila, advised by Kirill Neklyudov. My research develops new machine learning methods for scalable ab initio electronic structure calculations. My current goal is to scale density functional theory to large biomolecules. Chemistry that is predictable at protein scale would translate into better drugs and designed catalysts.

Before Mila I did an MPhil at the University of Amsterdam with Jan-Willem van de Meent on variational inference in flow-based generative models, and wrote my thesis at the University of Toronto with Alán Aspuru-Guzik on Bayesian and manifold optimization. My BSc was at the Vrije Universiteit Amsterdam with Andre Lucas, on statistical models for particle tracking.

Recently

7 more

Research

2026

Preprint · 2026

Scaling Density Functional Theory with Gaussian Splatting

Andrés Guzmán-Cordero, Cindy Zhang, Majdi Hassan, Marta Skreta, Kirill Neklyudov, Matija Medvidović

  • GS-DFT replaces the fixed atom-centered basis with a cloud of Gaussians whose positions, shapes and mixing coefficients are optimized jointly by gradient descent to minimize the energy, using no training data. Conceptually it is 3D Gaussian splatting with the renderer replaced by quantum mechanics.
  • The optimized basis reaches the accuracy of the largest conventional basis sets at a fraction of the parameter count, converging systematically in energy, density and nuclear forces. At equal parameter count it captures the stretched-bond and anion physics that fixed bases recover only with specialized augmentation.
  • Adaptive density fitting with screening and a regularized differentiable orthogonalization give quadratic peak memory scaling in the cloud size, reaching 2,742 atoms and 10,406 electrons at triple-zeta scale on a single four-GPU node.
More
Density functional theory (DFT) strikes a practical balance between accuracy and computational cost in many problems of computational chemistry and materials science. However, many DFT calculations are limited by fixed atom-centered basis sets, which dictate how accuracy and cost scale with system size. We propose Gaussian Splatting for Density Functional Theory (GS-DFT), which represents molecular orbitals as a cloud of Gaussians whose positions, shapes, and mixing coefficients are optimized jointly by gradient descent to minimize the energy without training data. Conceptually, GS-DFT is 3D Gaussian splatting with the renderer replaced by quantum mechanics. We introduce two key solver components: adaptive density fitting with screening for efficient evaluation of two-electron integrals, and a regularized differentiable orthogonalization of the molecular orbitals. Empirically, the optimized basis reaches the accuracy of the largest conventional basis sets with a fraction of the parameters, converging systematically in energy, density, and nuclear forces. At equal parameter count, it captures the stretched-bond and anion physics that fixed bases only recover with specialized basis augmentation. The resulting solver exhibits quadratic peak memory scaling in the cloud size, allowing us to simulate systems of up to 2,742 atoms (10,406 electrons) without any modifications at triple-zeta scale using a single four-GPU node.

2025

NeurIPS · 2025

Improving Energy Natural Gradient Descent through Woodbury, Momentum, and Randomization

Andrés Guzmán-Cordero, Felix Dangel, Gil Goldshlager, Marius Zeinhofer

  • The ENGD curvature matrix is built from residual Jacobians, so the Woodbury formula rewrites the solve in sample space instead of parameter space. This reaches the same L² error as the original ENGD up to 75× faster.
  • Adapting SPRING (Subsampled Projected-Increment Natural Gradient Descent) from the variational Monte Carlo literature accelerates convergence and removes the line search that standard ENGD requires.
  • Randomized solvers help in the early stages of training on low-dimensional problems. We identify the barriers to obtaining the same acceleration in other settings.
More
Natural gradient methods significantly accelerate the training of Physics-Informed Neural Networks (PINNs), but are often prohibitively costly. We introduce a suite of techniques to improve the accuracy and efficiency of energy natural gradient descent (ENGD) for PINNs. First, we leverage the Woodbury formula to dramatically reduce the computational complexity of ENGD. Second, we adapt the Subsampled Projected-Increment Natural Gradient Descent algorithm from the variational Monte Carlo literature to accelerate the convergence. Third, we explore the use of randomized algorithms to further reduce the computational cost in the case of large batch sizes. We find that randomization accelerates progress in the early stages of training for low-dimensional problems, and we identify key barriers to attaining acceleration in other scenarios. Our numerical experiments demonstrate that our methods outperform previous approaches, achieving the same L² error as the original ENGD up to 75× faster.

ICML · 2025

Exponential Family Variational Flow Matching for Tabular Data Generation

Andrés Guzmán-Cordero, Floor Eijkelboom, Jan-Willem van de Meent

  • Exponential Family Variational Flow Matching represents mixed continuous and discrete features with a single exponential family, so tabular data needs no continuous relaxation of its discrete columns.
  • The resulting objective is moment matching, which we connect to generalized flow matching objectives based on Bregman divergences. TabbyFlow, built on this formulation, reaches state of the art on tabular benchmarks.
More
While denoising diffusion and flow matching have driven major advances in generative modeling, their application to tabular data remains limited, despite its ubiquity in real-world applications. To this end, we develop TabbyFlow, a variational Flow Matching (VFM) method for tabular data generation. To apply VFM to data with mixed continuous and discrete features, we introduce Exponential Family Variational Flow Matching (EF-VFM), which represents heterogeneous data types using a general exponential family distribution. We hereby obtain an efficient, data-driven objective based on moment matching, enabling principled learning of probability paths over mixed continuous and discrete variables. We also establish a connection between variational flow matching and generalized flow matching objectives based on Bregman divergences. Evaluation on tabular data benchmarks demonstrates state-of-the-art performance compared to baselines.

Faraday Discussions · 2025

Spiers Memorial Lecture: How to do impactful research in artificial intelligence for chemistry and materials science

Austin H. Cheng, Cher Tian Ser, Marta Skreta, Andrés Guzmán-Cordero, Luca Thiede, Andreas Burger, Abdulrahman Aldossary, Shi Xuan Leong, Sergio Pablo-García, Felix Strieth-Kalthoff, Alán Aspuru-Guzik

  • Sets out how machine learning researchers view and approach problems in chemistry and materials science.
  • Gives our considerations for maximizing impact when doing machine learning research for chemistry.
More
We discuss how machine learning researchers view and approach problems in chemistry and provide our considerations for maximizing impact when researching machine learning for chemistry.

2024

AI4Mat @ NeurIPS · 2024

Dimension Debate: Is 3D a Step Too Far for Optimizing Molecules?

Andrés Guzmán-Cordero, Luca A. Thiede, Gary Tom, Alán Aspuru-Guzik, Felix Strieth-Kalthoff, Agustinus Kristiadi

  • Compares 1D, 2D and 3D molecular representations inside a Bayesian optimization loop, scored by optimization progress rather than by predictive accuracy. 1D and 2D representations generally win.
  • 3D surrogates need on the order of 10,000 observations before closing the gap to 2D baselines, which exceeds the budget Bayesian optimization is meant to operate under. Transfer learning narrows the gap but does not reverse the ordering.
More
Molecules are three-dimensional objects, so it is tempting to assume that 3D representations should dominate 1D and 2D ones for molecular discovery. We test that assumption inside a Bayesian optimization loop, where what matters is not representation faithfulness but whether the induced posterior finds high-value molecules in few expensive evaluations. Comparing Gaussian processes against neural surrogates with a linearized Laplace approximation, and scoring by normalized GAP, we find that richer geometry adds cost faster than it adds decision-making power.