![]() |
|
The course consists of five days of lectures, exercises, and presentations. The lectures will cover theoretical and practical aspects of tensor networks for machine learning. Technical aspects, programming and application of the covered concepts will be explored in tutorials and exercises. The course (2.5 ECTS points) passed by handing in a small report covering topics in the course. For further course details click here.
DTU Compute building 324 room 060
Mahito Sugiyama, Associate professor, National Institute of Informatics, Japan.
Evrim Acar, Chief Research Scientist/Research Professor, Simula Metropolitan, Oslo, Norway.
Antonio Vergari , Associate Professor, University of Edinburgh, UK.
Beatriz Quintanilla Casas, Assistant Professor, Department of Food Science, Copenhagen University.
Michael Kastoryano , Associate Professor, Department of Computer-science, Copenhagen University.
Sebastian Loeschcke , PhD, Department of Computer-science, Copenhagen University.Monday, 24 August
To gain insights about complex systems such as the human brain and human metabolome, and to detect early risk factors for various diseases, we often face the challenge of analyzing longitudinal, multimodal data and revealing interpretable patterns. For instance, how can we analyze metabolomics measurements of blood samples collected over time during a meal challenge test and capture the underlying patterns which may reveal markers of various phenotypes? How to jointly analyze such data with other datasets such as other omics data? How to reveal subject-specific temporal profiles? How to inform our analysis with prior information encapsulated in computational models of human metabolism? Multivariate measurements collected from multiple subjects over time can often be represented as a third-order tensor. For instance, metabolomics measurements collected over time from multiple subjects can be arranged as a subjects by metabolites by time tensor or neuroimaging data can be represented as a third-order tensor with modes: subjects, voxels and time. Tensor factorizations have been successfully used to reveal the underlying patterns in such higher-order datasets in many domains, and have been extended to joint analysis of multimodal data through coupled tensor factorizations. In this talk, we will discuss (i) models and algorithms for (coupled) tensor factorizations, (ii) how we can use (coupled) tensor factorizations as a general framework to extract interpretable patterns from longitudinal and/or multimodal data, and (iii) how we can bring together data and computational models using a knowledge-guided approach based on coupled tensor factorizations. Throughout the talk, we will cover applications from different fields, in particular, focusing on metabolomics and neuroimaging data analysis.
Tuesday, 25 August
We will review some of the basic tensor network constructions in physics, and discuss the core problems that people are interested in, as well as the main algorithms used to solve these problems. As much as possible, we will try to connect the topics back to more familiar data science formulations.
Although tensor methods are now widely used in machine learning and signal processing, several of the most influential tensor decomposition models were originally developed in the psychometrics field, as they allowed analysing multidimensional behavioural data. Models such as Tucker and PARAFAC were subsequently adopted and further developed in chemometrics, where they became powerful tools for the analysis of complex multiway data.
This lecture presents an application-oriented overview of the tensor models most widely used in chemometrics, with a particular focus on PARAFAC and Tucker models, together with their variations. Common multiway data structures encountered in food science, including excitation-emission fluorescence measurements, chromatographic-mass spectrometric profiles, and sensory evaluation data, will be used to illustrate the strengths and limitations of tensor decomposition methods. The lecture will discuss how these models exploit the underlying multiway structure to obtain chemically and sensorially meaningful representations, with emphasis on model assumptions, uniqueness and interpretability.
Current challenges and ongoing research approaches will also be presented, including the analysis of increasingly complex analytical data, robustness and automation of tensor-based workflows.
Wednesday, 26 August
Modern deep learning models are increasingly limited by memory. In this lecture, I will discuss how tensor networks and low-rank methods can be used to reduce memory usage during training while preserving model quality. I will present three lines of work. First, I will cover tensor networks for compact visual representations, showing how coarse-to-fine learning of quantized tensor trains can be used to fit images, 3D signals, and neural fields efficiently. Second, I will discuss memory-efficient LLM pretraining, focusing on how low-rank structure in gradients can be exploited to reduce the memory cost of optimizer states and gradient updates during training. Finally, I will extend these ideas beyond matrices to tensor-valued gradients, presenting recent work on memory-efficient optimization for neural operators used to solve large and complex PDEs. The talk will emphasize both the underlying structured-learning ideas and their practical implications for training expressive models under strict memory constraints.
Thursday, 27 August
This lecture will bridge the two often separate communities of tensor factorizations and circuit representations, which investigate concepts that are intimately related. By connecting these fields, I will highlight a series of opportunities that can benefit both communities. We will draw theoretical as well as practical connections, e.g., in efficient probabilistic inference, reliable neuro-symbolic AI and scalable statistical modeling. The tutorial will start from classical tensor factorizations and extend them to a hierarchical setting, where the connection to circuit representations will be highlighted. Then, we will list several opportunities by bridging the two communities, such as using hierarchical tensor factorizations for neuro-symbolic inference or exploiting algorithms from the tensor network communities to learn circuits. Then, I will introduce a modular "Lego block" approach to build tensorized circuit architectures in a unified way. This, in turn, allows us to systematically construct and explore novel circuit and therefore tensor factorization architectures in a breeze while maintaining tractability. Lastly, we will showcase how one can understand the many recent algorithms and representations to learn circuits from data as hierarchical tensor factorizations. At the end of the lecture, the audience will learn about the state-of-the-art in representing, learning and scaling tensor factorizations and circuits.
How can we estimate the true distribution underlying observed data? This is a fundamental question in machine learning. In this lecture, I will show that tensor decomposition is particularly well suited to discrete (categorical) density estimation, because it can naturally exploit the discreteness of tensor indices. We will then discuss recent developments in tensor-based density estimation beyond the Kullback-Leibler (KL) divergence, with a focus on improving robustness to noise and outliers. Although the KL divergence is often easier to optimize than other divergences, it is known to be less robust to outliers and noise. To address this issue, we will see efficient closed-form optimization methods based on a doubly bounded EM algorithm, as well as a relaxation approach to density estimation that uses deformed algebra to flatten the feasible set, thereby enabling iterative convex optimization. We will also discuss the limitations of conventional low-rank modeling approaches and introduce tensor many-body decomposition as an alternative energy-based modeling for density estimation.
Friday, 28 August
I introduce an information-geometric framework for modeling and learning for tensors. In this framework, tensors are treated as discrete probability distributions on partially ordered sets (posets). By modeling tensors with log-linear models, they can be parameterized using the (θ, η)-coordinate system, which forms the dually flat manifold canonically used in information geometry. This geometric structure enables a variety of tensor-processing tasks, including tensor approximation and tensor balancing, to be formulated as projection problems onto constrained spaces. These operations naturally correspond to "learning" in machine learning and can be solved through convex optimization. I will present practical methodologies based on this framework, including many-body tensor approximation and data augmentation, and demonstrate its usefulness, flexibility, and broad applicability.
General understanding of machine learning, statistical modeling, mathematics and computer science. Programming experience, ideally in Python. For the course you are required to bring your own laptop computer.
To registrer please send a CV to kazfu@dtu.dk no later than the 27th of June 2026. Confirmations will be sent out by the 30th of June.
After receiving a positive confirmation on the application, registration in DTUs systems must be carried out. For academics (masters and PhD students) there is no registration fee for the course. Students affiliated with DTU can use the course planner to register. PhD students outside of DTU have register via here. For all other participants a course fee will be charged and apart from signing up additional registration must be completed here.
This years course is supported by the Novo Nordisk Foundation funded project "Machine Learning for Tensor Networks: Stability, Efficiency and Explainability at Scale" and the Danish Data Science Academy.
2025 version of 02901 Advanced Topics in Machine Learning
2024 version of 02901 Advanced Topics in Machine Learning
2023 version of 02901 Advanced Topics in Machine Learning
2022 version of 02901 Advanced Topics in Machine Learning
2021 version of 02901 Advanced Topics in Machine Learning
2020 version of 02901 Advanced Topics in Machine Learning
2019 version of 02901 Advanced Topics in Machine Learning 2018 version of 02901 Advanced Topics in Machine Learning2017 version of 02901 Advanced Topics in Machine Learning
2016 version of 02901 Advanced Topics in Machine Learning
2015 version of 02901 Advanced Topics in Machine Learning
2014 version of 02901 Advanced Topics in Machine Learning
For further information, please contact:
|