Till innehåll på sidan
Till KTH:s startsida

Single-cell tumor phylogenetics

Probabilistic models and inference algorithms for tumor evolution from copy-number aberrations

Tid: Fr 2026-09-18 kl 15.00

Plats: Aula Buzano, Department of Mathematical Sciences, Politecnico di Torino, Corso Duca degli Abruzzi 24, Turin, Italy

Språk: Engelska

Ämnesområde: Datalogi

Respondent: Vittorio Zampinetti , Beräkningsvetenskap och beräkningsteknik, DISMA, Politecnico di Torino, Turin, Italy, Jens Lagergren

Opponent: Principal Research Fellow Simone Zaccaria, Princeton University, The Francis Crick Institute; University College London

Handledare: Professor Jens Lagergren, Beräkningsvetenskap och beräkningsteknik

Exportera till kalender

The thesis has been carried out under a co-tutelle agreement between KTH and Politecnico di Torino.

This work is licensed under a Creative Commons Attribution 4.0 International License( CC BY 4.0). You are free to share and adapt it for any purpose, provided that appropriate credit is given. The full license text is available at https://creativecommons.org/licenses/by/4.0/.The license above covers the introductory chapters of this thesis. The appended papers are included with permission from their respective copyright holders and remain subject to their own terms.

QC 20260824

Abstract

Tumors are highly heterogeneous populations of cells that evolve dynamically, acquiring mutations as they divide. Reconstructing this evolutionary history is essential for understanding cancer progression, metastasis, and therapy resistance. With the advent of single-cell DNA sequencing (scDNA-seq), we can now profile the genomic landscape of individual cells, which provides an unprecedented window into the formation and evolution of tumor cell populations. Specifically,single-cell whole-genome sequencing allows us to detect structural variations called \textit{copy number aberrations} (CNAs), which are known to play a critical role in cancer development and progression. Inferring the evolutionary trees, i.e., phylogenies, from single-cell DNA sequences presents immense computational challenges due to both the sparsity of the data inherent to the current state of sequencing technologies and the complexity of the underlying mutational processes.

In this thesis, we develop novel probabilistic models and algorithms to infer the evolutionary history of tumors from single-cell DNA sequencing data. The research addresses core methodological bottlenecks in tumor phylogenetics, moving from foundational distance-based methods to comprehensive joint Bayesian inference frameworks.

First, we introduce a method to estimate biologically meaningful evolutionary distances between single cells directly from noisy read counts, employing an original Hidden Markov Model to accommodate the unique noise profile of scDNA-seq and the interdependence of copy number states across the genome. Next, we extend distance-based tree inference from classical phylogenetics by presenting a scalable algorithm specifically designed for rooted trees, leveraging the biological premise that tumor evolution originates from a known healthy diploid ancestor. To enable rigorous uncertainty quantification over tree topologies, we then tackle the problem of sampling directed trees (arborescences). We present a stable, polynomial-time sampling algorithm capable of generating arborescences even on weakly connected graphs, which commonly arise when performing inference from single-cell sequences. Finally, we integrate these advancements into a comprehensive variational inference framework. This framework efficiently achieves joint inference over clonal tree structures, branch lengths, copy number profiles, and cell-to-clone assignments.

Collectively, this thesis contributes a suite of statistically grounded, highly scalable tools that bridge the gap between noisy single-cell sequencing reads and robust insights into cancer evolution, offering a foundation for future clinical applications and oncological research.

Link to DiVA