Master's project or research visits

April 1, 2026

We offer Master’s project opportunities in our lab for students admitted to a Master’s program at Chalmers University of Technology or the University of Gothenburg. On some occasions, our lab accepts candidates from other universities or other levels (internships, Erasmus, and PhD research visitors), which may be remote or in person. You are welcome to contact us directly at sina.majidian[at]chalmers.se to discuss opportunities further.

Students based at Chalmers or University of Gothenburg can check out these projects at Chalmers portal here.

Project 1: Computational methods for studying gene evolution using AlphaFold structures

One major challenge in evolutionary genomics is identifying relationships between genes across different species. Genes that descend from a common ancestral gene through speciation are called orthologs and often retain similar biological functions, while genes that arise through duplication events are known as paralogs and may evolve distinct or specialized functions. Correctly distinguishing between orthologs and paralogs is essential for transferring biological knowledge from model organisms to human. We previously developed a scalable orthology inference framework using protein sequences. While sequence-based methods are effective for closely related species, protein structures are more useful for highly divergent organisms, due to their higher conservation. Recent advances such as AlphaFold have revolutionized structural biology by making high-quality protein structure predictions available at scale. In this project, we will explore how protein structural information can improve orthology prediction across distant species. We leverage AI-powered predicted protein structure to improve upon state-of-the-art methods. Depending on the student’s interests and background, the work may involve machine learning, structural bioinformatics, or large-scale computational analysis.

  1. Majidian, S., et. al. (2025). Orthology inference at scale with FastOMA. Nature Methods, 22. doi.org/10.1038/s41592-024-02552-8
  2. Nevers, Y., et. al (2020). Orthology: promises and challenges. doi.org/10.1007/978-3-030-57246-4_9
  3. Jumper, J., et. al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature 596. doi.org/10.1038/s41586-021-03819-2
  4. Gilchrist, C. L., et. al. (2026). Multiple protein structure alignment at scale with FoldMason. Science, 391. doi.org/10.1126/science.ads6733
Master Project on protein structure 

Project 2: Tokenization techniques for DNA language model

Natural language processing techniques have been recently used to model DNA sequences as a biological “language”. The question is how to define “words” in DNA, since DNA sequences appear simply as strings such as “ATCGATCGATCGATCGATCG.” The idea is to apply deep learning models on small chunks of DNA to learn its patterns. One important development is DNABERT, which adapts the transformer architecture to DNA by representing sequences as fixed-length k-mers and learning contextual embeddings for downstream tasks such as predicting different DNA regions. However, k-mer tokenization is rigid and may not fully capture variable-length biological motifs or efficiently represent repetitive genomic structures. This limitation has motivated the exploration of more adaptive tokenization strategies, such as Byte Pair Encoding (BPE). In this project, we develop new tokenization strategies that can dynamically learn patterns and potentially provide more compact and informative representations of DNA data for improved modeling performance.

  1. Zhou, Z., et al. (2024). DNABERT-2: Efficient foundation model and benchmark for multi-species genomes. ICLR.
  2. Dilip, R., et al. (2026). Adaptive Protein Tokenization. arXiv:2602.06418.
  3. Lindsey, L. et al. (2025) The impact of tokenizer selection in genomic language models. Bioinformatics. doi.org/10.1093/bioinformatics/btaf456 Bioinformatic.s
  4. Avsec, Z, et al. (2026). Advancing regulatory variant effect prediction with AlphaGenome. Nature, 649.

Project 3: Mathematical analysis of phylogenetic gene trees

In this project, we aim to develop a graph-based framework for combining multiple partially overlapping gene trees into a more complete phylogenetic tree with branch lengths. Unlike classical consensus tree methods that assume identical leaf sets across all input trees, this project addresses a more realistic and challenging setting where several gene trees contain different taxa with only partial overlap. In a targeted scenario, the dataset consists of 10 gene trees with approximately 200 leaves each and around 30% overlap of leaves across trees, while the topology of output tree is assumed to be fixed. The goal is to estimate branch length. The proposed method will extend graph-based consensus approaches by integrating topological compatibility, clade frequency, edge support, and branch-length information to infer a unified tree representation. The project will investigate how overlapping taxa can guide the reconciliation and alignment of gene trees while preserving meaningful evolutionary distance information encoded in branch lengths. The resulting framework is expected to provide a more robust and biologically informative reconstruction method for large-scale phylogenetic datasets with incomplete taxon overlap. Final output includes a software package and theoretical bounds on the estimation error. The student should have a strong background in computer science and mathematics, with interest in algorithms and graph theory.

  1. Makarenkov, V., et. al. (2023). Inferring multiple consensus trees and supertrees using clustering: A review. doi.org/10.1007/978-3-031-31654-8_13
  2. Torquet, E., et. al. (2025). Graph-based method for constructing consensus . EDP Sciences. doi.org/10.1051/bioconf/202516301004
  3. Xia, X. (2018). Imputing missing distances in molecular phylogenetics. PeerJ, doi.org/10.7717/peerj.5321

Project 4: Multi-k-mer metagenomics clasiisfcaition for gut health Metagenomics studies genetic material from environments like gut microbiota or soil. High-throughput long-read sequencing enables large-scale metagenomic studies with high accuracy. This allows accurate characterization of microbial biodiversity and identification of bacterial species in a sample, enabling the assignment of taxonomy to each metagenomic read. This involves comparing DNA sequences from the sample against a reference database of tens of thousands of genomes. However, indexing and querying vast microbial genome collections present significant computational challenges. A common strategy involves breaking genomes into fixed-length sequences, k-mers. Kraken uses LCA (Lowest Common Ancestor) to assign reads to the deepest shared node in a taxonomic graph, based on their matching k-mers. In our previous work, the EvANI benchmarking pipeline revealed that different k-mer lengths can have complementary strengths for estimating evolutionary distances. Building on this, we are interested in investigating the effectiveness of using multiple k-mer lengths for metagenomic classification using machine learning approaches.

  1. Lu, Jennifer, et al. “Metagenome analysis using the Kraken software suite.” Nature protocols 17.12 (2022): 2815-2839.
  2. Wood, D, et al. “Kraken: ultrafast metagenomic sequence classification using exact alignments.” Genome biology 15.3 (2014): R46.
  3. Majidian, S, et al. “EvANI benchmarking workflow for evolutionary distance estimation.” Briefings in Bioinformatics 26.3 (2025): bbaf267.

Project 5: Impact of NCBI Database Growth on Metagenomics Association Studies

Metagenomics enables the study of microbial communities by sequencing genetic material directly from complex environments such as the human gut and soil. High-throughput sequencing has made it possible to characterize microbial diversity at large scale and to assign taxonomic identities to sequencing reads. A major objective of quantitative microbiome analysis is to identify microorganisms or community features associated with metadata such as human age, health, diet, or environmental conditions. Several statistical methods have been developed for differential abundance analysis, including ANCOM-BC (Analysis of Compositions of Microbiomes with Bias Correction) and MaAsLin (Microbiome Multivariable Associations with Linear Models). MaAsLin uses several statistical models, such as log-transformed linear models, count models (Negative Binomial), zero-adjusted models (Compound Poisson), and Zero-Inflated Negative Binomial (ZINB). However, there is still limited consensus on how best to perform differential abundance testing for microbiome data, particularly because of sparsity, zero observations, and potential confounding factors. One challenge that is less considered is the classification itself. One study of metagenomes has shown that different classifiers (MetaPhlAn4 and Kraken2) can produce different estimates of beta diversity and different age-associated taxa. Even when using the same classification method, changes in the reference database can affect classification. One study shows that when new species are added to the NCBI RefSeq database, more reads are classified, but fewer are classified at the species level. This project will investigate how metagenomic taxonomic classification and statistical methodology influence differential abundance results. We will explore which statistical models are less affected by changes in the reference database and investigate ways to mitigate these effects to achieve more robust statistical analysis.

  1. Nickols, William A., et al. “MaAsLin 3: refining and extending generalized multivariable linear models for meta-omic association discovery.” Nature Methods 23.3 (2026): 554-564.
  2. Mallick, Himel, et al. “Multivariable association discovery in population-scale meta-omics studies.” PLoS computational biology 17.11 (2021): e1009442.
  3. Nasko, Daniel J., et al. “RefSeq database growth influences the accuracy of k-mer-based lowest common ancestor species identification.” Genome biology 19.1 (2018): 165.
  4. Karagiannis, Tanya T., et al. “Integrative analysis across metagenomic taxonomic classifiers: A case study of the gut microbiome in aging and longevity in the Integrative Longevity Omics Study.” PLOS Computational Biology 22.1 (2026): e1013883.

Project 6: Vision foundation models for identifying genetic variations

Humans share most of their DNA, but small genetic differences contribute to variation in traits and susceptibility to disease. DNA can be extracted from human samples and sequenced to identify genetic variations that may contribute to genetic disorders. Short DNA reads are mapped to a reference genome, allowing variants such as single-nucleotide variants (SNVs) and small insertions and deletions to be identified. DeepVariant and ARCLID are deep-learning-based variant callers that represent sequencing data as images of aligned reads, which are then analyzed by neural networks to distinguish true genetic variants from sequencing errors. In this project, we will fine-tune DINOv3, a general-purpose vision foundation model, for genetic variant calling and investigate its ability to learn visual patterns in sequencing data for accurate variant detection.

  1. Poplin, Ryan, et al. “A universal SNP and small-indel variant caller using deep neural networks.” Nature biotechnology 36.10 (2018): 983-987.
  2. Tavakoli, Sajad, et al. “Accurate and Robust Characterization of Structural Variants at Low Coverages with ARCLID.” bioRxiv (2025): 2025-10.
  3. Sedlazeck, Fritz J., et al. “Piercing the dark matter: bioinformatics of long-range sequencing and mapping.” Nature Reviews Genetics 19.6 (2018): 329-346.
  4. Siméoni, Oriane, et al. “Dinov3.” arXiv preprint arXiv:2508.10104 (2025).
Master Project on genetic variations