Portfolio
Research
Interpreting bias in genomics: high-occupancy target regions
We used machine-learning models - elastic-net regression and principal component analysis (PCA) - to investigate genomic regions known as high-occupancy target (HOT) regions. These regions attract unusually high numbers of proteins and are likely technical artifacts of chromatin immunoprecipitation followed by sequencing (ChIP-seq) experiments.
While antibody quality and chromatin interactions are known to affect ChIP-seq reliability, our study found that GC- and CpG-rich sequences, DNA methylation, and RNA:DNA hybrids (R-loops) also contribute to these artifacts across species. This work shows how machine learning can uncover hidden biases in genomic data and improve experimental interpretation.

Liquid-biopsy epigenetics in disease
DNA methylation biomarkers in acute coronary syndrome
We investigated circulating cell-free DNA (cfDNA) methylation as a non-invasive biomarker for acute coronary syndrome (ACS), based on the principle that damaged tissues release DNA into the bloodstream. Methylation profiles distinguished ACS subtypes and identified cell-type-specific markers that can help trace the tissue origin of cfDNA. Hundreds of markers linked to cardiovascular conditions and inflammation were validated in an independent cohort.

Publication: Cuadrat et al., NAR Genomics and Bioinformatics, 2023
DNA methylation profiling in neuroblastoma
In collaboration with Charité Hospital in Berlin, we analyzed primary neuroblastoma tumors and urine-derived cfDNA using bisulfite sequencing and RNA-seq. We identified methylation patterns that distinguish high- and low-risk tumors and linked MYCN-driven methylation changes to disrupted transcription-factor networks, highlighting potential therapeutic targets.

Open-source software

genomation
genomation is a Bioconductor R package for genomic feature and interval analysis. It provides tools to read BED/GFF files as GRanges, summarize genomic features across regions, create enrichment plots and heatmaps, and annotate regions with exons, introns, or promoters. I was a co-developer and maintainer.
GitHub repository · Developed with the MDC-BIMSB Bioinformatics and Omics Data Science Platform

PiGx
PiGx is a collection of genomics pipelines implemented with Snakemake, Python, and R. Pipelines are configured with a sample sheet and settings file, then generate interactive HTML reports summarizing results. My contribution included co-implementing the PiGx BS-seq pipeline.
GitHub repository · Developed with the MDC-BIMSB Bioinformatics and Omics Data Science Platform

motifActivity
motifActivity is an R package for identifying transcription factors that may explain changes in gene expression or epigenetic marks. It predicts transcription-factor activity from RNA-seq, BS-seq, ChIP-seq, ATAC-seq, and related data, together with DNA motif collections.
GitHub repository · Developed with the MDC-BIMSB Bioinformatics and Omics Data Science Platform
Freelance projects
Endometriosis prediction from scRNA-seq data
I built an single-cell RNA-seq workflow using menstrual-effluent samples to detect endometriosis through endometrium-atlas-based cell mapping and composition profiling. It identifies disease-associated shifts in stromal and immune populations and derives a sample-level similarity score.
- Processed data with nf-core/scrnaseq for standardized quality control, alignment, and count-matrix generation.
- Used AWS resources to scale preprocessing, reference mapping, and downstream modeling.
- Mapped donor cells to an endometrium reference atlas using scArches transfer learning, scvi-tools, and PyTorch.
- Performed analysis in Python with AnnData, Scanpy, scvi-tools, NumPy, pandas, Matplotlib, and seaborn.
- Derived a sample-level score to classify donors as control-like or endometriosis-like.

Prioritizing therapeutic targets in clinical trials
Biomarker visualization and survival analysis
We developed interactive visualizations, including oncoprints, to highlight biomarkers in patients with limited treatment options. These summaries reveal genomic alterations and support the identification of potential therapeutic targets. Survival analyses helped assess the clinical relevance of nominated targets in patient cohorts with poor outcomes or few effective therapies.

Machine learning for target identification
To prioritize therapeutic targets, we applied positive-unlabeled (PU) learning, which is useful when only confirmed targets are labeled. PU classifiers combined gene expression, mutation, and therapy annotations to identify potential targets. We also used autoencoders to uncover hidden patterns and prioritize molecular features without labels.


Multi-omics and Enformer for an Alzheimer’s disease biomarker
In this project, I investigated glial-to-neuron reprogramming through activation of a specific transcription factor (TF) in Alzheimer’s disease.
I integrated multi-omics data—including RNA-seq, ATAC-seq, and H3K4me2/H3K27me3 ChIP-seq—to identify differentially expressed genes and pathways (GO, GSEA), and examine their association with Alzheimer’s disease risk variants (GWAS), regulatory enhancers, DNA motifs, and TF binding sites linked to neuronal differentiation.
Large-scale analyses were executed via Nextflow pipelines on Kubernetes and AWS, ensuring scalable and reproducible processing of NGS datasets.
Using Enformer, I mapped the regulatory landscape around the target transcription factor and identified sequence regions predicted to influence its expression in neurons and glia. Enformer predicts regulatory activity from DNA sequence; gradient-based attribution highlights regions with strong effects on gene expression. I analyzed an approximately 400 kb region, applying cell-type-specific masks and signal smoothing to identify candidate enhancers.






IGV web-app feature development
I contributed to the IGV web application, an interactive tool for exploring genomic data (source code). The work used JavaScript and Python and enabled visualization of public and in-house datasets.
- Enabled dynamic visualization of new in-house genomic datasets.
- Added highlighting of genomic regions of interest, including variants.
- Added RefSeq and GENCODE transcript controls to collapse or expand isoforms, extend selected gene isoforms, and adjust track widths.
- Linked visualized tracks to their source databases.
- Implemented a command-line tool to generate snapshots of specified genes or regions.

Proteomic signatures for Crohn’s disease
In this project, I worked with biobank data from IBD Plexus, a resource of the Crohn’s & Colitis Foundation, to explore machine-learning approaches for identifying proteomic signatures to support blood-based testing for Crohn’s disease. I also developed AWS infrastructure to support the associated data workflows.


Independent projects
Machine learning and AI
End-to-end machine-learning APIs
Gene-type prediction from DNA sequence
Built classical machine-learning models, convolutional networks, and a custom nucleotide-level Transformer to predict gene type from raw DNA sequence. The deployment work explored ONNX Runtime inference, FastAPI and BentoML serving, Docker Compose, Kubernetes with Kind, and AWS EKS.

Molecular solubility prediction
Built a Flask API to predict whether a compound dissolves in water from molecular descriptors such as size, polarity, solvation energy, and charge distribution. I compared Partial Least Squares, Elastic Net, Random Forest, and XGBoost models and deployed the service to AWS Elastic Beanstalk.

Immune-cell image classifier
Evaluated transfer learning for classifying immune-cell types in H&E-stained blood microscopy images. A pretrained Xception model served as a fixed feature extractor with a custom MLP classification head. I also trained logistic regression and XGBoost on Xception embeddings as interpretable baselines, tuning learning rate and dropout; the MLP-based model performed strongly on held-out data. The API was deployed with Docker and AWS Lambda.

Diffusion-based image inpainting
Explored diffusion models for reconstructing missing regions in H&E-stained blood-cell images. The compact diffusion model restores masked areas while preserving observed context and is presented through a simple Streamlit interface, with AWS Batch used for deployment.

Survival analysis using gene expression & clinical data (Cox models)
Developed models to predict mortality or relapse risk in newly diagnosed multiple-myeloma patients using baseline clinical and gene-expression data. The workflow included RNA-seq preprocessing, PCA and clustering, Cox regression, random survival forests, LASSO-based feature selection, and pathway-informed models, evaluated using the concordance index (C-index).

Deep learning for imaging and omics
CNNs and transfer learning for chest X-ray classification
I applied convolutional neural networks (CNNs) to classify chest X-ray images using both 224×224 and 64×64 pixel inputs, exploring whether lightweight models can retain diagnostic performance. Alongside a baseline CNN trained from scratch, I used transfer learning with pretrained convolutional backbones such as ResNet to assess whether pretraining could improve classification.


Autoencoder for scRNA-seq dimensionality reduction and data imputation
I developed a simple autoencoder with a custom loss function for imputing missing values in single-cell RNA-seq data. The approach was inspired by the method proposed by Badsha et al.


Federated VAE for scRNA-seq batch correction
Explored federated training of a scVI model using the Flower framework and SecAgg+ secure aggregation, and compared it with centralized training to study mitigation of batch effects in single-cell RNA-seq.



cfDNA cell-type deconvolution
Studied cfDNA fragments released by tissues into the blood as a potential source of early disease signals. Applied regression-based methods (NNLS, Lasso, Ridge, Elastic Net) to estimate cell-type proportions from bulk DNA methylation data, and developed:
- A variational autoencoder (VAE) that reconstructs CpG profiles while jointly predicting cell-type proportions.
- A semi-supervised NMF model anchored to known reference signatures.
- A lightweight transformer that treats CpG regions as tokens and uses embeddings and self-attention to capture genomic dependencies.

GNN for spatial transcriptomics
This project demonstrates how graph neural networks (GNNs) can capture spatially coherent patterns in gene expression, comparing their learned embeddings with traditional PCA and k-means clustering.
Spatial transcriptomics measures gene expression while preserving tissue architecture, enabling the study of cellular organization and microenvironments. Identifying spatial domains—regions with similar expression and spatial context—remains challenging.
Because GNNs model both gene-expression features and spatial neighborhood relationships, I implemented a mini graph autoencoder (GAE) from scratch in PyTorch to learn unsupervised embeddings of tissue spots. The project uses Squidpy’s toy 10x Genomics Visium H&E mouse-brain dataset (about 2,700 spots and 33,000 genes).


LLM assistant for bioinformatics queries
Developed an AI assistant that translates plain-English biology questions into SPARQL queries against UniProt (proteins and annotations), OMA (orthology), and Bgee (gene expression). The assistant combines Mistral or Llama models, available through Groq or Ollama, with retrieval-augmented generation using Qdrant and FastEmbed. Researchers can use a command-line interface or a Chainlit web app to validate and run queries and receive plain-language summaries.

Computational biology and algorithms
Protein folding with the HP model and replica Monte Carlo
Implemented simulated annealing and replica-exchange Monte Carlo in Python and NumPy for protein folding in the hydrophobic-polar (HP) lattice model. The model represents amino acids as hydrophobic (H) or polar (P) residues on a square lattice; Metropolis–Hastings sampling explores conformations according to the Boltzmann distribution.

Genome assembly with de Bruijn graphs
Implemented de Bruijn graph-based genome assembly using an Eulerian walk to reconstruct DNA sequences from k-mers. Nodes represent k-mer prefixes and suffixes; edges represent the k-mers. Finding an Eulerian cycle reconstructs a genome by joining successive k-mers with a one-base shift, avoiding the cost of searching for a Hamiltonian cycle.


Phylogenetic tree estimation with Felsenstein pruning and NNI
Implemented Felsenstein’s tree-pruning algorithm for evaluating evolutionary-tree likelihoods from nucleic-acid sequences, together with nearest-neighbor interchange (NNI) for rearranging rooted binary phylogenetic trees under the Jukes–Cantor substitution model.

Regulatory DNA discovery with comparative genomics
Bio Motif Ensembl is a Python tool for discovering candidate regulatory DNA regions across related mammalian genomes using Ensembl’s public MySQL databases. It retrieves orthologous sequences (for example, human, mouse, and rat), aligns upstream regions, detects conserved non-coding segments, and analyzes them with motif-discovery tools such as MEME and AlignACE. A binomial model tests motif over-representation to identify potentially functional regulatory elements.

Data engineering
Streaming analytics for urban bike sharing
This project ingests GBFS bike-station data every minute using Kestra, writes raw events to MinIO and Kafka, loads curated records into PostgreSQL, transforms them with dbt, and serves a two-tile Streamlit dashboard in Docker Compose. AWS infrastructure and CI/CD are provisioned with Terraform.

Web projects
Spatial transcriptomics explorer
An interactive platform for exploring gene expression in tissue context. It includes a chat interface, AI-powered gene summaries, and spatial-domain analysis through an MCP server.
- Gene summaries use Ollama, with Mistral 7B as the default model.
- The chatbox uses the same LLM backend.
- Spatial-domain analysis uses ChatSpatial MCP with SpaGCN when available, and falls back to Scanpy.

Sudoku
A simple Sudoku game implemented in JavaScript and jQuery.

Minesweeper
A classic Minesweeper game implemented in Java using Swing and AWT.

Django web services
Django-based server for visualizing multiple sequence alignments.
MSA visualization project on GitHub
Mobile application built with Django, a manifesto app, and localStorage.
