Independent projects

Machine learning and AI

End-to-end machine-learning API projects

Gene-type prediction from DNA sequence

Built classical machine-learning models, convolutional networks, and a custom nucleotide-level transformer to predict gene type from raw DNA sequence. The deployment work explored ONNX Runtime inference, FastAPI and BentoML serving, Docker Compose, Kubernetes with Kind, and AWS EKS.

Gene and nucleotide sequence example
Example gene and its nucleotide sequence.

Molecular solubility prediction

Built a Flask API to predict whether a compound dissolves in water from molecular descriptors such as size, polarity, solvation energy, and charge distribution. I compared Partial Least Squares, Elastic Net, Random Forest, and XGBoost models and deployed the service to AWS Elastic Beanstalk.

Conceptual illustration of molecular solubility prediction
Conceptual view of molecular solubility prediction.

Immune-cell image classifier

Evaluated transfer learning for classifying immune-cell types in H&E-stained blood microscopy images. A pretrained Xception model served as a fixed feature extractor with a custom MLP classification head. I also trained logistic regression and XGBoost on Xception embeddings as interpretable baselines, tuning learning rate and dropout; the MLP-based model performed strongly on held-out data. The API was deployed with Docker and AWS Lambda.

Examples of blood-cell types
Examples of blood-cell types: erythroblast, monocyte, and platelet.

Diffusion-based image inpainting

Explored diffusion models for reconstructing missing regions in H&E-stained blood-cell images. The compact diffusion model restores masked areas while preserving observed context and is presented through a simple Streamlit interface, with AWS Batch used for deployment.

Diffusion-based inpainting results
Inpainting results: reconstructing masked regions while retaining surrounding image context.

Survival analysis using gene expression & clinical data (Cox models)

Developed models to predict mortality or relapse risk in newly diagnosed multiple-myeloma patients using baseline clinical and gene-expression data. The workflow included RNA-seq preprocessing, PCA and clustering, Cox regression, random survival forests, LASSO-based feature selection, and pathway-informed models, evaluated using the concordance index (C-index).

Comparison of survival models by C-index Kaplan–Meier plot for the best-performing survival model
C-index comparison of multiple survival models (left) and Kaplan–Meier plot of the best-performing model (right).

Deep learning for imaging and omics

CNNs and transfer learning for chest X-ray classification

I applied convolutional neural networks (CNNs) to classify chest X-ray images using both 224×224 and 64×64 pixel inputs, exploring whether lightweight models can retain diagnostic performance. Alongside a baseline CNN trained from scratch, I used transfer learning with pretrained convolutional backbones such as ResNet to assess whether pretraining could improve classification.

Example healthy chest X-ray
Healthy example.
Example chest X-ray showing pneumonia
Pneumonia example.

Autoencoder for scRNA-seq dimensionality reduction and data imputation

I developed a simple autoencoder with a custom loss function for imputing missing values in single-cell RNA-seq data. The approach was inspired by the method proposed by Badsha et al.

Imputed single-cell RNA-seq data
Imputed scRNA-seq data.
Observed and imputed gene expression values
Model output compared with true gene-expression values. Non-imputed data (blue) show reconstructed known values; imputed data (orange) show predictions for missing (masked) values.

Federated VAE for scRNA-seq batch correction

Explored federated training of a scVI model using the Flower framework and SecAgg+ secure aggregation, and compared it with centralized training to study mitigation of batch effects in single-cell RNA-seq.

Baseline gene-expression UMAP
Baseline gene-expression UMAP.
UMAP after centralized scVI batch correction
Centralized scVI model.
UMAP after federated scVI batch correction
Federated scVI model.

cfDNA cell-type deconvolution

Studied cfDNA fragments released by tissues into the blood as a potential source of early disease signals. Applied regression-based methods (NNLS, Lasso, Ridge, Elastic Net) to estimate cell-type proportions from bulk DNA methylation data, and developed:

  • A variational autoencoder (VAE) that reconstructs CpG profiles while jointly predicting cell-type proportions.
  • A semi-supervised NMF model anchored to known reference signatures.
  • A lightweight transformer that treats CpG regions as tokens and uses embeddings and self-attention to capture genomic dependencies.
Cell-type deconvolution of blood DNA methylation data
Deconvolution of blood DNA methylation measured by bisulfite sequencing.

Computational biology and algorithms

Protein folding with the HP model and replica Monte Carlo

Implemented simulated annealing and replica-exchange Monte Carlo in Python and NumPy for protein folding in the hydrophobic-polar (HP) lattice model. The model represents amino acids as hydrophobic (H) or polar (P) residues on a square lattice; Metropolis–Hastings sampling explores conformations according to the Boltzmann distribution.

Two-dimensional HP model protein-folding schematic
Example HP-model conformation with optimal energy −2; the two hydrophobic contacts are between residues 4 and 13, and 5 and 12 (Thachuk et al., 2007).

Genome assembly with de Bruijn graphs

Implemented de Bruijn graph-based genome assembly using an Eulerian walk to reconstruct DNA sequences from k-mers. Nodes represent k-mer prefixes and suffixes; edges represent the k-mers. Finding an Eulerian cycle reconstructs a genome by joining successive k-mers with a one-base shift, avoiding the cost of searching for a Hamiltonian cycle.

De Bruijn graph built from a sequence containing repeats
Example de Bruijn graph from a sequence with repeats (Compeau et al., 2011).
Comparison of Hamiltonian and Eulerian genome-assembly strategies
Genome assembly strategies, from Hamiltonian to Eulerian cycles (Pevzner et al., 2001); focus on the Eulerian-cycle approach.

Phylogenetic tree estimation with Felsenstein pruning and NNI

Implemented Felsenstein’s tree-pruning algorithm for evaluating evolutionary-tree likelihoods from nucleic-acid sequences, together with nearest-neighbor interchange (NNI) for rearranging rooted binary phylogenetic trees under the Jukes–Cantor substitution model.

Nearest-neighbor interchange on a subsplit directed acyclic graph
Example NNI on a subsplit directed acyclic graph.

Regulatory DNA discovery with comparative genomics

Bio Motif Ensembl is a Python tool for discovering candidate regulatory DNA regions across related mammalian genomes using Ensembl’s public MySQL databases. It retrieves orthologous sequences (for example, human, mouse, and rat), aligns upstream regions, detects conserved non-coding segments, and analyzes them with motif-discovery tools such as MEME and AlignACE. A binomial model tests motif over-representation to identify potentially functional regulatory elements.

Over-represented motif discovery concept
Concept related to comparative regulatory-motif discovery.

Data engineering

Streaming analytics for urban bike sharing

This project ingests GBFS bike-station data every minute using Kestra, writes raw events to MinIO and Kafka, loads curated records into PostgreSQL, transforms them with dbt, and serves a two-tile Streamlit dashboard in Docker Compose. AWS infrastructure and CI/CD are provisioned with Terraform.

Bike-sharing analytics dashboard
Dashboard running on an AWS EC2 instance provisioned with Terraform.

Web projects

Spatial transcriptomics explorer

An interactive platform for exploring gene expression in tissue context. It includes a chat interface, AI-powered gene summaries, and spatial-domain analysis through an MCP server.

  • Gene summaries use Ollama, with Mistral 7B as the default model.
  • The chatbox uses the same LLM backend.
  • Spatial-domain analysis uses ChatSpatial MCP with SpaGCN when available, and falls back to Scanpy.
Spatial transcriptomics app preview
Explore precomputed 10x Genomics Visium datasets directly in tissue context, without running computationally intensive pipelines online.

Sudoku

A simple Sudoku game implemented in JavaScript and jQuery.

Sudoku game screenshot

Minesweeper

A classic Minesweeper game implemented in Java using Swing and AWT.

Minesweeper game screenshot

Django web services

Django-based server for visualizing multiple sequence alignments.

Mobile application built with Django, a manifesto app, and localStorage.