Senior Data Scientist
Data Scientist (8+ years) working at the intersection of statistical genetics, population genomics, and applied machine learning in biomedical research. MEng in Big Data Analytics, PhD in Computational Physics.
Population Genomics & Statistical Genetics
- GWAS, polygenic risk scores (PRS), multi-ancestry analysis
- PRS tools: SBayesR (via GCTB), LDpred2, PRS-CS, PLINK (
--score) - Causal inference: Mendelian randomization-style approaches using pharmacogenomic variants (e.g. CYP2D6, CYP2C19, CES1) as genetic instruments
- Martingale residual transformation for Cox-to-linear GWAS
- Large-scale cohort analysis (e.g. iPSYCH2015, N=105,477; ~10,000 PGS computed from FinnGen/UK Biobank summary statistics)
Machine Learning & Predictive Modeling
- Programming: Python, R, SQL, Bash
- Statistical modeling: generalized linear models, multivariate regression, time-series analysis (scikit-learn, statsmodels, pandas, numpy)
- Predictive modeling: LASSO, Random Forest, XGBoost, ensemble/Super Learner methods — applied to clinical prediction modeling in oncology trial data
- Applied transformer models via the Hugging Face pipeline API (BERT, DistilBERT, Sentence-BERT, BERTopic) and PyTorch for GPU-accelerated inference on HPC systems — applying pretrained models, not training architectures from scratch
Infrastructure & Pipelines
- HPC cluster pipelines, Snakemake (basic level)
- Containerization: Docker, Singularity
- Version control: Git
- Data visualization: Matplotlib, Seaborn, ggplot2
Research & Project Experience
- Phenotype QC, ancestry-stratified association testing, and PRS analysis across large genomic cohorts
- Predictive modeling for clinical trial data, contributing to early-phase trial insight generation
- Cross-institutional collaboration on pharmacogenomic causal inference projects
Soft Skills
- Written and verbal communication in English
- Collaboration in interdisciplinary, multicultural, and cross-institutional teams
- Independent project management