Protein Evolution with ESM Models
Benchmarks ESM-1v and ESM-2 on APP deep-mutational scanning data: reconstruct variants, score MLLR/PLL/embedding distance, then ask where the signal holds up for epistasis and synthetic fitness landscapes.
Work
GitHub-backed work from protein language models, docking, single-cell perturbation checks, and reproducible genomics workflows.
Benchmarks ESM-1v and ESM-2 on APP deep-mutational scanning data: reconstruct variants, score MLLR/PLL/embedding distance, then ask where the signal holds up for epistasis and synthetic fitness landscapes.
A Linux-first molecular docking workstation that takes receptor and ligand files through cleanup, pocket choice, Vina/GNINA docking, result ranking, 3D viewing, validation cases, and exportable reports.
FastQC-style guardrails for single-cell perturbation benchmarks. It audits AnnData datasets, splits, model claims, leakage, controls, confounding, weak target effects, and turns the warnings into HTML/CSV/JSON reports.
A breast-cancer subtype project that combines TCGA BRCA clinical data, RNA expression, DNA methylation, epigenetic age acceleration, immune scores, and multiclass models for Basal/Luminal subtype analysis.
A reproducible bacterial whole-genome sequencing workflow from paired-end reads to assembly, polishing, QC, Bakta annotation, optional AMR/virulence/plasmid screening, and core-genome phylogeny.
A Snakemake pipeline for adenovirus reads: QC, optional host decontamination, de novo assembly, reference selection, scaffolding, consensus calling, variability plots, and bootstrapped phylogeny.
A Snakemake RNA-seq workflow with reference and de novo branches, trimming/QC, count matrices, RSEM or kallisto quantification, and MultiQC reports collected into a repeatable results folder.