Computational Biology / Scientific Machine Learning
Bioinformatics ML and Data Pipeline
A large-scale machine-learning workflow for genomic and patient data.
- Institution
- University of California, San Francisco
- Role
- Machine Learning / Bioinformatics Research Assistant
- Timeline
- Aug 2025 — Present
- Status
- Ongoing
R-based ingestion, preprocessing, unsupervised learning, outlier detection, and cross-study analysis.
RPCAClusteringData pipelinesNeural networks
01Data ingestion02Preprocessing03Unsupervised ML04Marker analysis
Large biological datasets contain technical variation, missingness, and high-dimensional signals.
Implement cleaning, feature selection, hierarchical clustering, PCA, network outlier detection, and batch-effect correction.
Dhruva built the pipeline, ran it across more than 125,000 patient rows, and is developing data scraping and neural cross-study analysis.
The system supports lung-cancer gene-marker research with Professor Oldham.
Data quality, reproducibility, and pipeline design are foundational ML-engineering concerns.