Back to Work
Computational Biology / Scientific Machine Learning

Bioinformatics ML and Data Pipeline

A large-scale machine-learning workflow for genomic and patient data.

Institution
University of California, San Francisco
Role
Machine Learning / Bioinformatics Research Assistant
Timeline
Aug 2025 — Present
Status
Ongoing
OVERVIEW

R-based ingestion, preprocessing, unsupervised learning, outlier detection, and cross-study analysis.

RPCAClusteringData pipelinesNeural networks
TECHNICAL FLOW
01Data ingestion02Preprocessing03Unsupervised ML04Marker analysis
PROBLEM

Large biological datasets contain technical variation, missingness, and high-dimensional signals.

APPROACH

Implement cleaning, feature selection, hierarchical clustering, PCA, network outlier detection, and batch-effect correction.

CONTRIBUTION

Dhruva built the pipeline, ran it across more than 125,000 patient rows, and is developing data scraping and neural cross-study analysis.

RESULTS

The system supports lung-cancer gene-marker research with Professor Oldham.

REFLECTION

Data quality, reproducibility, and pipeline design are foundational ML-engineering concerns.