Modular LLM Evaluation Framework
Backend infrastructure for orchestrating LLM experiments across multiple model environments.
- Institution
- Carnegie Mellon University
- Role
- AI Safety Research Assistant
- Timeline
- Mar 2025 — Present
- Status
- Ongoing
A modular experimentation pipeline integrating inference APIs, experiment configuration, and automated analysis.
LLM behavior needed to be tested consistently across model environments and compared with game-theory baselines.
Create a modular backend that separates experiment configuration, inference integration, orchestration, and analysis.
Dhruva developed the experimentation pipeline while collaborating with Professor Sarah H. Cen on bounded rationality in institutional settings.
The framework supports repeatable model experiments and automated downstream analysis.
Reliable evaluation systems require reproducibility, clear configuration, and comparable outputs across providers.