Back to Work
AI Infrastructure / Backend Engineering

Modular LLM Evaluation Framework

Backend infrastructure for orchestrating LLM experiments across multiple model environments.

Institution
Carnegie Mellon University
Role
AI Safety Research Assistant
Timeline
Mar 2025 — Present
Status
Ongoing
OVERVIEW

A modular experimentation pipeline integrating inference APIs, experiment configuration, and automated analysis.

PythonLLM APIsExperiment pipelinesAutomated evaluation
TECHNICAL FLOW
01Experiment config02Model orchestration03Inference APIs04Automated analysis
PROBLEM

LLM behavior needed to be tested consistently across model environments and compared with game-theory baselines.

APPROACH

Create a modular backend that separates experiment configuration, inference integration, orchestration, and analysis.

CONTRIBUTION

Dhruva developed the experimentation pipeline while collaborating with Professor Sarah H. Cen on bounded rationality in institutional settings.

RESULTS

The framework supports repeatable model experiments and automated downstream analysis.

REFLECTION

Reliable evaluation systems require reproducibility, clear configuration, and comparable outputs across providers.