Experimental Framework
Agent Eval Lab
Synthetic policy-evaluation framework with a custom PPO-style trainer, distribution-shift evaluation, threshold-based checkpoint judging and reproducibility controls.
GitHub ↗Overview
Studies policy behavior and checkpoint acceptance under controlled synthetic distribution shifts.
Architecture
- Seeded synthetic retrieval-shift and label-noise environments supply one-step tasks (environments/).
- A PyTorch policy/value MLP is trained with PPO-style updates (training/ppo_trainer.py).
- Checkpoint judges and shift evaluation measure explicit thresholds, accuracy gaps and action-distribution KL (environments/retrieval_shift/judge.py and evaluation/ood_eval.py).
Engineering decisions
- Make checkpoint acceptance criteria explicit through configured thresholds.
- Record seeds and configuration hashes to support repeatable experiments.
- Monitor entropy, gradients and KL to abort problematic training runs.
What is implemented
- Custom PPO-style rollout collection and clipped policy updates.
- Synthetic retrieval-shift and label-noise datasets.
- Entropy, gradient and KL stability checks; checkpoint judging and out-of-distribution metrics.
- Seeding and configuration hashing in core/seed.py.
Current limitations
- The environments are synthetic one-step contextual-bandit/classification tasks, not LLM-agent workflows.
- Threshold passing reflects configured experimental criteria rather than a general quality guarantee.
- Deterministic-algorithm configuration uses warn-only behavior; cross-platform determinism is not guaranteed.