← All projects

Experimental Framework

Agent Eval Lab

Synthetic policy-evaluation framework with a custom PPO-style trainer, distribution-shift evaluation, threshold-based checkpoint judging and reproducibility controls.

GitHub ↗

Overview

Studies policy behavior and checkpoint acceptance under controlled synthetic distribution shifts.

Architecture

  • Seeded synthetic retrieval-shift and label-noise environments supply one-step tasks (environments/).
  • A PyTorch policy/value MLP is trained with PPO-style updates (training/ppo_trainer.py).
  • Checkpoint judges and shift evaluation measure explicit thresholds, accuracy gaps and action-distribution KL (environments/retrieval_shift/judge.py and evaluation/ood_eval.py).

Engineering decisions

  • Make checkpoint acceptance criteria explicit through configured thresholds.
  • Record seeds and configuration hashes to support repeatable experiments.
  • Monitor entropy, gradients and KL to abort problematic training runs.

What is implemented

  • Custom PPO-style rollout collection and clipped policy updates.
  • Synthetic retrieval-shift and label-noise datasets.
  • Entropy, gradient and KL stability checks; checkpoint judging and out-of-distribution metrics.
  • Seeding and configuration hashing in core/seed.py.

Current limitations

  • The environments are synthetic one-step contextual-bandit/classification tasks, not LLM-agent workflows.
  • Threshold passing reflects configured experimental criteria rather than a general quality guarantee.
  • Deterministic-algorithm configuration uses warn-only behavior; cross-platform determinism is not guaranteed.

Stack

  • Python
  • PyTorch
  • PPO
  • Evaluation
  • Distribution Shift
  • Synthetic Data