Runs agent evaluations with benchmarks, headless orchestration, trace metrics, and LLM judges.
dsh plugin --profile web add dsh-eval
View the source on GitHub