eval-engineering
Inspect an agent repository and optional traces, interview the user, write reviewed Task Specs, build and audit Harbor tasks, and bootstrap reusable project World Knowledge Skills. Use for agent evals, benchmark design, Task generation, controlled Environments, synthetic data, Verifiers, Harbor runs
By langchain-ai · 2,704 installs
npx skills add langchain-ai/langchain-skills --skill eval-engineering