arize-evaluator
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and continuous monitoring. Use when the user mentions create evaluator, LLM judge, hallucination, faithfulne
By github · 1,152 installs
npx skills add github/awesome-copilot --skill arize-evaluator
Source repository · Upstream listing
Arize Evaluator Skill
SPACE — All space flags and the ARIZE SPACE env var accept a space name (e.g., my workspace ) or a base64 space ID (e.g., U3BhY2U6... ). Find yours with ax spaces list .
This skill covers designing, creating, and running LLM as judge evaluators on Arize. An evaluator defines the judge; a task is how you run it against real data.
Prerequisites
Proceed directly with the task — run the ax command you need. Do NOT check versions, env vars, or profiles upfront.
If an ax command fails, troubleshoot based on the error:
command not found or version error → see references/ax setup.md
401 Unauthorized / missing API key → run ax profiles show to inspect the current profile. If the profile is missing or the API key is wrong, follow references/ax profiles.md to create/update it. If the user doesn't have their key, direct them to https://app.arize.com/admin API Keys
Space unknown → run ax spaces list to pick by name, or ask the user
LLM provider call fails (missing OPENAI API KEY / ANTHROPIC API KEY) → run ax ai integrations list space SPACE to check for platform managed credentials. If none exist, ask the user to provide the key or create an integration via the arize ai provider integration skill
Security: Never read .env files or search the filesystem for credentials. Use ax profiles for Arize credentials and ax ai integrations for LLM provider keys. If credentials are not available through these channels, ask the user.
CRITICAL — Never fabricate evaluation results: If an evaluation task fails, is cancelled, or produces no scores, report the failure clearly and explain what went wrong. Do NOT perform a "manual evaluation," invent quality scores, estimate percentages, or present any agent generated analysis as if it came from the Arize evaluation system. Instead suggest: (1) fix the identified issue and retry, (2) try running from the Arize UI, (3) verify integration credentials with ax ai integrations list , (4) contact support at https://arize.com/support
Concepts
What is an Evaluator?
An evaluator is an LLM as judge definition. It contains:
Field Description
Template The judge prompt. Uses {variable} placeholders (e.g. {input} , {output} , {context} ) that get filled in at run time via a task's column mappings.
Classification choices The set of allowed output labels (e.g. factual / hallucinated ). Binary is the default and most common. Each choice can optionally carry a numeric score.
AI Integration Stored LLM provider credentials (OpenAI, Anthropic, Bedrock, etc.) the evaluator uses to call the judge model.
Model The specific judge model (e.g. gpt 4o , claude sonnet 4 5 ).
Invocation params Optional JSON of model settings like {"temperature": 0} . Low temperature is recommended for reproducibility.
Optimization direction Whether higher scores are better ( maximize ) or worse ( minimize ). Sets how the UI renders trends.
Data granularity Whether the evaluator runs at the span , trace , or session level. Most evaluators run at the span level.
Evaluators are versioned — every prompt or model change creates a new immutable version. The most recent version is active.
What is a Task?
A task is how you run one or more evaluators against real data. Tasks are attached to a project (live traces/spans) or a dataset (experiment runs). A task contains:
Field Description
Evaluators List of evaluators to run. You can run multiple in one task.
Column mappings Maps each evaluator's template variables to actual field paths on spans or experiment runs (e.g. "input" → "attributes.input.value" ). This is what makes evaluators portable across projects and experiments.
Query filter SQL style expression to select which spans/runs to evaluate (e.g. "span kind = 'LLM'" ). Optional but important for precision.
Continuous For project tasks: whether to automatically score new spans as they arrive.
Sampling rate For continuous project tasks: fraction of new spans to evaluate (0–1).
Data Granularity
The data granularity flag controls what unit of data the evaluator scores. It defaults to span and only applies to project tasks (not dataset/experiment tasks — those evaluate experiment runs directly).
Level What it evaluates Use for Result column prefix
span (default) Individual spans Q&A correctness, hallucination, relevance eval.{name}.label / .score / .explanation
trace All spans in a trace, grouped by context.trace id Agent trajectory, task correctness — anything that needs the full call chain trace eval.{name}.label / .score / .explanation
session All traces in a session, grouped by attributes.session.id and ordered by start time Multi turn coherence, overall tone, conversation quality session eval.{name}.label / .score / .explanation
How trace and session aggregation works
For trace granularity, spans sharing the same context.trace id are grouped together. Column values used by the evaluator template are comma joined into a single string (each value truncated to 100K characters) before being passed to the judge model.
For session granularity, the same trace level grouping happens first, then traces are ordered by start time and grouped by attributes.session.id . Session level values are capped at 100K characters total.
The {conversation} template variable
At session granularity, {conversation} is a special template variable that renders as a JSON array of {input, output} turns across all traces in the session, built from attributes.input.value / attributes.llm.input messages (input side) and attributes.output.value / attributes.llm.output messages (output side).
At span or trace granularity, {conversation} is treated as a regular template variable and resolved via column mappings like any other.
Multi evaluator tasks
A task can contain evaluators at different granularities. At runtime the system uses the highest granularity (session trace span) for data fetching and automatically splits into one child run per evaluator . Per evaluator query filter in the task's evaluators JSON further narrows which spans are included (e.g., only tool call spans within a session).
Basic CRUD
AI Integrations
AI integrations store the LLM provider credentials the evaluator uses. For full CRUD — listing, creating for all providers (OpenAI, Anthropic, Azure, Bedrock, Vertex, Gemini, NVIDIA NIM, custom), updating, and deleting — use the arize ai provider integration skill.
Quick reference for the common case (OpenAI):
Copy the returned integration ID — it is required for ax evaluators create ai integration id .
Evaluators
Key flags for create :
Flag Required Description
name yes Evaluator name (unique within space)
space yes Space name or ID to create in
template name yes Eval column name — alphanumeric, spaces, hyphens, underscores
commit message yes Description of this version
ai integration id yes AI integration ID (from above)
model name yes Judge model (e.g. gpt 4o )
template yes Prompt with {variable} placeholders (single quoted in bash)
classification choices yes JSON object mapping choice labels to numeric scores e.g. '{"correct": 1, "incorrect": 0}'
description no Human readable description
include explanations no Include reasoning alongside the label
use function calling no Prefer structured function call output
invocation params no JSON of model params e.g. '{"temperature": 0}'
data granularity no span (default), trace , or session . Only relevant for project tasks, not dataset/experiment tasks. See Data Granularity section.
direction no Optimization direction: maximize or minimize . Sets how the UI renders trends.
provider params no JSON object of provider specific parameters
Tasks
PROJECT NAME , DATASET NAME , and evaluator id all accept a name or base64 ID.
Time format for trigger run: 2026 03 21T09:00:00 — no trailing Z .
Additional trigger run flags:
Flag Description
max spans Cap processed spans (default 10,000)
override evaluations Re score spans that already have labels
wait / w Block until the run finishes
timeout Seconds to wait with wait (default 600)
poll interval Poll interval in seconds when waiting (default 5)
Run status guide:
Status Meaning
completed , 0 spans The eval index lags 1–2 hours — spans ingested recently may not be indexed yet. Shift the window to data at least 2 hours old, or widen the time range to cover more historical data.
cancelled ~1s Integration credentials invalid
cancelled ~3min Found spans but LLM call failed — check model name or key
completed , N 0 Success — check scores in UI
Workflow A: Create an evaluator for a project
Use this when the user says something like "create an evaluator for my Playground Traces project" .
Step 1: Confirm the project name
ax spans export accepts a project name directly — no ID lookup needed. If you don't know the project name, list available projects:
Find the entry whose "name" matches (case insensitive) and use that name as PROJECT in subsequent commands. If you later hit a validation error with a name, fall back to using the project's "id" (a base64 string) instead.
Step 2: Understand what to evaluate
If the user specified the evaluator type (hallucination, correctness, relevance, etc.) → skip to Step 3.
If not, sample recent spans to base the evaluator on actual data:
Inspect attributes.input , attributes.output , span kinds, and any existing annotations. Identify failure modes (e.g. hallucinated facts, off topic answers, missing context) and propose 1–3 concrete evaluator ideas . Let the user pick.
Each suggestion must include: the evaluator name (bold), a one sentence description of what it judges, and the binary label pair in parentheses. Format each like:
1. Name — Description of what is being judged. ( label a / label b )
Example:
1. Response Correctness — Does the agent's response correctly address the user's financial query? ( correct / incorrect )
2. Hallucination — Does the response fabricate facts not grounded in retrieved context? ( factual / hallucinated )
Step 3: Confirm or create an AI integration
If a suitable integration exists, note its ID. If not, create one using the arize ai provider integration skill. Ask the user which provider/model they want for the judge.
Step 4: Create the evaluator
Use the template design best practices below. Keep the evaluator name and variables generic — the task (Step 6) handles project specific wiring via column mappings .
Step 5: Ask — backfill, continuous, or both?
Recommended approach: Always start with a small backfill (~100 historical spans) to validate the evaluator before turning on continuous monitoring. This lets you catch column mapping errors, wrong span kinds, and template issues on known data before scoring all future production spans. Only enable continuous after a backfill confirms correct scoring.
Before creating the task, ask:
"Would you like to:
(a) Run a backfill on historical spans (one time)?
(b) Set up continuous evaluation on new spans going forward?
(c) Both — backfill first to validate, then keep scoring new spans automatically? (recommended)"
Step 6: Determine column mappings from real span data
Do not guess paths. Pull a sample and inspect what fields are actually present:
For each template variable ( {input} , {output} , {context} ), find the matching JSON path. Common start