monitor-experiment
Monitor running experiments, check progress, collect results. Use when user says "check results", "is it done", "monitor", or wants experiment output.
By wanshuiyin · 421 installs
npx skills add wanshuiyin/auto-claude-code-research-in-sleep --skill monitor-experiment
Source repository · Upstream listing
Monitor Experiment Results
⏱ External cadence is appropriate here. This skill waits on an external
fact (job completion / progress), so it is a natural /loop / CronCreate
surface: the wake reads status and self judges only machine checkable
completion (exit code, file exists, epoch logged) — never quality. This is
the additive external wait shape in
[ shared references/external cadence.md ](../shared references/external cadence.md).
If a scheduled wait here ends in a verdict step (e.g. then audit results),
run that verdict once after the wait clears — not re entered per tick.
Monitor: $ARGUMENTS
Workflow
Step 1: Check What's Running
SSH server:
Vast.ai instance (read ssh host , ssh port from vast instances.json ):
Also check vast.ai instance status:
Modal (when gpu: modal in CLAUDE.md):
Modal apps auto terminate when done — if it's not in the list, it already finished. Check results via modal volume ls <volume or local output.
Step 2: Collect Output from Each Screen
For each screen session, capture the last N lines:
If hardcopy fails, check for log files or tee output.
Step 3: Check for JSON Result Files
If JSON results exist, fetch and parse them:
Step 3.5: Pull W&B Metrics (when wandb: true in CLAUDE.md)
Skip this step entirely if wandb is not set or is false in CLAUDE.md.
Pull training curves and metrics from Weights & Biases via Python API:
What to extract:
Training loss curve — is it converging? diverging? plateauing?
Eval metrics — loss, PPL, accuracy at latest checkpoint
Learning rate — is the schedule behaving as expected?
GPU memory — any OOM risk?
Run status — running / finished / crashed?
W&B dashboard link (include in summary for user):
This gives the auto review loop richer signal than just screen output — training dynamics, loss curves, and metric trends over time.
Step 4: Summarize Results
Present results in a comparison table:
Step 5: Interpret
Compare against known baselines
Flag unexpected results (negative delta, NaN, divergence)
Suggest next steps based on findings
Step 6: Feishu Notification (if configured)
After results are collected, check ~/.claude/feishu.json :
Send experiment done notification: results summary table, delta vs baseline
If config absent or mode "off" : skip entirely (no op)
Key Rules
Always show raw numbers before interpretation
Compare against the correct baseline (same config)
Note if experiments are still running (check progress bars, iteration counts)
If results look wrong, check training logs for errors before concluding
Vast.ai cost awareness : When monitoring vast.ai instances, report the running cost (hours $/hr from vast instances.json ). If all experiments on an instance are done, remind the user to run /vast gpu destroy <instance id to stop billing
Modal cost awareness : Modal auto scales to zero — no idle billing. When reporting results from Modal runs, note the actual execution time and estimated cost (time $/hr from the GPU tier used). No cleanup action needed