meta-optimize
Analyze ARIS usage logs and propose optimizations to SKILL.md files, reviewer prompts, and workflow defaults. Outer-loop harness optimization inspired by Meta-Harness (Lee et al., 2026). Use when user says "优化技能", "meta optimize", "improve skills", "分析使用记录", or wants to optimize ARIS's own harness c
By wanshuiyin · 357 installs
npx skills add wanshuiyin/auto-claude-code-research-in-sleep --skill meta-optimize
Source repository · Upstream listing
Meta Optimize: Outer Loop Harness Optimization for ARIS
Analyze accumulated usage logs and propose optimizations for: $ARGUMENTS
Privilege boundary — this skill is a READ ONLY PRODUCER
meta optimize proposes ; it does not land . The mutation of the skill corpus
is the exclusive job of a separate, human invoked skill: [ /meta apply ](../meta apply/SKILL.md).
This split is structural, not advisory — it is why a missed instruction cannot let
this loop apply its own patch (the self acquittal failure mode):
No Write / Edit tool. This skill cannot edit a SKILL.md / shared reference /
any corpus file with the frictionless mutators. Its only outputs are the REPORT and
staged patch files, written under .aris/meta/ (a scratch area, never the corpus).
No apply step. There is no in skill "apply the patch" path (see Step 6). The
producer ends by staging approved patches for /meta apply ; a human must then
invoke /meta apply to land them. That human action is the landing gate.
Bash writes to the corpus are filtered, not impossible — be honest about the
layers. What IS fully closed: the accidental / in flow self acquittal — this skill
has no Write / Edit and no apply step, so an honest run cannot slip into editing the
corpus. Defense in depth: install the
[ corpus write guard ](../../templates/claude hooks/corpus write guard.json) PreToolUse
hook (like meta logging.json ), which DENIES the common Bash shell writes ( , tee ,
sed i , cp / mv , touch , open(...,'w') ) to corpus paths. This is a blacklist,
NOT a complete sandbox — a deliberately obscured Bash write ( git apply , patch ,
$var /absolute paths, language file APIs) is not all caught. Full structural
prevention requires either removing this skill's Bash or an FS sandbox — over built
for a not yet load bearing producer, so deferred to when the gate carries real
auto modification volume (a brick 3 trigger). The intended backstop against a deliberate
write is detection, not prevention — a corpus change with no valid/current
provenance stamp (content hash mismatch) would be catchable in a pre push integrity
check — but that verifier is NOT yet built ( provenance.py has content hash but no
integrity check subcommand, and no pre push hook runs one). So today the deliberate write
case is neither prevented nor actively detected; track the integrity verifier as a
follow up before this producer goes load bearing. Its legitimate Bash writes go only to
.aris/meta/ .
See [ shared references/acceptance gate.md ](../shared references/acceptance gate.md):
a loop can DRIVE (propose, review) same model, but the ACQUITTAL that lands a change
must be cross model (Step 4 jury) and the landing must be a separate human gated
act ( /meta apply ).
Context
ARIS is a research harness — a system of skills, bridges, workflows, and artifact contracts that wraps around LLMs to orchestrate research. This skill implements a prototype outer loop that observes how the harness is used and proposes improvements to the harness itself (not to the research artifacts it produces).
Inspired by Meta Harness (Lee et al., 2026): the key insight is that harness design matters as much as model weights, and harness engineering can be partially automated by logging execution traces and using them to guide improvements.
What This Skill Optimizes (Harness Components)
Component Example Optimizable?
: :
SKILL.md prompts Reviewer instructions, quality gates, step descriptions Yes
Default parameters difficulty: medium , MAX ROUNDS: 4 , threshold: 6/10 Yes
Convergence rules When to stop the review loop, retry counts Yes
Workflow ordering Skill chain sequence within a workflow Yes
Artifact schemas What fields go in EXPERIMENT LOG.md, idea stage/IDEA REPORT.md Cautious
MCP bridge config Which reviewer model, routing rules No (infra)
Not optimized : The research artifacts themselves (papers, code, experiments). That's what the regular workflows do.
Prerequisites
1. Logging must be active. Copy templates/claude hooks/meta logging.json into your project's .claude/settings.json (or merge the hooks section).
2. Sufficient data. At least 5 complete workflow runs logged in .aris/meta/events.jsonl . The skill will check and warn if insufficient.
Workflow
Step 0: Check Data Availability
If a prior bottleneck entry exists, open the report (Step 5) by stating whether
that named bottleneck was resolved (and by which landed patches) and what it
has now moved to — bottleneck SUCCESSION, not just existence, is the signal
this ledger exists to carry.
Step 1: Analyze Usage Patterns
Read .aris/meta/events.jsonl and compute:
Frequency analysis:
Which skills are invoked most often?
Which slash commands do users type most?
What parameter overrides are most common? (These suggest bad defaults.)
Failure analysis:
Which tools fail most often? In which skills?
What error patterns repeat? (OOM, import, compilation, timeout)
How many auto debug retries per workflow run?
Convergence analysis (for auto review loop):
Average rounds to reach threshold
Score trajectory shape (fast improvement? plateau? oscillation?)
Which review round catches the most critical issues?
Do users override difficulty mid run?
Human intervention analysis:
Where do users interrupt with manual prompts during workflows?
What manual corrections do users make most? (These indicate skill gaps.)
Model delta analysis (harness diet):
Has the session model ( session start events' model field) or the pinned
reviewer model changed since a skill's SKILL.md was last touched?
( git log 1 format=%cs skills/<skill /SKILL.md vs the model bump date.)
A model bump is a trigger to re read, not evidence by itself . For each
reasoning scaffolding step or worked example in that SKILL.md, a deletion
proposal must cite TARGET SPECIFIC evidence that the new model no longer
needs it: a capability specific release note, or repeated observed behavior
in the event log (e.g. zero failures/interventions in the guarded step since
the bump). "The model got newer" alone never justifies a deletion.
Never deletion candidates , regardless of model: privilege boundaries,
acceptance/review gates, corpus and provenance integrity rules, output
contracts, and safety checks. The diet targets model compensation scaffolding
only — a capability the new model has natively is pure overhead (context
weight, drift surface, reading cost). A harness that only ever grows is a
harness nobody is re reading.
Trigger rate analysis (optional, measured — not from the event log):
The event log shows which skills were USED, not which were WANTED but omitted
— the omission failure mode (Claude Code passing over the right skill when the
installed list is long) is invisible to it. tools/meta opt/trigger eval.py
measures it directly: claude p probes with paraphrased intent queries run
from a neutral cwd (so the realistic long installed corpus is loaded), scored
as trigger / confusion(→which skill) / miss.
Run it when a specific skill is suspected of under or mis triggering, or as a
before/after check around a description edit:
python3 tools/meta opt/trigger eval.py eval file tools/meta opt/trigger evals.sample.json skills <name samples 2
The confusion matrix is the signal , not just the rate: a query that keeps
landing on a sibling skill means the two descriptions overlap on that intent —
the fix is disambiguation, not "make the description pushier".
Measure only, evidence not verdict. A low trigger rate is an INPUT to a
Step 2 proposal (which lands only via /meta apply ), never a self applied
description rewrite. Trigger rate is model dependent, so compare like with
like (record the probe model) and treat it as a proxy — it measures selection
under a query set, not the full long list omission problem.
Present findings as a structured summary table.
Step 1.5: Name the Current Bottleneck
Synthesize the Step 1 analyses into one sentence naming the single
most limiting pipeline stage right now — e.g. "planning", "verification
quality", "experiment execution reliability", "writing polish" — with the
supporting evidence. The bottleneck always moves: when coding stops being the
constraint, planning becomes it; when planning is solved, verification; when
verification is automated, taste. This step exists to make the CURRENT
constraint visible, so Step 2's ranked table reads as sub fixes for one named
constraint instead of scattered tweaks.
Append the verdict to the append only ledger .aris/meta/bottleneck log.jsonl
(same never mutate discipline as .aris/runs/<run id .iterations.jsonl ):
Never edit or delete prior lines — succession history is the point.
Step 2: Identify Optimization Targets
Based on Step 1, rank optimization opportunities by expected impact:
The Proposed Change column is explicitly allowed to be a deletion — "DELETE
step N, new model does this for free" is a first class optimization, ranked by
the same impact logic as additions.
If $ARGUMENTS specifies a target skill, focus analysis on that skill only.
If $ARGUMENTS is empty or "all", analyze all skills with sufficient data.
Step 3: Generate Patch Proposals
For each optimization target, generate a concrete diff:
Rules for patch generation:
One patch per optimization target
Each patch must include a comment explaining WHY (with data from the log)
Patches must be minimal — change only what the data supports
Never change artifact schemas or MCP bridge config in v1
Never change behavior that would break existing user workflows
Anti self poisoning screen (see [ shared references/capture antipatterns.md ](../shared references/capture antipatterns.md)):
run a proposed patch's rationale through tools/capture filter.py (resolve via
the canonical chain). NEVER propose a change that encodes a negative
tool capability claim ("codex can't…", "gemini is broken") or a one off /
transient failure as a durable rule — those harden into self cited refusals.
Encode the fix / the flag needed / the workaround , not "X can't do Y".
Step 4: Cross Model Review of Patches (ADVISORY pre screen)
This review is advisory — it sharpens the Step 5 REPORT so the human can decide
what to stage. It is not the landing verdict. The binding cross model jury runs
later, at landing, inside [ /meta apply ](../meta apply/SKILL.md), on the actual staged
diff (a producer relayed verdict would be forgeable). Record this result as
advisory screen only.
Send each patch to GPT 6 Astra xhigh for adversarial review:
Step 5: Present Results
Output a structured report:
Step 6: Stage approved patches for /meta apply (NO in skill apply)
This skill does not apply anything. After the user has read the Step 5 REPORT and
indicated which changes to land, stage them for the privileged applier:
1. For each approved change N , write its unified diff to
.aris/meta/pending/<NN <skill .diff and append a row to
.aris/meta/pending/manifest.jsonl :
{patch: "<NN <skill .diff", target: "<corpus path ", author model: "<executor ",
advisory screen: "pass kill", advisory reason: "<one line "} .
The advisory screen (your Step 4 codex pre review) is advisory only — it helps
the human read the REPORT. It is NOT the landing verdict and /meta apply does not
trust it: a producer written verdict would be forgeable. The binding cross model jury
runs at landing, inside /meta apply , on the actual staged diff.
2. Tell the user: "Staged M patches. Run /meta apply to judge & land them."
The backup → fresh jury at landing → apply → provenance stamp → log all happen
inside [ /meta apply ](../meta apply/SKILL.md). meta