meta-optimize

Analyze ARIS usage logs and propose optimizations to SKILL.md files, reviewer prompts, and workflow defaults. Outer-loop harness optimization inspired by Meta-Harness (Lee et al., 2026). Use when user says "优化技能", "meta optimize", "improve skills", "分析使用记录", or wants to optimize ARIS's own harness c

By wanshuiyin · 357 installs

npx skills add wanshuiyin/auto-claude-code-research-in-sleep --skill meta-optimize

Source repository · Upstream listing

Meta Optimize: Outer Loop Harness Optimization for ARIS Analyze accumulated usage logs and propose optimizations for: $ARGUMENTS Privilege boundary — this skill is a READ ONLY PRODUCER meta optimize proposes ; it does not land . The mutation of the skill corpus is the exclusive job of a separate, human invoked skill: [ /meta apply ](../meta apply/SKILL.md). This split is structural, not advisory — it is why a missed instruction cannot let this loop apply its own patch (the self acquittal failure mode): No Write / Edit tool. This skill cannot edit a SKILL.md / shared reference / any corpus file with the frictionless mutators. Its only outputs are the REPORT and staged patch files, written under .aris/meta/ (a scratch area, never the corpus). No apply step. There is no in skill "apply the patch" path (see Step 6). The producer ends by staging approved patches for /meta apply ; a human must then invoke /meta apply to land them. That human action is the landing gate. Bash writes to the corpus are filtered, not impossible — be honest about the layers. What IS fully closed: the accidental / in flow self acquittal — this skill has no Write / Edit and no apply step, so an honest run cannot slip into editing the corpus. Defense in depth: install the [ corpus write guard ](../../templates/claude hooks/corpus write guard.json) PreToolUse hook (like meta logging.json ), which DENIES the common Bash shell writes ( , tee , sed i , cp / mv , touch , open(...,'w') ) to corpus paths. This is a blacklist, NOT a complete sandbox — a deliberately obscured Bash write ( git apply , patch , $var /absolute paths, language file APIs) is not all caught. Full structural prevention requires either removing this skill's Bash or an FS sandbox — over built for a not yet load bearing producer, so deferred to when the gate carries real auto modification volume (a brick 3 trigger). The intended backstop against a deliberate write is detection, not prevention — a corpus change with no valid/current provenance stamp (content hash mismatch) would be catchable in a pre push integrity check — but that verifier is NOT yet built ( provenance.py has content hash but no integrity check subcommand, and no pre push hook runs one). So today the deliberate write case is neither prevented nor actively detected; track the integrity verifier as a follow up before this producer goes load bearing. Its legitimate Bash writes go only to .aris/meta/ . See [ shared references/acceptance gate.md ](../shared references/acceptance gate.md): a loop can DRIVE (propose, review) same model, but the ACQUITTAL that lands a change must be cross model (Step 4 jury) and the landing must be a separate human gated act ( /meta apply ). Context ARIS is a research harness — a system of skills, bridges, workflows, and artifact contracts that wraps around LLMs to orchestrate research. This skill implements a prototype outer loop that observes how the harness is used and proposes improvements to the harness itself (not to the research artifacts it produces). Inspired by Meta Harness (Lee et al., 2026): the key insight is that harness design matters as much as model weights, and harness engineering can be partially automated by logging execution traces and using them to guide improvements. What This Skill Optimizes (Harness Components) Component Example Optimizable? : : SKILL.md prompts Reviewer instructions, quality gates, step descriptions Yes Default parameters difficulty: medium , MAX ROUNDS: 4 , threshold: 6/10 Yes Convergence rules When to stop the review loop, retry counts Yes Workflow ordering Skill chain sequence within a workflow Yes Artifact schemas What fields go in EXPERIMENT LOG.md, idea stage/IDEA REPORT.md Cautious MCP bridge config Which reviewer model, routing rules No (infra) Not optimized : The research artifacts themselves (papers, code, experiments). That's what the regular workflows do. Prerequisites 1. Logging must be active. Copy templates/claude hooks/meta logging.json into your project's .claude/settings.json (or merge the hooks section). 2. Sufficient data. At least 5 complete workflow runs logged in .aris/meta/events.jsonl . The skill will check and warn if insufficient. Workflow Step 0: Check Data Availability If a prior bottleneck entry exists, open the report (Step 5) by stating whether that named bottleneck was resolved (and by which landed patches) and what it has now moved to — bottleneck SUCCESSION, not just existence, is the signal this ledger exists to carry. Step 1: Analyze Usage Patterns Read .aris/meta/events.jsonl and compute: Frequency analysis: Which skills are invoked most often? Which slash commands do users type most? What parameter overrides are most common? (These suggest bad defaults.) Failure analysis: Which tools fail most often? In which skills? What error patterns repeat? (OOM, import, compilation, timeout) How many auto debug retries per workflow run? Convergence analysis (for auto review loop): Average rounds to reach threshold Score trajectory shape (fast improvement? plateau? oscillation?) Which review round catches the most critical issues? Do users override difficulty mid run? Human intervention analysis: Where do users interrupt with manual prompts during workflows? What manual corrections do users make most? (These indicate skill gaps.) Model delta analysis (harness diet): Has the session model ( session start events' model field) or the pinned reviewer model changed since a skill's SKILL.md was last touched? ( git log 1 format=%cs skills/<skill /SKILL.md vs the model bump date.) A model bump is a trigger to re read, not evidence by itself . For each reasoning scaffolding step or worked example in that SKILL.md, a deletion proposal must cite TARGET SPECIFIC evidence that the new model no longer needs it: a capability specific release note, or repeated observed behavior in the event log (e.g. zero failures/interventions in the guarded step since the bump). "The model got newer" alone never justifies a deletion. Never deletion candidates , regardless of model: privilege boundaries, acceptance/review gates, corpus and provenance integrity rules, output contracts, and safety checks. The diet targets model compensation scaffolding only — a capability the new model has natively is pure overhead (context weight, drift surface, reading cost). A harness that only ever grows is a harness nobody is re reading. Trigger rate analysis (optional, measured — not from the event log): The event log shows which skills were USED, not which were WANTED but omitted — the omission failure mode (Claude Code passing over the right skill when the installed list is long) is invisible to it. tools/meta opt/trigger eval.py measures it directly: claude p probes with paraphrased intent queries run from a neutral cwd (so the realistic long installed corpus is loaded), scored as trigger / confusion(→which skill) / miss. Run it when a specific skill is suspected of under or mis triggering, or as a before/after check around a description edit: python3 tools/meta opt/trigger eval.py eval file tools/meta opt/trigger evals.sample.json skills <name samples 2 The confusion matrix is the signal , not just the rate: a query that keeps landing on a sibling skill means the two descriptions overlap on that intent — the fix is disambiguation, not "make the description pushier". Measure only, evidence not verdict. A low trigger rate is an INPUT to a Step 2 proposal (which lands only via /meta apply ), never a self applied description rewrite. Trigger rate is model dependent, so compare like with like (record the probe model) and treat it as a proxy — it measures selection under a query set, not the full long list omission problem. Present findings as a structured summary table. Step 1.5: Name the Current Bottleneck Synthesize the Step 1 analyses into one sentence naming the single most limiting pipeline stage right now — e.g. "planning", "verification quality", "experiment execution reliability", "writing polish" — with the supporting evidence. The bottleneck always moves: when coding stops being the constraint, planning becomes it; when planning is solved, verification; when verification is automated, taste. This step exists to make the CURRENT constraint visible, so Step 2's ranked table reads as sub fixes for one named constraint instead of scattered tweaks. Append the verdict to the append only ledger .aris/meta/bottleneck log.jsonl (same never mutate discipline as .aris/runs/<run id .iterations.jsonl ): Never edit or delete prior lines — succession history is the point. Step 2: Identify Optimization Targets Based on Step 1, rank optimization opportunities by expected impact: The Proposed Change column is explicitly allowed to be a deletion — "DELETE step N, new model does this for free" is a first class optimization, ranked by the same impact logic as additions. If $ARGUMENTS specifies a target skill, focus analysis on that skill only. If $ARGUMENTS is empty or "all", analyze all skills with sufficient data. Step 3: Generate Patch Proposals For each optimization target, generate a concrete diff: Rules for patch generation: One patch per optimization target Each patch must include a comment explaining WHY (with data from the log) Patches must be minimal — change only what the data supports Never change artifact schemas or MCP bridge config in v1 Never change behavior that would break existing user workflows Anti self poisoning screen (see [ shared references/capture antipatterns.md ](../shared references/capture antipatterns.md)): run a proposed patch's rationale through tools/capture filter.py (resolve via the canonical chain). NEVER propose a change that encodes a negative tool capability claim ("codex can't…", "gemini is broken") or a one off / transient failure as a durable rule — those harden into self cited refusals. Encode the fix / the flag needed / the workaround , not "X can't do Y". Step 4: Cross Model Review of Patches (ADVISORY pre screen) This review is advisory — it sharpens the Step 5 REPORT so the human can decide what to stage. It is not the landing verdict. The binding cross model jury runs later, at landing, inside [ /meta apply ](../meta apply/SKILL.md), on the actual staged diff (a producer relayed verdict would be forgeable). Record this result as advisory screen only. Send each patch to GPT 6 Astra xhigh for adversarial review: Step 5: Present Results Output a structured report: Step 6: Stage approved patches for /meta apply (NO in skill apply) This skill does not apply anything. After the user has read the Step 5 REPORT and indicated which changes to land, stage them for the privileged applier: 1. For each approved change N , write its unified diff to .aris/meta/pending/<NN <skill .diff and append a row to .aris/meta/pending/manifest.jsonl : {patch: "<NN <skill .diff", target: "<corpus path ", author model: "<executor ", advisory screen: "pass kill", advisory reason: "<one line "} . The advisory screen (your Step 4 codex pre review) is advisory only — it helps the human read the REPORT. It is NOT the landing verdict and /meta apply does not trust it: a producer written verdict would be forgeable. The binding cross model jury runs at landing, inside /meta apply , on the actual staged diff. 2. Tell the user: "Staged M patches. Run /meta apply to judge & land them." The backup → fresh jury at landing → apply → provenance stamp → log all happen inside [ /meta apply ](../meta apply/SKILL.md). meta