mandela

Audit any eval, metric, experiment, or benchmark for leakage — does external ground-truth enter independently, or are the model, scorer, and designer just confirming a result no outside truth ever produced? Use before trusting any 'how we'll know it worked' — an A/B, a holdout, a score, a validation

By lilmgenius · 1,085 installs

npx skills add lilmgenius/paperthin --skill mandela

Source repository · Upstream listing