mandela
Audit any eval, metric, experiment, or benchmark for leakage — does external ground-truth enter independently, or are the model, scorer, and designer just confirming a result no outside truth ever produced? Use before trusting any 'how we'll know it worked' — an A/B, a holdout, a score, a validation
By lilmgenius · 1,085 installs
npx skills add lilmgenius/paperthin --skill mandela