model-pruning

Reduce LLM size and accelerate inference using pruning techniques like Wanda and SparseGPT. Use when compressing models without retraining, achieving 50% sparsity with minimal accuracy loss, or enabling faster inference on hardware accelerators. Covers unstructured pruning, structured pruning, N:M s

By orchestra-research · 758 installs

npx skills add orchestra-research/ai-research-skills --skill model-pruning

Source repository · Upstream listing