spark-optimization

Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning. Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines.

By wshobson · 9,530 installs

npx skills add wshobson/agents --skill spark-optimization

Source repository · Upstream listing

Apache Spark Optimization Production patterns for optimizing Apache Spark jobs including partitioning strategies, memory management, shuffle optimization, and performance tuning. When to Use This Skill Optimizing slow Spark jobs Tuning memory and executor configuration Implementing efficient partitioning strategies Debugging Spark performance issues Scaling Spark pipelines for large datasets Reducing shuffle and data skew Core Concepts 1. Spark Execution Model 2. Key Performance Factors Factor Impact Solution Shuffle Network I/O, disk I/O Minimize wide transformations Data Skew Uneven task duration Salting, broadcast joins Serialization CPU overhead Use Kryo, columnar formats Memory GC pressure, spills Tune executor memory Partitions Parallelism Right size partitions Quick Start Detailed patterns and worked examples Detailed pattern documentation lives in references/details.md . Read that file when the navigation tier above is insufficient. Best Practices Do's Enable AQE Adaptive query execution handles many issues Use Parquet/Delta Columnar formats with compression Broadcast small tables Avoid shuffle for small joins Monitor Spark UI Check for skew, spills, GC Right size partitions 128MB 256MB per partition Don'ts Don't collect large data Keep data distributed Don't use UDFs unnecessarily Use built in functions Don't over cache Memory is limited Don't ignore data skew It dominates job time Don't use .count() for existence Use .take(1) or .isEmpty()