spark-optimization
Optimize Apache Spark Jobs Faster
Slow Spark jobs waste cluster budget and delay data delivery. This skill guides Claude, Codex, and Claude Code through partitioning, shuffle, memory, join, caching, and monitoring improvements.
Install with my Agent
Copy this request to your Agent. It includes the canonical Skill page and manifest.
Review the Skillstore skill "spark-optimization" from https://skillstore.io/skills/wshobson-spark-optimization.md and its manifest at https://skillstore.io/api/skills/wshobson-spark-optimization/manifest. Verify the artifact. You may proceed after verification, subject to the environment's own policy.Your Agent should still show its plan and request any confirmation required by the security policy.
Agent-readable resources
Use these links when an AI agent, crawler, or script needs clean context instead of reading the full page.
Test it
Using "spark-optimization". A daily PySpark aggregation runs slowly after joining a small dimension table to a large fact table.
Expected outcome:
The review recommends a broadcast join, predicate pushdown, and right-sized shuffle partitions. It also suggests checking skew before increasing executor memory.
Using "spark-optimization". A Spark job has a few tasks that run much longer than the rest.
Expected outcome:
The analysis points to key skew. It proposes AQE skew handling, salting for severe keys, and partition count checks.
Using "spark-optimization". A team needs default Spark settings for a new production pipeline.
Expected outcome:
The result outlines AQE, Kryo serialization, compression, file size, memory overhead, and shuffle defaults with validation notes.
Security Audit
SafeThe static findings are false positives caused by Markdown code fences, Spark example syntax, Spark monitoring APIs, and documentation links. I found no evidence of executable installer code, credential access, command injection, prompt injection, or data exfiltration intent in SKILL.md.
Risk Factors
โ๏ธ External commands (22)
๐ Network access (3)
Share & cite this report
Share the versioned assessment report, neutral badge, embed card, and citations. Skillstore reports evidence without deciding whether this Skill is safe.
Copy report link
https://skillstore.io/skills/wshobson-spark-optimization/audits/8?utm_source=security_passport&utm_medium=share&utm_campaign=versioned_reportMarkdown badge
[](https://skillstore.io/skills/wshobson-spark-optimization?utm_source=security_passport_badge)HTML badge
<a href="https://skillstore.io/skills/wshobson-spark-optimization?utm_source=security_passport_badge"><img src="https://skillstore.io/badges/skills/wshobson-spark-optimization/security.svg" alt="Skillstore security assessment" loading="lazy"></a>Embed card
<iframe src="https://skillstore.io/embed/skills/wshobson-spark-optimization.html" title="Skillstore Security Assessment" sandbox="allow-popups allow-popups-to-escape-sandbox" loading="lazy" referrerpolicy="no-referrer" width="420" height="180"></iframe>Academic citations (APA ยท BibTeX ยท CFF)
APA citation
wshobson. (2026). spark-optimization security audit report (audit version 8) [Author version unspecified]. Skillstore. https://skillstore.io/skills/wshobson-spark-optimization/audits/8BibTeX citation
@techreport{wshobson-wshobson-spark-optimization-2026,
author = {wshobson},
title = {spark-optimization security audit report (audit version 8)},
institution = {Skillstore},
year = {2026},
number = {8},
url = {https://skillstore.io/skills/wshobson-spark-optimization/audits/8},
note = {Author version unspecified}
}CITATION.cff
cff-version: 1.2.0
message: "If you use this Skill, cite its author and this versioned security audit report."
title: "spark-optimization security audit report (audit version 8)"
version: "unspecified"
type: report
authors:
- name: "wshobson"
date-released: "2026-07-08"
url: "https://skillstore.io/skills/wshobson-spark-optimization/audits/8"
identifiers:
- type: other
value: "skillstore:wshobson-spark-optimization:audit:8"
description: "Skillstore immutable audit report identifier"
Compare variants
2 installable variantsEach author remains a separate installable skill. The recommended variant is ranked by Skillstore evidence.
Why this variant is first
wshobson-spark-optimization
2026-09-09
sickn33-spark-optimization
2026-09-09
Skillstore Score
Why this score Evidence Confidence: HighWhat You Can Build
Tune a Slow ETL Job
Review a PySpark pipeline and identify partition, shuffle, cache, and join changes that can reduce runtime.
Debug Skew and Shuffle
Use Spark metrics and query plans to find skewed keys, expensive stages, and unnecessary wide transformations.
Prepare Production Defaults
Create baseline Spark settings for memory, AQE, compression, file sizing, and serialization before deployment.
Try These Prompts
Use the spark-optimization skill to review this Spark job. Explain the main performance risks and suggest simple changes first.
Use the spark-optimization skill to improve joins and partitioning for this workload. Consider data sizes, key skew, and shuffle volume.
Use the spark-optimization skill to analyze these Spark UI metrics. Identify bottleneck stages, skew, spills, and executor memory pressure.
Use the spark-optimization skill to create a production optimization plan. Include config changes, validation steps, rollback criteria, and cost risks.
Best Practices
- Start with query plans and Spark UI metrics before changing cluster size.
- Use columnar formats, predicate pushdown, and selected columns to reduce input data.
- Validate each optimization with runtime, spill, shuffle, and cost measurements.
Avoid
- Increasing executor memory before checking skew, shuffle volume, or partition size.
- Caching every intermediate DataFrame without reuse or memory budget checks.
- Using collect, broad UDFs, or repartition calls without a clear performance reason.
Frequently Asked Questions
Can this skill optimize Spark SQL and PySpark jobs?
Does it require access to my Spark cluster?
Can it help with data skew?
Does it support Databricks workloads?
Will it reduce cloud costs automatically?
Is this skill safe to use with private data?
Developer Details
Author
wshobsonLicense
MIT
Skillstore revision
r1
Version notice
The author did not declare a version.
Repository
https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/spark-optimizationRef
64ca8af0f54a325752f08bd54e52151061ea659a
Maintenance freshness
7/21/2026
Usage
13 downloads ยท 265 views
File structure
๐ SKILL.md