Waste Profiles & Optimizations
definity identifies where Spark pipelines waste resources and provides job-specific optimizations to eliminate that waste. Each identified waste profile is surfaced as a savings opportunity with a projected annualized impact, so teams can prioritize the highest-value optimizations first.
Detected waste falls into two categories:
- Unused resources - resources that were provisioned and paid for but not utilized.
- Inefficient use - resources that were utilized, but consumed by avoidable work such as spill, retries, or skew.
Waste Category 1: Unused Resources​
Resources provisioned but not used — allocations that can be safely reduced.
Over-Provisioned Executors​
More executors are allocated to the job than its workload utilizes, leaving executors idle or underused for significant portions of the run.
Over-Provisioned Cluster Machines​
The cluster is sized beyond what the workloads running on it consume, leaving machine capacity unused.
Over-Provisioned Memory​
Total memory allocation significantly exceeds what the job actually uses. Identified per memory region and component:
On-Heap Executor Memory​
The on-heap memory portion of executors is oversized relative to actual usage.
Off-Heap Executor Memory​
The off-heap memory portion of executors is oversized relative to actual usage.
On-Heap Driver Memory​
Driver memory allocation exceeds actual driver memory usage.
Orphaned vCores​
vCores on cluster machines left unusable because the job's high memory-to-vCore ratio exhausts machine memory before all vCores are allocated to executors. These cores are paid for but cannot run any work.
Over-Provisioned Driver Cores​
More driver cores are allocated than the driver workload utilizes.
Waste Category 2: Inefficient Use​
Resources that were consumed, but in a suboptimal manner. These profiles typically point to execution tuning or code-level improvements rather than resizing.
Inefficient Code Patterns​
Inefficiencies originating in the pipeline code itself, where rewriting queries or transformations can reduce the compute required to produce the same output.
Under-Utilized Executor CPU Time​
Executor cores are occupied by tasks but actual CPU usage is low, indicating time spent waiting rather than computing.
Long Skew Time​
Data imbalances cause a small number of tasks to run much longer than their peers, so the stage waits on stragglers while other resources sit idle.
Excessive Partition Listing​
Significant time spent listing partitions and files before or during execution, adding overhead independent of the actual data processing.
Retry & Failure Overhead​
Compute spent re-executing work that already ran once:
Task Retries & Failures​
Compute spent on tasks that fail due to application-level errors and are re-executed.
High Cost of Stage Retries​
Entire stages are recomputed after failures, repeating work that was already paid for.
Spark Internal Task Retries​
Task re-executions triggered by Spark's internal mechanisms rather than by application-level errors.
Excessive Storage Access​
A high volume of read and write requests to cloud storage, adding request costs and latency overhead to the job.
High Spill Overhead​
Tasks spill data to disk during execution, adding I/O overhead and extending run time.
Long Executor GC Time​
A high share of executor time is spent in JVM garbage collection instead of task execution.