Skip to main content
Unlisted page
This page is unlisted. Search engines will not index it, and only users having a direct link can access it.

Waste Profiles & Optimizations

definity identifies where Spark pipelines waste resources and provides job-specific optimizations to eliminate that waste. Each identified waste profile is surfaced as a savings opportunity with a projected annualized impact, so teams can prioritize the highest-value optimizations first.

Detected waste falls into two categories:

  1. Unused resources - resources that were provisioned and paid for but not utilized.
  2. Inefficient use - resources that were utilized, but consumed by avoidable work such as spill, retries, or skew.

Waste Category 1: Unused Resources​

Resources provisioned but not used — allocations that can be safely reduced.

Over-Provisioned Executors​

More executors are allocated to the job than its workload utilizes, leaving executors idle or underused for significant portions of the run.

Over-Provisioned Cluster Machines​

The cluster is sized beyond what the workloads running on it consume, leaving machine capacity unused.

Over-Provisioned Memory​

Total memory allocation significantly exceeds what the job actually uses. Identified per memory region and component:

On-Heap Executor Memory​

The on-heap memory portion of executors is oversized relative to actual usage.

Off-Heap Executor Memory​

The off-heap memory portion of executors is oversized relative to actual usage.

On-Heap Driver Memory​

Driver memory allocation exceeds actual driver memory usage.

Orphaned vCores​

vCores on cluster machines left unusable because the job's high memory-to-vCore ratio exhausts machine memory before all vCores are allocated to executors. These cores are paid for but cannot run any work.

Over-Provisioned Driver Cores​

More driver cores are allocated than the driver workload utilizes.

Waste Category 2: Inefficient Use​

Resources that were consumed, but in a suboptimal manner. These profiles typically point to execution tuning or code-level improvements rather than resizing.

Inefficient Code Patterns​

Inefficiencies originating in the pipeline code itself, where rewriting queries or transformations can reduce the compute required to produce the same output.

Under-Utilized Executor CPU Time​

Executor cores are occupied by tasks but actual CPU usage is low, indicating time spent waiting rather than computing.

Long Skew Time​

Data imbalances cause a small number of tasks to run much longer than their peers, so the stage waits on stragglers while other resources sit idle.

Excessive Partition Listing​

Significant time spent listing partitions and files before or during execution, adding overhead independent of the actual data processing.

Retry & Failure Overhead​

Compute spent re-executing work that already ran once:

Task Retries & Failures​

Compute spent on tasks that fail due to application-level errors and are re-executed.

High Cost of Stage Retries​

Entire stages are recomputed after failures, repeating work that was already paid for.

Spark Internal Task Retries​

Task re-executions triggered by Spark's internal mechanisms rather than by application-level errors.

Excessive Storage Access​

A high volume of read and write requests to cloud storage, adding request costs and latency overhead to the job.

High Spill Overhead​

Tasks spill data to disk during execution, adding I/O overhead and extending run time.

Long Executor GC Time​

A high share of executor time is spent in JVM garbage collection instead of task execution.