AWS EKS Cost Optimization: Karpenter, Spot, Graviton and Pod Rightsizing
Reduce Amazon EKS cost through pod allocation, resource rightsizing, Karpenter NodePools, consolidation, Spot diversification, Graviton and commitment-aware capacity design.
Amazon EKS cost optimization is a systems problem. Pod requests influence scheduling. Scheduling determines node demand. Node provisioning determines instance family, architecture, Availability Zone and purchase model. Those choices then interact with disruption budgets, topology, autoscaling and service reliability.
Optimizing only one layer can move waste elsewhere. Smaller nodes do not help when requests are inflated. Spot does not help when workloads cannot tolerate interruption. Karpenter cannot choose economical capacity when NodePools and pod constraints leave it one eligible instance type.
A durable program proceeds in order: allocate cost, understand demand, rightsize pods, coordinate autoscaling, simplify constraints, improve node economics and verify the result against service objectives.
Establish complete EKS cost visibility
Worker nodes are usually the largest line item, but the cluster also consumes EKS control planes, EBS, EFS, load balancers, NAT gateways, cross-zone transfer, public egress, CloudWatch, security agents, backups and shared platform services.
AWS split cost allocation data can add pod-level CPU and memory cost records to the Cost and Usage Report. Kubernetes dimensions include cluster, namespace, deployment, workload and node. This gives FinOps and engineering a common billing basis, although a complete model still needs shared costs and non-compute services.
Define an allocation hierarchy such as product, service, environment, team, cluster, namespace and workload. Publish idle capacity and shared platform cost rather than hiding them inside an arbitrary application. The multi-cloud Kubernetes cost optimization playbook covers allocation design in more depth.
Rightsize requests before changing nodes
The scheduler uses requests, not average observed consumption, to place pods. If a deployment requests four CPU cores but normally needs one, the unused reservation prevents other workloads from using the node and can trigger unnecessary scale-out.
Build a workload profile from representative windows. Include ordinary demand, known peaks, deployments, background jobs, garbage collection and failure recovery. For CPU, review usage percentiles and throttling. For memory, review working set, peaks, OOM kills and restart behavior. Include sidecars because their requests contribute to the pod total.
Make changes in stages. A request is a reliability control as well as a cost input. Test lower environments, limit production rollout, monitor pending pods, latency, throttling, OOM events and SLOs, and keep a rollback path.
The Horizontal Pod Autoscaler, Vertical Pod Autoscaler and node autoscaler operate at different layers. Their policies must be compatible. VPA recommendation mode can provide evidence without automatically restarting production workloads. HPA should scale on a signal that represents demand, not blindly on a CPU target that conflicts with frequently changing requests.
Decide where Karpenter fits
Karpenter provisions nodes from the scheduling needs of pending pods. It can evaluate resource requests, architecture, zones, taints, labels and other constraints, then choose compatible EC2 capacity. It is well suited to clusters with varied or rapidly changing demand.
Managed node groups can remain appropriate for stable baseline or special operational requirements. Many production designs use a small, predictable managed group for critical cluster services and Karpenter for dynamic application capacity.
Keep NodePools purposeful. A separate pool is justified when workloads have a genuinely different security, architecture, interruption or hardware requirement. A pool per team or application often fragments capacity and weakens consolidation.
Give Karpenter room to choose
Cost and availability improve when the provisioner can select from multiple compatible instance families, sizes and Availability Zones. Excessive node selectors, required affinities and narrow instance allowlists reduce those choices.
Review constraints from both directions:
| Constraint | Question |
|---|---|
| Architecture | Can the image and dependencies run on arm64 as well as amd64? |
| Instance family | Does the workload need a specific family, or only CPU and memory characteristics? |
| Availability Zone | Is strict placement required by storage or resilience, or inherited from an old template? |
| Capacity type | Can the workload use Spot, or does it require On-Demand capacity? |
| Taints and tolerations | Does the isolation provide a real control or simply reduce bin packing? |
| Topology spread | Does the policy reflect the actual failure objective? |
Karpenter NodePool limits can prevent an unexpected configuration from creating unbounded capacity. Pair limits with budgets and cloud cost anomaly detection, because an autoscaler that works correctly can still respond to unintended demand.
Configure consolidation around disruption tolerance
Consolidation can replace or remove underused nodes when workloads fit more efficiently elsewhere. It saves money only when voluntary disruption is safe.
Set PodDisruptionBudgets, topology rules, termination grace periods and application shutdown behavior according to service requirements. Confirm that replicas are genuinely independent. Stateful workloads need tested storage and failover behavior. Critical platform services need capacity and topology that survive node turnover.
Use disruption budgets in Karpenter to control the rate and timing of voluntary changes. Review blocked consolidation events as useful evidence. They may reveal restrictive scheduling, a single replica, a local volume or a budget that no longer reflects service design.
Use Spot as a workload capability
Spot capacity can reduce rates significantly, but it is interruptible. Suitable workloads can retry, reschedule or checkpoint without unacceptable customer impact. Batch jobs, CI workers, stateless services with sufficient replicas and some data processing often fit. Single-replica stateful services and tightly coupled work usually do not.
Enable Karpenter interruption handling and diversify eligible instance types and zones. Keep On-Demand fallback where the service requires it. Test termination behavior rather than assuming Kubernetes will make interruption harmless.
Segment capacity by behavior. Stable critical demand may run on On-Demand capacity covered by Savings Plans. Variable fault-tolerant demand may use Spot. Uncertain or newly migrated demand can remain On-Demand until the baseline is understood.
Evaluate Graviton with application evidence
AWS Graviton instances use the arm64 architecture and can offer attractive price performance, but the migration decision belongs to the complete application. Container images, native libraries, agents and build pipelines must support the architecture.
Build multi-architecture images and test representative performance. Compare cost per successful request or job, not instance price alone. A lower hourly rate is not valuable if throughput falls or operational complexity rises.
Karpenter can provision both architectures when the workload declares compatibility. This expands choice and can improve availability. Roll out by service and retain an amd64 path until production evidence supports the change.
Manage commitments after the workload is efficient
Savings Plans can cover a stable eligible compute baseline. Buy against the floor that remains after rightsizing and autoscaling, not against peak node count. Forecast architecture changes, Graviton adoption, migrations and Spot usage before committing.
Track coverage and utilization separately. A commitment can be fully utilized while covering an inefficient fleet. The AWS Savings Plans and Reserved Instances guide explains portfolio design and governance.
Do not ignore network, storage and observability
EKS optimization often stops at compute even when other costs are material. Review orphaned volumes and snapshots, storage class and IOPS, idle load balancers, NAT paths, cross-zone traffic and public egress.
For telemetry, inspect log volume, metric cardinality, duplicate collection and retention. Keep security evidence and service-level signals. Reduce data that has no defined diagnostic, reliability or compliance purpose.
Run optimization as a controlled program
A practical sequence is:
- Enable cost allocation and identify the most expensive clusters and workloads.
- Correct missing ownership and profile representative demand.
- Rightsize requests and align HPA or VPA behavior.
- Simplify NodePools and unnecessary scheduling constraints.
- Enable consolidation with disruption controls and SLO monitoring.
- Introduce diversified Spot and Graviton where workloads qualify.
- Cover only the stable residual baseline with commitments.
- Verify realized cost, reliability and unit economics each month.
CloudForge provides EKS, AKS and GKE cost optimization consulting and AWS FinOps consulting. The objective is not the smallest cluster. It is the most economical capacity system that still meets the workload's delivery and reliability requirements.
Sources
Want this applied to your cloud environment?
Send CloudForge your requirements and the company will identify the highest-impact next step for your cost, delivery or reliability goals.
Contact CloudForge