Senior Kubernetes Engineer

Strategio Inc.

New York (NY)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work model

Job summary

Strategio in the DMV/NY area seeks a Senior Kubernetes Engineer to design, deploy, and operate production-grade Amazon EKS clusters within air-gapped/private environments handling petabyte-scale workloads.

You will architect highly available Kubernetes environments, manage Karpenter-based scaling, and implement robust storage and monitoring to support thousands of concurrent jobs. The role is hybrid, requires deep expertise in Spark workloads.

Qualifications

  • Extensive hands-on experience designing and operating production Kubernetes environments.
  • Deep expertise with Amazon EKS and AWS cloud infrastructure.
  • Strong ability to troubleshoot complex performance and reliability issues in large clusters.

Responsibilities

  • Design, deploy, and operate production-grade Amazon EKS clusters.
  • Architect highly available and fault-tolerant Kubernetes environments for thousands of concurrent jobs.
  • Operate infrastructure within air-gapped private VPC environments.
  • Investigate and resolve distributed systems issues across EKS, Karpenter, and Spark.
  • Troubleshoot scheduling bottlenecks and node scaling delays under heavy workloads.
  • Define and enforce resource quotas, limits, and priority classes.
  • Configure persistent volumes with EBS and EFS CSI drivers.
  • Implement logging, monitoring, and anomaly detection to prevent failures.

Skills

Kubernetes
Amazon EKS
Karpenter
Apache Spark
Distributed systems
Troubleshooting
Resource management
Air-gapped environments
Cloud infrastructure
Spot/On-Demand strategies

Tools

Kubernetes storage
EKS tooling
AWS CLI

Job description

Senior Kubernetes Engineer

Location: DMV area, NYC, NJ - Hybrid

The Opportunity

Strategio is supporting an enterprise client in building and operating mission‑critical Amazon EKS infrastructure designed to power petabyte‑scale data processing workloads. This is an opportunity to work on highly complex distributed systems where Kubernetes performance, scalability, resiliency, and infrastructure design are critical to supporting thousands of concurrent data processing jobs.

You will take ownership of large‑scale Kubernetes environments operating within secure, air‑gapped private VPCs and solve complex engineering challenges across Amazon EKS, Karpenter, and Apache Spark. This role is ideal for an experienced Kubernetes engineer who enjoys working deep within infrastructure, troubleshooting distributed systems under heavy load, and designing highly available platforms at significant scale.

About You

You are a highly experienced Kubernetes engineer with deep technical expertise in Kubernetes internals, Amazon EKS, and large‑scale distributed systems. You have operated complex production environments where availability, performance, scalability, and fault tolerance are critical.

You are comfortable troubleshooting infrastructure issues beyond the surface level, including scheduling bottlenecks, node scaling delays, resource contention, executor failures, and cascading failures under heavy workloads. You combine strong Kubernetes engineering expertise with an understanding of large‑scale data processing environments and Apache Spark workloads.

As a Senior Kubernetes Engineer, you will:
  • Design, deploy, and operate production‑grade Amazon EKS clusters supporting petabyte‑scale data processing workloads.
  • Architect highly available and fault‑tolerant Kubernetes environments capable of supporting thousands of concurrent jobs.
  • Operate infrastructure within air‑gapped private VPC environments without internet access, including the management of secure package repositories and container registries.
  • Investigate and resolve complex distributed systems issues across Amazon EKS, Karpenter, and Apache Spark.
  • Troubleshoot scheduling bottlenecks, node scaling delays, resource contention, and cascading failures under heavy workloads.
  • Implement and optimize Karpenter consolidation and disruption policies while balancing infrastructure costs with workload resiliency.
  • Develop and manage effective Spot and On‑Demand instance strategies, including robust node interruption handling.
  • Define and enforce ResourceQuotas, LimitRanges, and PriorityClasses to ensure fair resource allocation and prevent resource starvation.
  • Configure persistent volume claims and container‑native storage using the Amazon EBS CSI and Amazon EFS CSI drivers.
  • Optimize storage architecture for resilient, high‑throughput data processing workloads.
  • Implement comprehensive logging, monitoring, alerting, and anomaly detection to proactively identify provisioning failures, executor loss, and throughput degradation.
  • Design fault‑tolerant architectures using retry strategies, checkpointing, and graceful degradation patterns to minimize re‑computation following system failures.
  • Develop, manage, and maintain infrastructure configurations within highly secure private cloud environments.
  • Continuously improve platform scalability, reliability, performance, and operational efficiency.
Core Skills
  • Extensive hands‑on experience designing, deploying, and operating production Kubernetes environments.
  • Deep expertise with Amazon EKS and AWS cloud infrastructure.
  • Strong understanding of Kubernetes internals, architecture, scheduling, resource management, and cluster operations.
  • Experience managing large‑scale Kubernetes clusters supporting high‑volume or data‑intensive workloads.
  • Hands‑on experience with Karpenter, including node provisioning, consolidation, disruption policies, and scaling strategies.
  • Experience supporting and optimizing Apache Spark workloads running on Kubernetes.
  • Strong understanding of distributed systems and the ability to troubleshoot complex performance and reliability issues.
  • Experience designing highly available and fault‑tolerant infrastructure.
  • Strong knowledge of Kubernetes resource management, including ResourceQuotas, LimitRanges, and PriorityClasses.
  • Experience implementing Spot and On‑Demand instance strategies and handling node interruptions.
  • Hands‑on experience with Kubernetes storage architecture, persistent volumes, Amazon EBS CSI, and Amazon EFS CSI drivers.
  • Experience implementing infrastructure monitoring, logging, alerting, and anomaly detection.
  • Experience operating within secure private VPC or air‑gapped environments.
  • Strong troubleshooting and root‑cause analysis capabilities.
Nice‑to‑haves
  • Certified Kubernetes Administrator (CKA), Certified Kubernetes Application Developer (CKAD), or Certified Kubernetes Security Specialist (CKS) certification.
  • AWS certifications.
  • Active contributions to open‑source Kubernetes, Karpenter, or Apache Spark projects.
  • Experience implementing FinOps practices and optimizing cloud infrastructure costs at scale.
  • Previous experience working with large‑scale data engineering or analytics platforms.

At Strategio, we are committed to creating a diverse and inclusive work environment where everyone feels valued and respected. We work hard to create a culture of kindness, honesty, and respect, and we are dedicated to creating higher‑performing organizations by addressing the lack of diversity in the workforce. Individuals seeking employment are considered without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, status as a protected veteran, or disability. You are being given the opportunity to provide the following information in order to help us comply with federal and state Equal Employment Opportunity/Affirmative Action record‑keeping, reporting, and other legal requirements.

Completing this form is voluntary, and we hope that you will choose to do so. Your answer is confidential. No one who makes hiring decisions will see it. Your decision to complete the form and your answer will not harm you in any way. If you want to learn more about the law or this form, visit the U.S. Department of Labor’s Office of Federal Contract Compliance Programs (OFCCP) website at www.dol.gov/ofccp.

I agree to the information submitted above *

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Kubernetes Engineer for Petabyte-Scale Data Ops
Senior Kubernetes Engineer for Petabyte-Scale Data Ops

Strategio Inc. • New York (NY)

Hybrid
USD 180,000 - 240,000
Hybrid work model
Sr. Solutions Architect – Kubernetes Platform
Sr. Solutions Architect – Kubernetes Platform

Gravity IT Resources • Plantation (FL)

On-site
USD 150,000 - 210,000
Referral bonus
Full-Stack Engineer
Full-Stack Engineer

Strategio Inc. • New York (NY)

On-site
USD 90,000 - 130,000
Staff Site Reliability Engineer - Kubernetes
Staff Site Reliability Engineer - Kubernetes

Segment (Twilio) • Chicago (IL), New York (NY), Bellevue (WA)

On-site
USD 174,000 - 214,000
Equity
Health insurance
Paid leave
+2
Senior DevOps Engineer
Senior DevOps Engineer

Accenture Federal Services • Colorado Springs (CO)

On-site
USD 117,000 - 244,000
Senior DevOps Engineer
Senior DevOps Engineer

Newton Research • Boston (MA)

On-site
USD 150,000 - 175,000
Equity
Competitive salary
Benefits
SRE/Devops Engineer
SRE/Devops Engineer

INSPYR Solutions • Sunnyvale (CA)

Hybrid
USD 120,000 - 180,000
Work-life balance
No on-call requirements
Standard business hours
Lead Associate Principal, Cloud Engineering
Lead Associate Principal, Cloud Engineering

The Options Clearing Corporation • Chicago (IL)

On-site
USD 143,000 - 229,000
Tuition Reimbursement
Student Loan Repayment Assistance
Technology Stipend
+3
Principal Kubernetes Solutions Architect
Principal Kubernetes Solutions Architect

Accenture Federal Services • Tampa (FL)

On-site
USD 139,000 - 171,000
Staff Site Reliability Engineer - Kubernetes
Staff Site Reliability Engineer - Kubernetes

Okta • Bellevue (WA)

On-site
USD 174,000 - 214,000