Senior Kubernetes Platform Engineer Lead

TechDigital Group

Georgia

Hybrid

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Job summary

TechDigital Group is seeking a Senior Site Reliability Engineer (SRE) - Technical Leader to design, operate, and scale a Kubernetes-based platform for highly regulated environments, including FedRAMP High and DoD IL5. You will bridge software engineering and infrastructure to ensure resilience, observability, security, and developer velocity.

In this Georgia-based on-site role, you will lead production-grade Kubernetes implementations, drive reliability with automation, define SLOs/SLIs, design

Qualifications

  • 10+ years in SRE, DevOps, or infra engineering.
  • Hands-on production Kubernetes in regulated environments.
  • Experience with FedRAMP High and/or DoD IL5.
  • Strong cloud, Linux, and networking fundamentals.
  • Infrastructure as Code (Terraform preferred).
  • Proficient in scripting (Python/Go) and observability stacks.

Responsibilities

  • Design, build, and operate production-grade Kubernetes platforms in regulated environments.
  • Improve system reliability through automation and thoughtful design.
  • Define and drive SLOs, SLIs, and error budgets.
  • Build and evolve secure, scalable CI/CD pipelines.
  • Implement robust observability (metrics, logs, traces).
  • Collaborate with security/compliance to meet FedRAMP IL5.
  • Support ATO processes with documentation and controls.
  • Participate in on-call rotations and incident response.

Skills

Kubernetes
CI/CD
FedRAMP IL5
Observability
Python/Go scripting
Terraform
GitOps workflows

Tools

GitHub Actions
Jenkins
ArgoCD
Prometheus
Grafana
OpenTelemetry
ELK
Kubernetes (EKS/AKS/GKE)

Job description

Top Skills Required
  • Kubernetes (GKE / AKS / EKS) – Production-grade, regulated environments
  • CI/CD & Automation (GitHub Actions, Jenkins, ArgoCD, GitOps)
  • FedRAMP High / IL5 Compliance & Observability (Prometheus, Grafana, ELK, OpenTelemetry)
About the Role

We build technology that simply works—reliable, secure, and easy to use. We\'re looking for a Senior Site Reliability Engineer (SRE) - Technical Leader to help us design, operate, and scale a Kubernetes-based platform supporting highly regulated environments, including FedRAMP High and DoD IL5.

This role sits at the intersection of software engineering and infrastructure. You\'ll work closely with engineers across the stack to ensure our platform is resilient, observable, compliant, and developer-friendly—without slowing teams down.

What You\'ll Do
  • Design, build, and operate production-grade Kubernetes platforms in regulated environments
  • Improve system reliability through automation, thoughtful design, and continuous iteration
  • Define and drive SLOs, SLIs, and error budgets to guide reliability decisions
  • Build and evolve CI/CD pipelines that are secure, scalable, and easy to use
  • Implement robust observability (metrics, logs, traces) to make systems understandable and actionable
  • Reduce operational toil by automating repetitive processes and improving workflows
  • Partner with security and compliance teams to meet FedRAMP High and IL5 requirements without sacrificing developer velocity
  • Support ATO processes, including documentation, controls implementation, and audit readiness
  • Participate in on-call rotations supporting customer requests and paging alerts
  • Participate in incident response, blameless postmortems, and continuous improvement efforts
  • Help shape a platform that engineers enjoy using
What You Bring
  • 10+ years of experience in SRE, DevOps, or infrastructure engineering
  • Strong experience running Kubernetes in production (EKS, AKS, GKE, or upstream)
  • Hands-on experience working in FedRAMP High and/or DoD IL5 environments
  • Solid understanding of cloud infrastructure, Linux systems, and networking fundamentals
  • Experience with Infrastructure as Code (Terraform preferred)
  • Familiarity with CI/CD systems (GitHub Actions, GitLab CI, Jenkins, ArgoCD)
  • Proficiency in scripting or programming (Python, Go)
  • Experience building or operating observability platforms (Prometheus, Grafana, OpenTelemetry, ELK)
  • Working knowledge of compliance frameworks (e.g., NIST 800-53, STIGs, RMF)
Nice to Have
  • Experience with service mesh technologies (Istio, Linkerd)
  • Familiarity with policy-as-code (OPA/Gatekeeper, Kyverno)
  • Experience with GitOps workflows
  • Exposure to multi-cluster or hybrid cloud architectures
  • Knowledge of FIPS-compliant systems or DoD Cloud SRG
  • Relevant certifications (CKA, CKS, cloud provider certs, Security+)
How We Work
  • We value simplicity, transparency, and collaboration
  • We believe in blameless culture and learning from incidents
  • We focus on building tools and platforms that empower other engineers
  • We balance reliability, security, and developer experience—not one at the expense of the others
What Success Looks Like
  • Our platform is reliable, scalable, and easy to operate
  • Engineers can deploy confidently in high-compliance environments
  • Observability provides clear, actionable insights
  • Operational overhead is minimized through automation
  • Compliance requirements are met seamlessly as part of the platform
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE Engineer
Senior SRE Engineer

Astreya • California (MO)

On-site
USD 140,000 - 200,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Hobbsnews • Chandler (AZ), Northern (KY)

Hybrid
USD 120,000 - 180,000
Senior Software Engineer
Senior Software Engineer

CoreWeave • Livingston (NJ)

On-site
USD 140,000 - 200,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New Jersey

On-site
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior AWS Cloud Platform Analyst
Senior AWS Cloud Platform Analyst

TechDigital Group • Georgia

Hybrid
USD 120,000 - 180,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New York (NY)

Hybrid
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+1
Site Reliability Engineer
Site Reliability Engineer

Flanksource Inc. • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work
Flexible hours
Opportunity to work with cutting-edge technology
Lead Kubernetes SRE
Lead Kubernetes SRE

TechDigital Group • Minneapolis (MN)

On-site
USD 80,000 - 120,000
Senior Platform Engineer – Core Infrastructure
Senior Platform Engineer – Core Infrastructure

Jobtailor • California (MO)

On-site
USD 150,000 - 190,000