Senior Site Reliability Engineer, Forward Deployed - Remote USA ONLY

Ardan Labs

United States

Remote

USD 140,000 - 210,000

Full time

46 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Ardan Labs is seeking a Senior Site Reliability Engineer to join a platform team and take ownership of reliability for an internal developer platform and for the experiences of application teams building on it.

This hands-on senior individual contributor role focuses on operating Kubernetes at scale, solving complex infrastructure problems, and working directly with teams to diagnose and resolve issues, including HIPAA-compliant security hygiene and patching as ongoing priorities.

Qualifications

  • >=8 years in infrastructure, platform engineering or SRE
  • >3 years production Kubernetes experience
  • >experience with GitOps in production
  • Terraform across cloud accounts or environments
  • Understanding of cloud identity and resource policies
  • Ability to lead technical work without direct authority
  • Strong troubleshooting and communications skills
  • Comfort working with application teams rather than behind the scenes

Responsibilities

  • Own health and reliability of Kubernetes management and workload clusters across sandbox, development, QA, and production environments
  • Diagnose and permanently resolve GitOps delivery and reconciliation failures
  • Manage version and component adoption to ensure fixes/deployments across clusters
  • Develop and maintain Terraform and infrastructure pipelines across a multi-account AWS environment
  • Work with cloud identity, resource policies, and key management to resolve permission issues
  • Partner with application teams to troubleshoot manifests, secrets, certificates, ingress, and other platform issues
  • Build documentation, guides, and guardrails to enable self-sufficiency
  • Own and improve telemetry path into the observability platform
  • Partner with security, networking, vendors, and other teams during incidents
  • Participate in incident response and drive durable resolutions
  • Lead technical initiatives across teams without direct reporting lines

Skills

8+ years experience
production Kubernetes
GitOps
Terraform
cloud identity & policies
lead without authority
troubleshooting & systems thinking
cross-team communication

Tools

Terraform

Job description

We are looking for a Senior Site Reliability Engineer to join a platform team and take ownership of both the reliability of an internal developer platform and the experience of the application teams building on it.

This is a hands-on senior individual contributor role for an engineer who enjoys solving complex infrastructure problems, operating Kubernetes at scale, and working directly with teams to diagnose and resolve issues.

You will work across Amazon EKS and Red Hat OpenShift on AWS, a curated Helm chart catalog, GitOps pipelines, Terraform, and a multi-account AWS environment built on Control Tower. Our workloads operate under HIPAA requirements, making security hygiene, patch currency, and operational reliability ongoing engineering priorities.

Approximately 60% of your time will focus on reliability engineering and 40% on forward-deployed work with application, security, network, and other technical teams.

What You'll Do
  • Own the health and reliability of Kubernetes management and workload clusters across sandbox, development, QA, and production environments.
  • Diagnose and permanently resolve GitOps delivery and reconciliation failures.
  • Manage version and component adoption to ensure fixes and improvements are consistently deployed across clusters.
  • Develop and maintain Terraform and infrastructure pipelines across a multi-account AWS environment.
  • Work with cloud identity, resource policies, and key management to diagnose and resolve complex permission issues.
  • Partner directly with application teams to troubleshoot manifests, managed resource claims, secrets, certificates, ingress, and other platform issues.
  • Build documentation, guides, and guardrails that help application teams become increasingly self-sufficient.
  • Own and improve the telemetry path into the observability platform.
  • Partner with security, networking, vendors, and other technical teams when the platform is involved in an incident.
  • Participate in incident response and drive problems through to durable resolution.
  • Lead technical initiatives across teams that do not report to you.
What We're Looking For
  • 8+ years of experience in infrastructure, platform engineering, or site reliability engineering.
  • At least 3 years of production Kubernetes experience supporting teams beyond your own.
  • Strong hands-on experience with production GitOps, including diagnosing conflicts between declared and live state.
  • Experience using Terraform across multiple cloud accounts or environments.
  • Deep understanding of cloud identity and resource policies, including troubleshooting permissions that appear correct but are still denied.
  • Demonstrated ability to lead technical work and influence teams without direct authority.
  • Strong troubleshooting, systems thinking, and communication skills.
  • Comfortable working directly with application and engineering teams rather than operating solely behind the scenes.
Nice to Have
  • Red Hat OpenShift / ROSA experience
  • Experience working in regulated environments
  • Production experience with Istio or another service mesh
  • Forward-deployed, field engineering, implementation, solutions, or customer-facing engineering experience
  • Ownership of OpenTelemetry, Prometheus, or commercial APM platforms
  • Go or another compiled language used for infrastructure/tooling
  • Published technical writing, such as incident reports, design documents, or technical/customer-facing documentation
Why This Role Is Different

This isn't a role where success is measured by how many tools you know. The platform is highly automated and largely self-healing. The difficult problems are often about figuring out where the problem actually belongs, determining the right durable fix, and building agreement between teams that don't report to one another.

We're looking for an engineer who can operate at both levels: deep technical infrastructure expertise and strong cross-team engagement.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Remote Senior Site Reliability Lead - AWS, Kubernetes, CI/CD
Remote Senior Site Reliability Lead - AWS, Kubernetes, CI/CD

Empower • United States

On-site
USD 114,000 - 166,000
401(k) with company matching
Tuition reimbursement
Paid volunteer time
+1
Site Reliability Engineer
Site Reliability Engineer

Veritas Search Group • Tustin (CA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

Hybrid
USD 140,000 - 170,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Hobbsnews • Chandler (AZ), Northern (KY)

Hybrid
USD 120,000 - 180,000
Senior Site Reliability Engineer, AI Agents & Automation
Senior Site Reliability Engineer, AI Agents & Automation

ServiceTitan • United States

On-site
USD 140,000 - 190,000
Flexible time off
Fully paid medical, dental, and vision
HSA/FSA programs
+7