Senior Devops Engineer/lead

Scalence

Morristown (NJ)

Hybrid

USD 140,000 - 190,000

Full time

11 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Scalence is seeking a Senior DevOps Engineer to join our Infrastructure team, focusing on Kubernetes operations and automation at scale. You will own the health, scalability, and reliability of our clusters and mentor junior engineers while driving the development of tooling used by platform teams.

You will lead cluster upgrades, implement automation, and partner with service owners and leadership to improve observability and incident response across environments.

Qualifications

  • 7-9 years in DevOps, SRE, or Infra Eng.
  • Deep Kubernetes ops expertise across large multi-cluster environments.
  • Strong programming/scripting in Python, Go, Bash.
  • Advanced skills reading logs, metrics, and system state for diagnosis.
  • Extensive experience with cloud infrastructure (AWS, GCP, Azure).
  • Expert debugging skills across distributed systems.
  • Excellent written and verbal communication and ability to document and influence.
  • Mentoring engineers and leading technical initiatives.

Responsibilities

  • Own and drive the roadmap for Kubernetes cluster operations — lead upgrades, manage node groups, and maintain cluster health across environments at scale.
  • Lead troubleshooting of production and non-production issues across multiple interdependent systems.
  • Architect and lead development of internal tooling — automation (scripts, CLIs, controllers, operators).
  • Drive automation strategy for operational workflows — replace manual runbooks with scripts and tools.
  • Diagnose and resolve complex infrastructure blockers across clusters.
  • Define and improve observability strategy with logging, metrics, and alerting.
  • Set standards for issue tracking and triage; drive resolution and reporting practices.
  • Partner cross-functionally with service owners, platform teams, and leadership.
  • Lead on-call rotations and incident response; drive post-incident reviews and improvements.
  • Mentor and upskill junior and mid-level engineers.

Skills

Kubernetes operations
Automation scripting
Cloud architecture
Mentoring
Cross-team collaboration
Communication

Tools

Terraform
Helm
Ansible
Jenkins
GitHub Actions
Spinnaker
ArgoCD
Flux
Prometheus
Grafana
Datadog

Job description

ABOUT THE ROLE

We are looking for a Senior DevOps Engineer to join our Infrastructure team, focused on Kubernetes operations and automation at scale. You will drive the strategy and own the health, scalability, and reliability of our Kubernetes clusters, while architecting and leading the development of tooling that platform and service teams rely on to manage infrastructure safely and efficiently. You will work closely with service owners, platform engineers, and tooling teams — and mentor junior engineers — to keep clusters running smoothly and to automate away manual, error-prone operational work.

WHAT YOU’LL DO
  • Own and drive the roadmap for Kubernetes cluster operations — lead cluster upgrades, manage node groups (scaling, draining, replacement), and maintain overall cluster health across environments at scale.
  • Lead troubleshooting of complex production and non-production issues — use kubectl, logs, metrics, and other diagnostics to identify root cause and resolve workload, node, and networking failures, often across multiple interdependent systems.
  • Architect and lead development of internal tooling — design, implement, and maintain automation (scripts, CLIs, controllers, operators) that helps teams manage infrastructure more reliably and with less manual effort.
  • Drive automation strategy for operational workflows — replace manual runbooks with scripts and tools that handle upgrades, remediation, scaling, and routine maintenance.
  • Diagnose and resolve complex infrastructure blockers — debug deployment failures, node/pod scheduling issues, resource constraints, and misconfigurations across clusters.
  • Define and improve observability strategy — instrument logging, metrics, and alerting to increase visibility into cluster health and reduce time-to-detect/resolve.
  • Set standards for issue tracking and triage — author detailed bug reports capturing root cause, repro steps, and impact; drive issues to resolution and improve team-wide reporting practices.
  • Partner cross-functionally with service owners, platform teams, and leadership — align on operational requirements, capacity planning, and upgrade/maintenance schedules.
  • Lead on-call rotations and incident response, own post-incident reviews, and drive continuous improvement initiatives across the infrastructure org.
  • Mentor and upskill junior and mid-level engineers, providing technical guidance and code/design reviews.
WHAT WE’RE LOOKING FOR
Required
  • 7-9 years of DevOps, SRE, or Infrastructure Engineering experience.
  • Deep, hands‑on Kubernetes operations expertise — cluster upgrades, node group management, and complex troubleshooting across large‑scale, multi‑cluster environments.
  • Strong programming/scripting skills (Python, Go, Bash, or similar) with a proven track record of building, scaling, and maintaining automation and internal tooling used by multiple teams.
  • Advanced skills in reading and interpreting logs, metrics, and system state to diagnose complex, cross‑system infrastructure issues.
  • Extensive experience with cloud infrastructure (AWS, GCP, or Azure), including architecture decisions and cost/performance tradeoffs.
  • Expert‑level debugging skills across distributed systems — able to trace failures from symptom through to root cause in highly complex environments.
  • Excellent written and verbal communication — able to produce clear status updates, bug reports, technical documentation, and influence technical direction across teams.
  • Demonstrated experience mentoring engineers and leading technical initiatives.
Preferred
  • Deep experience with infrastructure‑as‑code tools (Terraform, Helm, Ansible), including designing reusable modules/patterns for org‑wide use
  • Strong familiarity with CI/CD systems (Jenkins, GitHub Actions, Spinnaker, or similar), including pipeline architecture Hands‑on experience with GitOps workflows (ArgoCD, Flux) at scale
  • Advanced experience with observability stacks (Prometheus, Grafana, Datadog), including designing alerting/SLO frameworks
  • Proven experience operating Kubernetes at scale across multiple clusters, regions, or environments , including capacity planning and disaster recovery
  • Experience contributing to or leading architectural decisions for infrastructure platforms
TECHNOLOGIES YOU’LL WORK WITH
JD Focus on:

AWS, Linux, Kubernetes, python – individual go-getter, not a follower

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Devops Lead
Devops Lead

Zealogics.com • West Carteret (NJ)

Remote
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Flanksource Inc. • Mission (KS)

On-site
USD 90,000 - 130,000
100% remote work
Flexible hours
Opportunity to work with cutting-edge technology
Platform Engineering Manager
Platform Engineering Manager

Good co India • United States

Remote
USD 140,000 - 200,000
Senior DevOps Engineer
Senior DevOps Engineer

PayZen • United States

Hybrid
USD 140,000 - 190,000
Hybrid schedule: In-office Mon Tue Thu
Remote Wed and Fri
Discretionary time-off
+2
Manager, DevOps
Manager, DevOps

1 O.C. Tanner Company • Salt Lake City (UT)

On-site
USD 120,000 - 150,000
DevOps Lead
DevOps Lead

Resolve Tech Solutions, LLC • Town of Texas (WI)

On-site
USD 120,000 - 150,000
DevOps Engineer DevOps Engineer
DevOps Engineer DevOps Engineer

Kurai • Seattle (WA)

On-site
USD 120,000 - 180,000
Sr DevOps Platform Engineer - Bilingual
Sr DevOps Platform Engineer - Bilingual

CNX • United States

Remote
USD 130,000 - 190,000
DevOps Engineer / Architect
DevOps Engineer / Architect

MM International, LLC • Austin (TX)

On-site
USD 120,000 - 150,000
Lead Devops Platform Engineer
Lead Devops Platform Engineer

Techblocks • New York (NY)

Hybrid
USD 140,000 - 200,000