On-Prem Reliability Engineer

OpsMill

United States

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpsMill is seeking a Production Engineer / SRE to help ensure Infrahub deployments are reliable across diverse Kubernetes environments. You will act as a bridge between customer escalations and engineering, building diagnostics and automation to prevent recurring issues.

The role emphasizes deep debugging, collaboration with cross-functional teams, and delivering robust, production-grade tooling in Python, Go, or Rust.

Qualifications

  • 4-7 years of experience in production engineering, SRE, platform engineering, or similar roles.
  • Strong software engineering fundamentals including design, debugging, testing, code review.
  • Practical Kubernetes expertise to debug real deployments across distributions.
  • Deep troubleshooting instincts and observability using logs, metrics, and traces.
  • Experience with at least one of Python, Go, or Rust for tooling and product code.
  • Excellent problem decomposition and communication skills.
  • Self-directed remote work capability with strong async communication.
  • Collaborative mindset with cross-functional partnerships to drive improvements.

Responsibilities

  • Partner directly with customers and with our Solution Architecture/Customer Success teams on L2/L3 escalations across heterogeneous Kubernetes environments.
  • Drive issues to resolution by reproducing problems locally and coordinating fixes with engineering.
  • Build and maintain diagnostics tooling including support bundles, health checks, and environment validators.
  • Own the test automation infrastructure roadmap, improving CI stability and reducing flaky tests.
  • Establish and maintain performance baselines and regression tests as actionable gates.
  • Improve installation and upgrade robustness by identifying recurring failure modes and automating fixes.
  • Write production-quality code in Python, Go, or Rust for internal tooling and product improvements.
  • Close the reliability feedback loop by turning field issues into better tests and documentation.

Skills

Production engineering
SRE
Platform engineering
Kubernetes
Observability
Python
Go
Rust
Communication
Remote work
Collaboration

Job description

OpsMill is seeking a Production Engineer / SRE to help ensure Infrahub deployments are reliable across diverse Kubernetes environments. You will act as a bridge between customer escalations and engineering, building diagnostics and automation to prevent recurring issues.

The role emphasizes deep debugging, collaboration with cross-functional teams, and delivering robust, production-grade tooling in Python, Go, or Rust.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Product Reliability Engineer for Infra Automation
Remote Product Reliability Engineer for Infra Automation

OpsMill SAS • United States

On-site
USD 100,000 - 140,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Senior Platform Engineer (SRE) — Production Reliability
Senior Platform Engineer (SRE) — Production Reliability

SentiLink • United States

On-site
USD 160,000 - 210,000
Health insurance
401(k) plan with employer match
Flexible paid time off
+2
Site Reliability Engineer
Site Reliability Engineer

Triwill Group • United States

Remote
USD 120,000 - 180,000
Senior Production Engineer: Reliability Platform & SRE
Senior Production Engineer: Reliability Platform & SRE

Weights & Biases • Livingston (NJ)

On-site
USD 139,000 - 185,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
SRE – SaaS Platform Reliability (Kubernetes, CI/CD)
SRE – SaaS Platform Reliability (Kubernetes, CI/CD)

Obsidian • Palo Alto (CA)

On-site
USD 165,000 - 190,000
Equity + 401k
Healthcare coverage
Flexible PTO
+2
SRE Engineer
SRE Engineer

ALLTECH CONSULTING SVC INC • Oregon (WI)

On-site
USD 90,000 - 120,000
Service Reliability Engineer
Service Reliability Engineer

NVIDIA Corporation • United States

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000