Senior Site Reliability Engineer

Megaport

Abbeyville (CO)

On-site

USD 130,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Contractor (PJ)
Paid Time Off
Competitive Compensation
Wellhub (former Gympass)
Annual Bonus
Flexible work hours
Growth opportunities

Job summary

Latitude.sh is seeking a Senior Site Reliability Engineer to focus on building reliable, observable, and self-healing systems at scale for our bare metal cloud platform.

You will design and implement tools to automate operations, improve incident response, and enhance system observability, ensuring the platform can handle customer workloads with high availability. This role emphasizes reliability, automation, and cloud-native security practices.

Qualifications

  • Must have strong experience with Linux/Unix production environments.
  • Experience designing and operating Kubernetes-based systems.
  • Proficiency with Terraform and Ansible for automation.
  • Experience with observability stacks (Prometheus, Grafana, Loki, ELK).
  • Solid scripting knowledge (Bash, Python, Go, or Ruby).
  • Working knowledge of Git and CI/CD pipelines.
  • Strong incident management and RCA skills.
  • Knowledge of cloud-native reliability and security best practices.
  • Excellent written and spoken English communication.

Responsibilities

  • Continuously improve platform reliability and performance.
  • Design, build, and maintain automation tools for operations and incident response.
  • Implement observability solutions: monitoring, alerting, tracing.
  • Collaborate with engineering and platform teams to design scalable systems.
  • Participate in on-call rotations and lead post-incident reviews.
  • Develop and document runbooks and processes for operational excellence.
  • Contribute to SLOs/SLIs and reliability metric adoption.

Skills

English communication
Linux/Unix production
Kubernetes
Infrastructure automation
Observability stacks
Scripting languages
Git & CI/CD
Incident management
Cloud-native reliability & security

Tools

Terraform
Ansible

Job description

About Latitude.sh

Latitude.sh's global computing platform was launched in 2019, enabling businesses to programmatically deploy single-tenant Bare Metal instances in different parts of the world. We are a team of passionate individuals about hardware, software, and network infrastructure looking to build the fastest, easiest-to-use, developer-centric single-tenant Cloud infrastructure. If you share this passion, join our growing team of talented people and help build the future of the Internet.

Why Latitude.sh?

We're a lean, agile team of passionate professionals who believe in the power of innovation and creative problem-solving. As part of our team, you won't be lost in the crowd – you'll be an essential contributor, making a real impact from day one.

Our values at Latitude.sh guide us in all our work and partnerships. We're proud to be an inclusive company, and we welcome all applicants for our open positions, regardless of their background, religion, sexual orientation, gender identity, age, nationality, or disability. If these values speak to you, we'd love for you to become a part of our team.

The Role

At Latitude.sh, the Reliability team is responsible for the health and resilience of the infrastructure that powers our global bare metal cloud. As a Senior Site Reliability Engineer (SRE), you’ll focus on building reliable, observable, and self-healing systems at scale.

SREs at Latitude.sh work at the intersection of software engineering and infrastructure. You’ll design and implement tools that automate operations, improve incident response, and enhance system observability—ensuring our platform is always ready for the workloads of our customers.

This might be a good opportunity if you’re passionate about reliability, automation, and creating cloud-like experiences for bare metal infrastructure.

What You'll Be Doing
  • Continuously improve Latitude.sh’s platform reliability and performance
  • Design, build, and maintain tools to automate operational tasks and incident response
  • Implement and improve observability solutions, including monitoring, alerting, and tracing
  • Collaborate with engineering and platform teams to design scalable and resilient systems
  • Participate in on-call rotations and lead post-incident reviews with a focus on learning
  • Develop and document processes and runbooks that ensure operational excellence
  • Contribute to SLOs/SLIs definition and reliability metrics adoption across teams
What We're Looking For
  • Strong verbal and written English communication skills
  • Advanced knowledge of Linux/Unix systems in production environments
  • Experience with Kubernetes and container orchestration
  • Proficiency with infrastructure automation tools (e.g., Terraform, Ansible)
  • Experience with observability stacks (e.g., Prometheus, Grafana, Loki, ELK)
  • Familiarity with scripting and programming languages such as Bash, Python, Go, or Ruby
  • Working knowledge of Git and CI/CD pipelines
  • Solid understanding of incident management and root cause analysis processes
  • Knowledge of cloud-native reliability and security best practices
What We Offer
  • Contractor (PJ)
  • Paid Time Off
  • Competitive Compensation
  • Wellhub (former Gympass)
  • Annual Bonus based on company and team performance
  • Flexible work hours
  • Opportunities for professional growth and development
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Latitude.sh • United States

Remote
USD 120,000 - 150,000
Paid Time Off
Competitive Compensation
Wellhub (former Gympass)
+2
Network Engineer
Network Engineer

Megaport • Abbeyville (CO)

On-site
USD 110,000 - 170,000
Contractor (PJ)
Paid Time Off
Competitive Compensation
+5
Senior SRE: Scale Reliability & Observability
Senior SRE: Scale Reliability & Observability

Megaport • Abbeyville (CO)

On-site
USD 130,000 - 190,000
Contractor (PJ)
Paid Time Off
Competitive Compensation
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Latitude AI • Palo Alto (CA)

On-site
USD 179,000 - 269,000
Health insurance
401(k) match
Unlimited vacation
+2
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

Remote
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Oowlish • United States

Remote
MXN 1,222,000 - 1,573,000
Home office setup
Competitive compensation
Career growth plans
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New York (NY)

Hybrid
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+1
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New Jersey

On-site
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000