Site Reliability Engineer

NETFORCE Group

United States

On-site

USD 90,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

10 days of paid vacation
5 days of sick leave
Opportunity for relocation to office countries

Job summary

A technology-driven healthcare logistics provider in the United States is seeking a Cloud Infrastructure Specialist to automate and enhance infrastructure processes. The role involves troubleshooting production incidents, managing service level objectives, and supporting multiple development teams. Ideal candidates will have strong experience in GCP, Terraform, and Kubernetes, and thrive in a collaborative and remote-friendly environment. Benefits include 10 days of paid vacation, sick leave, and a focus on US market challenges.

Qualifications

  • Experience with cloud-native infrastructure and automation.
  • Proven competence in troubleshooting production incidents.
  • Strong communication skills and a collaborative mindset.

Responsibilities

  • Automate infrastructure processes to eliminate toil.
  • Troubleshoot production incidents and define resolutions.
  • Manage SLOs and error budgets for production services.
  • Guide developers to maintain shipping confidence.

Skills

GCP
Terraform
Kubernetes
Go (Golang)
Prometheus
Grafana
SQL / NoSQL
Helm
eBPF
Vault
OpsGenie
PagerDuty

Job description

Overview

A technology-driven Non-Emergency Medical Transportation (NEMT) service focused on providing patients with safe, reliable, and timely rides to medical appointments. The goal is to eliminate transportation barriers in accessing healthcare. The team works in a collaborative, remote-friendly environment with hours aligned to the US Eastern Time zone. The focus is on developing and supporting systems for dispatch, scheduling, and driver coordination — directly impacting people’s daily lives. A great fit for those who want to use technology to solve real challenges in healthcare logistics.

Technical Requirements
  • Platform & Infrastructure GCP, Terraform (IaC), Kubernetes, Service Mesh (Istio / Linkerd), Linux distributed systems at scale
  • Language & Observability Go (Golang), Prometheus / Grafana, Distributed tracing, SQL / NoSQL at scale
  • Competencies Cloud native mindset · Automate Everything · Speaks developers language · Fantastic communication · SLO / error quota ownership
  • Nice to Have Helm, eBPF, Vault / Crossplane, OpsGenie / PagerDuty
Responsibilities
  • Eliminate toil through automation, re-architecting, and refactoring — not just patching symptoms.
  • Approach every incident with an "Automate Everything" mindset so the same problem never fires twice.
  • Pair with software engineers to troubleshoot and resolve production incidents down to root cause.
  • Drive complex infrastructure changes with full transparency, clear communication, and zero downtime.
  • Design and implement self-healing, reliable, and scalable infrastructure in a cloud-native environment.
  • Guide and unblock developers across multiple teams so they can keep shipping with confidence.
  • Define SLOs and error quotas for production services; own and manage the error budget.
  • Own the GitOps workflow via ArgoCD — every deployment is Git-defined, automated, and reproducible.
  • Write or review postmortems after incidents; track corrective actions to completion. Participate in the follow-the-sun on-call rota and actively champion our DevOps culture.
  • We’re currently looking for someone who is comfortable working evening hours and can be reliably online from 6:00 PM to 2:00 AM.
What does an average day look like?

You’ll proactively support production workloads, troubleshoot issues to their root cause, and write or review postmortems once incidents are resolved. You’ll continuously identify weaknesses in infrastructure and observability and feed them into the improvement backlog.

Benefits
  • Opportunity to work on a live product focused on the US market
  • 10 days of paid vacation and 5 days of sick leave
  • Possibility of relocation to countries where we have offices
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Fullstack Software Engineer
Sr. Fullstack Software Engineer

American Logistics • United States

Hybrid
USD 100,000 - 130,000
Product Engineering Team - Software Engineer (Mid-Level)
Product Engineering Team - Software Engineer (Mid-Level)

Sprinter Health • Menlo Park (CA)

Hybrid
USD 180,000 - 240,000
Meaningful pre‑IPO equity
Medical, dental, and vision plans
Flexible PTO + holidays
+4
Product Engineering Team - Software Engineer (Mid-Level)
Product Engineering Team - Software Engineer (Mid-Level)

Sprinter Health • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Meaningful pre‑IPO equity
Medical, dental, and vision plans 100%
Flexible PTO + 10 paid holidays per yr
+5
Product Engineering Team - Software Engineer (Senior)
Product Engineering Team - Software Engineer (Senior)

Sprinter Health • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Meaningful pre-IPO equity
Medical, dental, and vision plans 100%
Flexible PTO
+6
Product Engineering Team - Software Engineer (Senior)
Product Engineering Team - Software Engineer (Senior)

Sprinter Health • Menlo Park (CA)

Hybrid
USD 140,000 - 210,000
Equity
Medical, dental, and vision plans
Flexible PTO
+4
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Kontakt.io • New York (NY)

Hybrid
USD 200,000 - 250,000
Hybrid work 3 days/week in NYC office.
Equity in a high-growth company
Health, dental, vision insurance
+1
Site Reliability Engineer — Remote, Evening Hours, Healthcare
Site Reliability Engineer — Remote, Evening Hours, Healthcare

NETFORCE Group • United States

On-site
USD 90,000 - 120,000
10 days of paid vacation
5 days of sick leave
Opportunity for relocation to office countries
Staff Software Engineer
Staff Software Engineer

Metriport • United States

Hybrid
USD 180,000 - 260,000
Competitive equity
Health insurance
401(k) matching
+3
Staff Software Engineer
Staff Software Engineer

Metriport • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Competitive equity and compensation
Full health insurance (family)
401(k) plan + matching
+3
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

Remote
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3