Senior Site Reliability Engineer

Vimo

California (MO)

On-site

USD 150,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401k
Paid time off
Holidays
Education assistance

Job summary

Vimo is seeking a Senior Site Reliability Engineer to drive reliability for its production platform powering health insurance exchanges and safety-net programs.

You will design automation, enforce SLOs/SLIs, and lead incident response while collaborating with development, security, and infrastructure teams to reduce toil and improve resilience.

Qualifications

  • 5+ years in Site Reliability Engineering, DevOps, or related roles.
  • Strong programming in Python, Go, Java, or Bash.
  • Deep AWS cloud services experience (EC2, EKS, RDS, S3, VPC, Lambda).
  • Proficient with Docker and Kubernetes/EKS; Terraform IaC.
  • Solid CI/CD knowledge and tooling; observability stack expertise.
  • Networking fundamentals and Linux/Unix troubleshooting.
  • On-call incident management experience.

Responsibilities

  • Design, build, and maintain tooling and automation for high availability.
  • Define and monitor SLOs/SLIs and error budgets for critical services.
  • Build and improve CI/CD pipelines with automated rollback.
  • Develop observability across monitoring, logging, tracing, alerting.
  • Lead incident response and blameless postmortems with follow‑ups.
  • Automate toil reduction; implement self‑healing, runbook remediation.
  • Manage AWS infrastructure with cost efficiency and security focus.
  • Implement IaC via Terraform; manage Kubernetes/EKS orchestration.
  • Plan capacity and conduct load testing for peak enrollment periods.
  • Collaborate on architecture reviews and disaster recovery planning.

Skills

Strong programming skills (Python, Go,
Java
Bash
AWS cloud services
Kubernetes/EKS
Terraform / IaC
CI/CD tooling (Jenkins, GitLab CI, Git
GitHub Actions, ArgoCD
Observability stack (Loki, Prometheus,
Grafana, ELK/OpenSearch, PagerDuty
Linux/Unix administration

Education

Bachelor’s degree in Computer Science, Engineering, or related field
Equivalent practical experience

Tools

Docker
Kubernetes/EKS
Terraform
Jenkins
GitLab CI
GitHub Actions
ArgoCD
Prometheus
Grafana
Loki
OpenSearch/ELK
PagerDuty

Job description

Vimo® started as the “Expedia” of health insurance and has evolved into a leader in transforming government IT infrastructure with its proven SaaS and AI technology. Our innovative approach to health insurance shopping and enrollment has expanded beyond exchanges, and we are now reinventing how states administer safety net programs such as Medicaid, SNAP (food stamps), child care, and unemployment insurance. With our cutting‑edge technology, we are helping agencies serve more people, faster, and transforming healthcare service delivery as we know it.

We are looking for a Senior Site Reliability Engineer (SRE) to join our Vimo team.
About The Role

As a Senior Site Reliability Engineer, you will be at the intersection of software engineering and systems operations, ensuring that Vimo’s production platform is reliable, performant, and scalable. Our systems power health insurance exchanges, Medicaid enrollment, and other safety‑net programs for state governments—meaning the services you keep running directly affect millions of people’s access to critical benefits. You will design and build automation, define and enforce SLOs, respond to and learn from incidents, and continuously reduce operational toil. You’ll work closely with development, infrastructure, and security teams to embed reliability into every layer of the stack.

Responsibilities
  • Design, build, and maintain the tooling, automation, and infrastructure that keeps Vimo’s production services highly available and performant.
  • Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical services.
  • Build and improve CI/CD pipelines to enable safe, fast, and repeatable deployments with automated rollback capabilities.
  • Develop and operate comprehensive observability solutions—monitoring, logging, tracing, and alerting—using tools such as Loki, Prometheus, Grafana, ELK/OpenSearch, and PagerDuty.
  • Lead incident response efforts: triage production issues in real time, coordinate cross‑team resolution, and author thorough blameless postmortems with actionable follow‑ups.
  • Identify and eliminate toil through automation; build self‑healing mechanisms and runbook‑driven remediation.
  • Manage and optimize cloud infrastructure on AWS (EC2, EKS, RDS, S3, VPC, CloudFront, Route 53, Lambda) with a focus on automation, cost efficiency and security.
  • Implement and maintain infrastructure‑as‑code using Terraform, and manage container orchestration with Kubernetes/EKS.
  • Perform capacity planning and load testing to ensure systems can handle peak enrollment periods and traffic surges.
  • Collaborate with application engineering teams on architecture reviews, resilience patterns (circuit breakers, retries, graceful degradation), and production readiness reviews.
  • Contribute to disaster recovery planning and testing, including automated failover and multi‑region strategies.
  • Support compliance and security requirements (HIPAA, FedRAMP, SOC 2) by ensuring infrastructure controls are in place and auditable.
  • Participate in a 24/7 on‑call rotation and continuously improve on‑call processes to reduce alert fatigue and mean time to resolution (MTTR).
Qualifications
Basic Qualifications/Skills
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a related systems‑focused role.
  • Strong software engineering skills in at least one language (Python, Go, Java, or Bash) with the ability to write production‑quality automation and tooling.
  • Deep hands‑on experience with AWS cloud services (EC2, EKS, RDS, S3, VPC, IAM, Lambda, CloudWatch).
  • Proficiency with container technologies (Docker) and orchestration platforms (Kubernetes/EKS).
  • Solid experience with infrastructure‑as‑code tools, particularly Terraform.
  • Strong understanding of CI/CD principles and tools (Jenkins, GitLab CI, GitHub Actions, ArgoCD, or similar).
  • Experience with observability and monitoring platforms (Loki, Prometheus, VictoriaMetrics, Grafana, ELK/OpenSearch, PagerDuty).
  • Solid understanding of networking fundamentals: TCP/IP, DNS, load balancing, CDN, TLS/SSL, and firewall configuration.
  • Experience with incident management processes, on‑call rotations, and blameless postmortem culture.
  • Strong Linux/Unix systems administration and troubleshooting skills.
Preferred Qualifications/Skills
  • Experience in healthcare technology, government IT, or benefits administration platforms.
  • Familiarity with compliance frameworks such as HIPAA, FedRAMP, or SOC 2 and their impact on infrastructure operations.
  • Experience with PostgreSQL, mongo, mysql administration, performance tuning, and high‑availability configurations.
  • Hands‑on experience with chaos engineering practices and tools (Gremlin, Litmus, or equivalent).
  • Experience with GitOps workflows and tools (ArgoCD, Flux).
  • Knowledge of service mesh technologies (Istio, Linkerd) and API gateway patterns.
  • Experience with configuration management tools (Ansible, Chef, or Puppet).
  • Familiarity with FinOps principles and cloud cost optimization strategies.
  • Experience with load testing and performance benchmarking tools (k6, Locust, JMeter).
  • Strong Plus:
  • AWS certifications (Solutions Architect, DevOps Engineer, or SysOps Administrator)
  • Experience with AWS security implementation through IaC / automation
Compensation and Benefits

Competitive compensation - All In range of ($150,000-$200,000). (Please note that compensation may vary based on factors such as skills, experience, performance and location.)

We offer a comprehensive benefits package, including but not limited to:

  • Health, Dental, Life, Disability, and Vision insurance
  • Healthcare spending or reimbursement accounts (HSA/FSA)
  • Retirement benefits (401k)
  • Paid time off
  • Holidays: 13 paid days per year
  • Education assistance or tuition reimbursement
  • Employee discounts for Gym memberships & commuting/travel assistance
Our Values
  • We believe that working hard, when it is imbued with purpose, can and should be fun.
  • You’ll find we are a "can do" place where people work together and roll up their sleeves to get the job done.
  • Everyone has a voice; everyone’s ideas count, and everyone is respected.
  • We have built a company, as well as a community of friends and colleagues, with respect for each other.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Vimo • California (MO)

On-site
USD 120,000 - 165,000
Health insurance
401k retirement plan
Education assistance
Product Manager
Product Manager

SupportFinity™ • Mountain View (CA)

On-site
USD 90,000 - 120,000
Health, Dental, Life, Disability, and Vision insurance
Retirement benefits (401k)
Paid time off
+3
QA Automation Engineer (SDET) AI-Enhanced Testing
QA Automation Engineer (SDET) AI-Enhanced Testing

GetInsured • Mountain View (CA)

On-site
USD 120,000 - 140,000
Health insurance
Dental insurance
Vision insurance
Sr. QA Automation Engineer (SDET) AI-Enhanced Testing
Sr. QA Automation Engineer (SDET) AI-Enhanced Testing

Vimo • Mountain View (CA)

On-site
USD 150,000 - 190,000
Health insurance
401(k) plan
Paid time off
+1
QA Lead (SDET)
QA Lead (SDET)

VIMO INC • Mountain View (CA)

On-site
USD 140,000 - 170,000
Health, Dental, Life, Disability, and Vision insurance
Retirement benefits (401k)
Paid time off
Site Reliability Engineer – Public Health & Government Platform
Site Reliability Engineer – Public Health & Government Platform

Vimo • California (MO)

On-site
USD 120,000 - 165,000
Health insurance
401k retirement plan
Education assistance
Group Product Manager
Group Product Manager

Vimo • Mountain View (CA)

On-site
USD 180,000 - 200,000
Health insurance
Dental insurance
Life insurance
+7
Group Product Manager
Group Product Manager

Vimo-Inc • Mountain View (CA)

On-site
USD 180,000 - 200,000
Health benefits
Security Analyst
Security Analyst

VIMO INC • Mountain View (CA)

Hybrid
USD 90,000 - 120,000
Health, dental, life, disability, and vision insurance
Health care spending accounts
Retirement benefits (401k)
+3
Data Analyst
Data Analyst

VIMO INC • Mountain View (CA)

Hybrid
USD 80,000 - 100,000
Health, dental, life insurance
401k retirement benefits
Paid time off