Site Reliability Engineer

Vimo

California (MO)

On-site

USD 120,000 - 165,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401k retirement plan
Education assistance

Job summary

Vimo is seeking a Site Reliability Engineer (SRE) to ensure the reliability, availability, and performance of its production platform powering health insurance exchanges and safety-net programs for state governments. You will build automation, improve observability, respond to incidents, and reduce toil in a mission-driven environment.

The role collaborates with developers and infra engineers, involves CI/CD, Terraform, Docker/Kubernetes, and AWS, and offers growth opportunities in reliability

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience.
  • 2+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or related ops role.
  • Proficiency in at least one programming/scripting language (Python, Bash, Go).
  • Hands-on experience with AWS services (EC2, RDS, S3, VPC, CloudWatch).
  • Familiarity with Docker and Kubernetes.
  • Experience with Terraform or similar IaC tools.
  • Understanding of CI/CD concepts and at least one pipeline tool.
  • Familiarity with monitoring/observability tools (Prometheus, Grafana, CloudWatch, ELK, PagerDuty).
  • Solid Linux/Unix administration, networking basics.
  • Networking fundamentals: TCP/IP, DNS, HTTP/HTTPS, load balancing, TLS/SSL.
  • Strong troubleshooting and analytical skills, good communication.

Responsibilities

  • Monitor, maintain, and troubleshoot production services for high availability and performance.
  • Respond to production incidents; triage alerts and drive resolution.
  • Contribute to blameless postmortems and track follow-up actions.
  • Build and maintain CI/CD pipelines for safe deployments.
  • Write automation scripts (Python, Bash, or Go) to reduce manual work.
  • Improve observability with dashboards and log aggregation.
  • Manage AWS infrastructure (EC2, EKS, RDS, S3, VPC, etc.).
  • Work with Terraform and Kubernetes/EKS to provision environments.
  • Assist with capacity planning and performance analysis for peak periods.
  • Collaborate with dev teams to improve reliability and resilience.
  • Support disaster recovery with backup validation and runbooks.
  • Ensure security/compliance controls per HIPAA, FedRAMP, SOC 2.

Skills

Python
Bash
Go
AWS
Docker
Kubernetes
Terraform
CI/CD
Monitoring
Linux
Networking
Communication

Education

Bachelor's degree in CS or related field

Tools

Datadog
Prometheus
Grafana
ELK
PagerDuty

Job description

Vimo® started as the “Expedia” of health insurance and has evolved into a leader in transforming government IT infrastructure with its proven SaaS and AI technology. Our innovative approach to health insurance shopping and enrollment has expanded beyond exchanges, and we are now reinventing how states administer safety net programs such as Medicaid, SNAP (food stamps), child care, and unemployment insurance. With our cutting‑edge technology, we are helping agencies serve more people, faster, and transforming healthcare service delivery as we know it.

We are looking for a Site Reliability Engineer (SRE) to join our Vimo team.

About The Role

As a Site Reliability Engineer, you will help ensure the reliability, availability, and performance of Vimo’s production platform. Our systems power health insurance exchanges, Medicaid enrollment, and other safety‑net programs for state governments—the services you support directly impact millions of people’s access to critical benefits. You will work alongside senior SREs, developers, and infrastructure engineers to build automation, improve observability, respond to incidents, and reduce operational toil. This role is ideal for an engineer who is passionate about systems thinking, enjoys solving problems at scale, and wants to grow their career in reliability engineering within a mission‑driven environment.

Responsibilities
  • Monitor, maintain, and troubleshoot production services to ensure high availability and performance across Vimo’s SaaS platform.
  • Respond to production incidents as part of the on‑call rotation; triage alerts, coordinate with engineering teams, and drive issues to resolution.
  • Contribute to blameless postmortem processes by documenting incidents, identifying root causes, and tracking follow‑up action items.
  • Build and maintain CI/CD pipelines to support safe, repeatable, and efficient application deployments.
  • Write automation scripts and tools (Python, Bash, or Go) to reduce manual operational work and improve reliability.
  • Support and improve observability infrastructure including monitoring dashboards, log aggregation, distributed tracing, and alerting using tools such as Datadog, Prometheus, Grafana, ELK/OpenSearch, and PagerDuty.
  • Manage cloud infrastructure on AWS (EC2, EKS, RDS, S3, VPC, CloudFront, Route 53, Lambda) following established standards and best practices.
  • Work with infrastructure‑as‑code tools (Terraform) and container orchestration (Docker, Kubernetes/EKS) to provision and manage environments.
  • Assist with capacity planning, load testing, and performance analysis to prepare for peak traffic periods such as open enrollment seasons.
  • Collaborate with application development teams to improve service reliability through architecture reviews, production readiness checklists, and resilience patterns.
  • Support disaster recovery procedures including backup validation, failover testing, and documentation of recovery runbooks.
  • Help maintain compliance with security and regulatory requirements (HIPAA, FedRAMP, SOC 2) by ensuring infrastructure controls are properly implemented and documented.
  • Continuously improve on‑call processes, runbooks, and operational documentation to reduce mean time to detection (MTTD) and mean time to resolution (MTTR).
Qualifications
Basic Qualifications/Skills
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 2+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or a related operations‑focused role.
  • Proficiency in at least one programming or scripting language (Python, Bash, Go, or similar) for writing automation and tooling.
  • Hands‑on experience with AWS cloud services (EC2, RDS, S3, VPC, IAM, CloudWatch, or equivalent).
  • Familiarity with container technologies (Docker) and orchestration platforms (Kubernetes).
  • Experience with infrastructure‑as‑code tools such as Terraform, or similar.
  • Understanding of CI/CD concepts and experience with at least one pipeline tool (Jenkins, GitLab CI, GitHub Actions, or similar).
  • Familiarity with monitoring and observability tools (Loki, Prometheus, VictoriaMetrics, Grafana, CloudWatch, ELK, or PagerDuty).
  • Solid understanding of Linux/Unix systems administration, including process management, file systems, and networking basics.
  • Understanding of networking fundamentals: TCP/IP, DNS, HTTP/HTTPS, load balancing, and TLS/SSL.
  • Strong troubleshooting and analytical skills with the ability to diagnose issues across the application and infrastructure stack.
  • Good communication skills and the ability to work collaboratively in a team‑oriented environment.
Preferred Qualifications/Skills
  • Experience in healthcare technology, government IT, or benefits administration platforms.
  • Familiarity with compliance frameworks such as HIPAA, FedRAMP, or SOC 2 and their operational implications.
  • Experience with PostgreSQL, Mongo or such relational & NoSQL database systems in a production environment.
  • Exposure to incident management frameworks and blameless postmortem practices.
  • Experience with GitOps workflows and tools (ArgoCD or similar).
  • Familiarity with configuration management tools (Puppet, Ansible, Chef or similar).
  • Experience with log management and analysis at scale.
  • Exposure to load testing or performance benchmarking tools (k6, Locust, JMeter).
  • AWS certifications (Cloud Practitioner, Solutions Architect Associate, or SysOps Administrator) are a plus.
  • Familiarity with SLO/SLI concepts and error budget‑driven development practices.
Compensation and Benefits

Competitive compensation – All In range of ($120,000–$165,000). (Please note that compensation may vary based on factors such as skills, experience, performance and location.)

We offer a comprehensive benefits package, including but not limited to:

  • Health, Dental, Life, Disability, and Vision insurance
  • Healthcare spending or reimbursement accounts (HSA/FSA)
  • Retirement benefits (401k)
  • Paid time off
  • Holidays: 13 paid days per year
  • Education assistance or tuition reimbursement
  • Employee discounts for Gym memberships & commuting/travel assistance
Our Values
  • We believe that working hard, when it is imbued with purpose, can and should be fun.
  • You'll find we are a "can do" place where people work together and roll up their sleeves to get the job done.
  • Everyone has a voice; everyone's ideas count, and everyone is respected.
  • We have built a company, as well as a community of friends and colleagues, with respect for each other.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Vimo • California (MO)

On-site
USD 150,000 - 200,000
Health insurance
401k
Paid time off
+2
Site Reliability Engineer – Public Health & Government Platform
Site Reliability Engineer – Public Health & Government Platform

Vimo • California (MO)

On-site
USD 120,000 - 165,000
Health insurance
401k retirement plan
Education assistance
Product Manager
Product Manager

SupportFinity™ • Mountain View (CA)

On-site
USD 90,000 - 120,000
Health, Dental, Life, Disability, and Vision insurance
Retirement benefits (401k)
Paid time off
+3
Sr. QA Automation Engineer (SDET) AI-Enhanced Testing
Sr. QA Automation Engineer (SDET) AI-Enhanced Testing

Vimo • Mountain View (CA)

On-site
USD 150,000 - 190,000
Health insurance
401(k) plan
Paid time off
+1
QA Automation Engineer (SDET) AI-Enhanced Testing
QA Automation Engineer (SDET) AI-Enhanced Testing

GetInsured • Mountain View (CA)

On-site
USD 120,000 - 140,000
Health insurance
Dental insurance
Vision insurance
QA Lead (SDET)
QA Lead (SDET)

VIMO INC • Mountain View (CA)

On-site
USD 140,000 - 170,000
Health, Dental, Life, Disability, and Vision insurance
Retirement benefits (401k)
Paid time off
Group Product Manager
Group Product Manager

Vimo • Mountain View (CA)

On-site
USD 180,000 - 200,000
Health insurance
Dental insurance
Life insurance
+7
Data Analyst
Data Analyst

VIMO INC • Mountain View (CA)

Hybrid
USD 80,000 - 100,000
Health, dental, life insurance
401k retirement benefits
Paid time off
Security Analyst
Security Analyst

VIMO INC • Mountain View (CA)

Hybrid
USD 90,000 - 120,000
Health, dental, life, disability, and vision insurance
Health care spending accounts
Retirement benefits (401k)
+3
Group Product Manager
Group Product Manager

Vimo-Inc • Mountain View (CA)

On-site
USD 180,000 - 200,000
Health benefits