Site Reliability Engineer

Evlo AI

Minneapolis (MN)

On-site

USD 120,000 - 180,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Evlo AI in Minneapolis seeks a Site Reliability Engineer to own the reliability, scalability, and security of distributed infrastructure powering millions of users globally.

You will collaborate with platform architects and software teams to design resilient cloud architectures, automate deployments, and minimize downtime while driving security best practices and robust monitoring. This role emphasizes hands-on engineering and operational excellence.

Qualifications

  • 3–6 years of experience in Site Reliability Engineering, DevOps, or systems engineering in cloud-native environments.
  • Strong hands-on experience with Kubernetes, Docker, and container orchestration at scale.
  • Proficiency in infrastructure automation and configuration management using Terraform and Ansible.
  • Deep understanding of Linux systems internals, networking fundamentals (TCP/IP, DNS, TLS), and load balancing.
  • Bonus: Experience with service meshes like Istio, chaos engineering practices, and software development proficiency in Go or Python

Responsibilities

  • Design, build, and maintain cloud infrastructure on AWS and Kubernetes using Terraform and Infrastructure as Code best practices
  • Define and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and comprehensive monitoring metrics using Prometheus, Grafana, and Datadog
  • Lead incident response and root cause analysis (RCA) for production outages, implementing automated remediation to prevent recurrence
  • Optimize cloud resource utilization, compute performance, and infrastructure security posture across all environments
  • Build and scale CI/CD deployment pipelines using GitHub Actions or ArgoCD to support rapid, safe software releases

Skills

Kubernetes
Docker
Linux
Networking
Go
Python

Tools

Terraform
Ansible
Prometheus
Grafana
Datadog
GitHub Actions
ArgoCD

Job description

About The Role

The role owns the reliability, scalability, and security of distributed infrastructure powering high-traffic production systems serving millions of users globally.

You will work alongside platform architects and software engineering teams to design resilient cloud architectures, automate deployment pipelines, and minimize system downtime.

Key Responsibilities
  • Design, build, and maintain cloud infrastructure on AWS and Kubernetes using Terraform and Infrastructure as Code best practices
  • Define and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and comprehensive monitoring metrics using Prometheus, Grafana, and Datadog
  • Lead incident response and root cause analysis (RCA) for production outages, implementing automated remediation to prevent recurrence
  • Optimize cloud resource utilization, compute performance, and infrastructure security posture across all environments
  • Build and scale CI/CD deployment pipelines using GitHub Actions or ArgoCD to support rapid, safe software releases
What We Are Looking For
  • 3–6 years of experience in Site Reliability Engineering, DevOps, or systems engineering in cloud-native environments
  • Strong hands-on experience with Kubernetes, Docker, and container orchestration at scale
  • Proficiency in infrastructure automation and configuration management using Terraform and Ansible
  • Deep understanding of Linux systems internals, networking fundamentals (TCP/IP, DNS, TLS), and load balancing
  • Bonus: Experience with service meshes like Istio, chaos engineering practices, and software development proficiency in Go or Python
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Evlo AI • San Francisco (CA)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

NextGen | GTA: A Kelly Telecom Company • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

JobCubby • Barrington (RI), Northern (KY)

On-site
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Flanksource Inc. • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work
Flexible hours
Opportunity to work with cutting-edge technology
DevOps Engineer DevOps Engineer
DevOps Engineer DevOps Engineer

Kurai • Seattle (WA)

On-site
USD 120,000 - 180,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000