Cloud SRE: Build Resilient, Secure, Scalable Infra

Evlo AI

Minneapolis (MN)

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Evlo AI in Minneapolis seeks a Site Reliability Engineer to own the reliability, scalability, and security of distributed infrastructure powering millions of users globally.

You will collaborate with platform architects and software teams to design resilient cloud architectures, automate deployments, and minimize downtime while driving security best practices and robust monitoring. This role emphasizes hands-on engineering and operational excellence.

Qualifications

  • 3–6 years of experience in Site Reliability Engineering, DevOps, or systems engineering in cloud-native environments.
  • Strong hands-on experience with Kubernetes, Docker, and container orchestration at scale.
  • Proficiency in infrastructure automation and configuration management using Terraform and Ansible.
  • Deep understanding of Linux systems internals, networking fundamentals (TCP/IP, DNS, TLS), and load balancing.
  • Bonus: Experience with service meshes like Istio, chaos engineering practices, and software development proficiency in Go or Python

Responsibilities

  • Design, build, and maintain cloud infrastructure on AWS and Kubernetes using Terraform and Infrastructure as Code best practices
  • Define and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and comprehensive monitoring metrics using Prometheus, Grafana, and Datadog
  • Lead incident response and root cause analysis (RCA) for production outages, implementing automated remediation to prevent recurrence
  • Optimize cloud resource utilization, compute performance, and infrastructure security posture across all environments
  • Build and scale CI/CD deployment pipelines using GitHub Actions or ArgoCD to support rapid, safe software releases

Skills

Kubernetes
Docker
Linux
Networking
Go
Python

Tools

Terraform
Ansible
Prometheus
Grafana
Datadog
GitHub Actions
ArgoCD

Job description

Evlo AI in Minneapolis seeks a Site Reliability Engineer to own the reliability, scalability, and security of distributed infrastructure powering millions of users globally.

You will collaborate with platform architects and software teams to design resilient cloud architectures, automate deployments, and minimize downtime while driving security best practices and robust monitoring. This role emphasizes hands-on engineering and operational excellence.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud SRE: Reliability, CI/CD & Automation
Senior Cloud SRE: Reliability, CI/CD & Automation

Electrolux Home Products, Inc. • Charlotte (NC)

Hybrid
USD 130,000 - 160,000
Flexible work hours
Hybrid work environment (80/20)
Health insurance
+8
Senior SRE Leader: AI-Powered Reliability for Cloud
Senior SRE Leader: AI-Powered Reliability for Cloud

Optum • Minnetonka (MN)

On-site
USD 120,000 - 160,000
Comprehensive benefits package
Incentive and recognition programs
401k contribution
Network SRE: Reliability & Automation for Cloud Infra
Network SRE: Reliability & Automation for Cloud Infra

Nebius • United States

Remote
USD 140,000 - 210,000
Remote Cloud SRE: Build Resilient Health-Tech Infra
Remote Cloud SRE: Build Resilient Health-Tech Infra

UnitedHealth Group • Eden Prairie (MN)

Remote
Confidential
Remote work flexibility
Comprehensive benefits package
Equity stock purchase
+1
Cloud SRE - Database & Distributed Systems
Cloud SRE - Database & Distributed Systems

Ll Oefentherapie • Seattle (WA)

On-site
USD 100,000 - 130,000
Senior Cloud Reliability Engineer & Security Architect
Senior Cloud Reliability Engineer & Security Architect

SAP Belgium NV/SA • Reston (VA)

Hybrid
USD 176,000 - 374,000
Senior Site Reliability Engineer - AI-Driven Cloud
Senior Site Reliability Engineer - AI-Driven Cloud

SAP • Reston (VA)

On-site
USD 176,000 - 374,000
SRE Manager: Reliability Leader for Scalable Cloud
SRE Manager: Reliability Leader for Scalable Cloud

Litera • Denver (CO)

Hybrid
USD 120,000 - 160,000
Cloud SRE — Build Resilient, Scalable Infra & Automations
Cloud SRE — Build Resilient, Scalable Infra & Automations

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000
Senior Cloud SRE & Platform Reliability
Senior Cloud SRE & Platform Reliability

Carrier • Town of Florida (NY)

On-site
USD 96,000 - 192,000
Health Care Benefits
Retirement Benefits
Paid time off
+1