Site Reliability Engineer (Contractor)

Varo

San Francisco, Salt Lake City, New York (CA, UT, NY)

Hybrid

USD 140,000 - 190,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Varo is seeking an SRE/DevOps engineer to design, build, and run large-scale, distributed systems in AWS and Kubernetes. You will scale production infrastructure, develop CI/CD pipelines, and collaborate with developers to reduce friction and improve reliability.

This role emphasizes automation, observability, and AI/ML-enabled runbooks, with on-call responsibilities and cost-optimization focus in a fast-growing fintech environment.

Qualifications

  • 3+ years in an SRE, DevOps, or Infrastructure Engineering role.
  • Strong hands-on experience with core AWS services including EKS, EC2, RDS Aurora, MSK, S3, IAM, VPC, and Direct Connect.
  • Deep production experience with Kubernetes, Helm, and GitOps tools like ArgoCD.

Responsibilities

  • Manage, upgrade, and autoscale EKS clusters across multiple environments and AWS accounts.
  • Write Terraform modules and Helm charts to support GitOps workflows using ArgoCD and GitLab CI/CD pipelines.
  • Maintain and troubleshoot Kafka (MSK) clusters, including broker health, connectors, and CDC pipelines.
  • Improve observability using Prometheus, Grafana, and ELK while proactively identifying cloud cost-optimization opportunities.
  • Automate operational tasks with Python and leverage AI/ML techniques for predictive alerting and intelligent runbooks.
  • Handle Platform Service Desk requests, including Terraform merge request reviews, access management, and deployment support.
  • Participate in the production on-call rotation, support incident response, and contribute to blameless post-mortems.

Skills

SRE / DevOps
AWS Expertise
Kubernetes
Terraform
ArgoCD
GitLab CI/CD
Python scripting

Tools

Terraform
Helm
ArgoCD
GitLab CI/CD
Kubernetes
Airflow
Databricks
EMR
Kafka/MSK

Job description

Varo is an entirely new kind of bank. All digital, mission-driven, FDIC insured and designed for the way our customers live their lives. A bank for all of us.

About the role

Varo’s SRE team is well established, designing, building, and running large-scale, distributed, fault-tolerant systems that power most of Varo's operations. We live and breathe AWS and Kubernetes, having an open source first and result oriented mindset.

We are an automation and observability focused team and we strive to automate ourselves out of manual / remedial tasks. We monitor and create dashboards to promote a data-driven approach to scale out our platform.

On a typical day, members of our team are hands-on scaling-out production infrastructure, building out CI/CD pipelines, and brainstorming with developers on how to make things better. We collectively strive to build and maintain a rapid-feedback platform that enables our engineers to accomplish their own goals instead of creating friction.

Responsibilities
  • EKS & Karpenter Management: Manage, upgrade, and autoscale EKS clusters across multiple environments (SIT, UAT, Prod) and AWS accounts.

  • Infrastructure as Code & GitOps: Write Terraform modules and Helm charts to support GitOps workflows using ArgoCD and GitLab CI/CD pipelines.

  • Kafka & Data Platform Support: Maintain and troubleshoot Kafka (MSK) clusters, including broker health, connectors, and CDC pipelines.

  • Observability & Cost Control: Improve observability using Prometheus, Thanos, Grafana, and ELK while proactively identifying cloud cost-optimization opportunities.

  • Automation & AIOps: Automate operational tasks with Python and leverage AI/ML techniques for predictive alerting and intelligent runbooks.

  • Service Desk & Support: Handle Platform Service Desk requests, including Terraform merge request reviews, access management, and deployment support.

  • Incident Response & On-Call: Participate in the production on-call rotation, support incident response, and contribute to blameless post-mortems.

Skills & Qualification
  • Experience: 3+ years of experience in an SRE, DevOps, or Infrastructure Engineering role, with the ability to work independently and manage multiple workstreams.

  • AWS Expertise: Strong hands-on experience with core AWS services, including EKS, EC2, RDS Aurora, MSK, S3, IAM, VPC, and Direct Connect.

  • Kubernetes & Deployment: Deep production experience with Kubernetes (upgrades, networking, RBAC) alongside Helm and GitOps tools like ArgoCD.

  • Infrastructure as Code: Advanced proficiency with Terraform, including writing modules and managing multi-account/multi-environment states.

  • Data Infrastructure: Experience supporting and maintaining data platforms such as Airflow, Databricks, EMR, Kafka/MSK, or CDC pipelines.

  • Networking & Automation: Solid understanding of networking (VPCs, security groups, Istio, DNS) paired with strong Python scripting skills for tooling and automation.

  • Observability & AI: Experience managing observability stacks (Prometheus, Grafana, ELK) and effectively leveraging AI/LLM tools for automation and incident analysis.

  • Nice to Have: Experience with Karpenter and KEDA, GitLab CI/CD pipeline experience, Hashicorp Vault for secrets management.

For cash compensation, we set standard ranges for all US-based roles based on function, level, and geographic location, benchmarked against similar‑stage growth companies. Final offer amounts are determined by multiple factors as well as candidate experience and expertise and may vary from the identified range.

Varo is an equal opportunity employer. Varo embraces diversity and we are committed to building teams that represent a variety of backgrounds, perspectives, and skills. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Director of Engineering
Director of Engineering

Neara • New York (NY)

On-site
USD 200,000 - 280,000
Director of Engineering
Director of Engineering

Neara • Salt Lake City (UT)

On-site
USD 200,000 - 280,000
Director of Engineering
Director of Engineering

Neara • Charlotte (NC)

On-site
USD 200,000 - 280,000
Director of Engineering
Director of Engineering

Neara • San Francisco (CA)

On-site
USD 200,000 - 280,000
Security Engineer
Security Engineer

Varo • United States

Remote
USD 120,000 - 180,000
Site Reliability Engineer: Cloud Automation & Observability
Site Reliability Engineer: Cloud Automation & Observability

Varo • San Francisco (CA), Salt Lake City (UT), New York (NY)

Hybrid
USD 140,000 - 190,000
FP&A Analyst
FP&A Analyst

Varo • Charlotte (NC), New York (NY), San Francisco (CA), Salt Lake City (UT)

On-site
USD 90,000 - 130,000
Sr. Staff Data Scientist, Machine Learning
Sr. Staff Data Scientist, Machine Learning

Varo Bank • Atlanta (GA)

On-site
USD 210,000 - 280,000
Bonus opportunities
Equity options
Competitive benefits
FP&A Analyst
FP&A Analyst

Varo Bank • Salt Lake City (UT)

On-site
USD 90,000 - 130,000
FP&A Analyst
FP&A Analyst

Varo Bank • Charlotte (NC)

On-site
USD 65,000 - 95,000