Senior Client-Facing SRE / Cloud Engineer (AWS, Kubernetes)

Salve.Lab

Chicago (IL)

On-site

USD 140,000 - 190,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Fully remote
B2B cooperation
Enterprise collaboration

Job summary

Salve.Lab is seeking a Senior Cloud Reliability Engineer for a fully remote role spanning EU and US markets. You will own production AWS and Kubernetes environments, manage on-call incidents, and drive permanent reliability improvements.

Expect strong collaboration with enterprise customers and cross-functional teams. The role emphasizes IaC with Terraform, GitOps with Argo CD, and advanced observability using Prometheus, Grafana, and ELK.

Qualifications

  • Senior hands-on AWS expertise and production Kubernetes experience.
  • Experience in on-call ownership and incident response.
  • Strong observability using Prometheus, Grafana, ELK.
  • CI/CD and GitOps expertise across environments.
  • Proficient with Python, Bash or Go for automation.
  • Excellent English communication with customers and teams.

Responsibilities

  • Own reliability and health of production AWS and Kubernetes environments.
  • Lead on-call incidents from detection to resolution.
  • Conduct root cause analysis and implement permanent fixes.
  • Identify and remediate architectural and operational weaknesses.
  • Improve deployments and GitOps with Argo CD and IaC.
  • Build monitoring, logging, tracing, dashboards, and alerts.
  • Scale workloads with Kubernetes autoscaling tools like Karpenter/KEDA.
  • Collaborate with enterprise customers on troubleshooting and architecture discussions.
  • Develop automation tooling in Python/Go/Bash and improve security posture.

Skills

SRE
Cloud Reliability
Kubernetes
Terraform
GitOps
Argo CD
Prometheus
Grafana
ELK
Python
Go
Bash
Ansible
Kafka
Linux Networking
Incident Response
English Communication
Argo CD Deployment

Education

AWS Certification

Tools

Argo CD
Kubernetes
Terraform
Go
Python
Bash
Ansible
Kafka
Prometheus
Grafana
ELK
AWS CLI

Job description

B2B Contract | EU or US - fully remote

Role Overview

We are looking for a Senior Cloud Reliability Engineer to take hands‑on ownership of highly available, cloud-native production environments.

This is not a traditional DevOps role focused primarily on building CI/CD pipelines or migrating infrastructure. We are looking for an engineer who has operated critical production systems, owned incidents while on call, and can identify weaknesses in an existing cloud environment and drive meaningful improvements.

The role combines deep AWS and Kubernetes engineering with SRE practices, infrastructure automation, observability, incident management, and direct technical interaction with enterprise customers.

You will have significant autonomy to challenge existing approaches, propose better solutions, and improve the reliability, scalability, security, and operational maturity of the platform.

Key Responsibilities
  • Own the reliability and operational health of production AWS and Kubernetes environments.
  • Participate in on-call rotations and take ownership of high-severity production incidents from detection through mitigation and resolution.
  • Lead root cause analysis and post-incident reviews, implementing permanent corrective actions rather than temporary fixes.
  • Identify architectural, reliability, security, performance, and operational weaknesses within existing cloud environments.
  • Propose and implement improvements based on AWS and Kubernetes best practices.
  • Design, maintain, and continuously improve AWS infrastructure and production Kubernetes/EKS environments.
  • Automate infrastructure provisioning and operational workflows using Terraform and configuration-management tools.
  • Improve deployment and GitOps processes using tools such as Argo CD.
  • Build and improve monitoring, logging, tracing, dashboards, and actionable alerting using Prometheus, Grafana, ELK and related observability technologies.
  • Improve scalability and workload management using Kubernetes autoscaling technologies such as Karpenter or KEDA.
  • Support distributed and event-driven environments, including technologies such as Kafka.
  • Develop automation and operational tooling using Python, Bash, Go, or similar languages.
  • Strengthen cloud security, resilience, disaster recovery, and production-readiness practices.
  • Work directly with enterprise customers when required, including technical troubleshooting, escalations, incident discussions, and explaining infrastructure or reliability issues.
  • Collaborate with engineering teams while bringing independent ideas and challenging existing technical approaches where improvements can be made.
Requirements
  • Strong professional experience in Site Reliability Engineering, Cloud Reliability, Platform Engineering, or a comparable production-focused role.
  • Senior-level hands-on AWS expertise, with the ability to understand, design, troubleshoot, and improve existing AWS architectures.
  • Senior-level Kubernetes experience, ideally operating Amazon EKS in production.
  • Strong hands-on experience with Terraform and Infrastructure as Code.
  • Proven experience participating in an on-call rotation and personally owning production incidents.
  • Demonstrable experience with high-severity incident response, root cause analysis, postmortems, MTTR reduction, and permanent remediation.
  • Proven experience communicating directly with external or enterprise customers in a technical capacity, particularly during troubleshooting, escalations, architecture discussions, or production incidents.
  • Strong observability experience with Prometheus, Grafana, ELK, or equivalent production observability stacks.
  • Strong Linux and cloud networking fundamentals.
  • Experience automating operational processes using Python, Bash, Go, or similar scripting/programming languages.
  • Experience with CI/CD and GitOps environments.
  • Ability to independently identify infrastructure weaknesses and translate them into practical technical improvements.
  • Strong communication skills and professional-level English.
  • Comfortable explaining complex technical issues to both engineering teams and customers.
  • Hands-on production experience with Argo CD and GitOps-based deployment workflows.
  • Hands-on experience with Kafka in distributed production environments.
  • Production experience with Kubernetes autoscaling using Karpenter and/or KEDA.
  • Strong hands-on experience with Ansible for infrastructure/configuration automation.
  • Practical experience with AWS security tooling and cloud security best practices.
  • A relevant AWS certification.
What's on Offer
  • Fully remote position.
  • Full-time B2B cooperation.
  • Opportunity to work on complex, production-critical cloud environments.
  • High level of technical ownership and autonomy.
  • Real influence over cloud architecture, reliability practices, automation, and platform improvements.
  • International engineering environment.
  • Direct collaboration with experienced technical teams and enterprise customers.
  • Long-term cooperation and opportunities to introduce new technologies and engineering practices.
Diversity and Inclusion Commitment

We are dedicated to creating and sustaining an inclusive, respectful workplace for all -regardless of gender, ethnicity, or background. We actively encourage applicants from all identities and experience levels to apply and bring your authentic self to our fast-paced, supportive team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New Jersey

On-site
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New York (NY)

Hybrid
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+1
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior/Staff Cloud Reliability Engineer
Senior/Staff Cloud Reliability Engineer

Cerebras • Mountain View (CA)

On-site
USD 190,000 - 240,000
Senior/Staff Cloud Reliability Engineer
Senior/Staff Cloud Reliability Engineer

ThoughtSpot • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior Site Reliability Engineer / DevOps Engineer
Senior Site Reliability Engineer / DevOps Engineer

Prophet Town • Mountain View (CA)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000
SRE/Devops Engineer
SRE/Devops Engineer

INSPYR Solutions • Sunnyvale (CA)

Hybrid
USD 120,000 - 180,000
Work-life balance
No on-call requirements
Standard business hours
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

Remote
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3