Cloud Infrastructure – Site Reliability Engineer (SRE)-Sunnyvale

Alibaba Cloud

Sunnyvale (CA)

On-site

USD 104,000 - 171,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Alibaba Cloud is seeking a Reliability Engineer for its Cloud Product Operations & Reliability team in Sunnyvale. The role focuses on stability, performance tuning, and high-availability design for messaging middleware and cloud services.

You will manage Kubernetes-based deployments, automations, and disaster recovery workflows. The ideal candidate has 2+ years in distributed systems reliability, proficiency in Python/Go/Java, and hands-on experience with Kubernetes, Terraform, and Helm.

Qualifications

  • Over 2 years of experience in distributed systems reliability engineering.
  • Familiar with high-availability architecture design.
  • Proficient in Python, Go, or Java.

Responsibilities

  • Oversee stability maintenance, performance tuning, and HA design for cloud middleware.
  • Manage containerized middleware lifecycles on Kubernetes clusters (deployments, auto-scaling, upgrades).
  • Lead incident response, RCA efforts, and log tracing for production issues.
  • Develop diagnostic tools in Python/Go to resolve bottlenecks and compatibility challenges.
  • Build automation tools (Python/Go/Shell) and IaC workflows to standardize deployment and DR.

Skills

Kubernetes
Python
Go
Java
Distributed systems
Terraform

Tools

Helm
Operator

Job description

Alibaba Cloud Native Message Middleware Team is responsible for message products, including RocketMQ and other messaging products. We are committed to creating a more stable, user-friendly, streaming, and large-scale messaging platform for the future.

Cloud Product Operations & Reliability

Oversee stability maintenance, performance tuning, and high-availability architecture design for cloud middleware, including messaging middleware (Kafka/RocketMQ).

Manage the containerized middleware lifecycle on Kubernetes clusters: implement deployments, auto-scaling, version upgrades, and resource optimization in K8s environments.

Incident Response & Root Cause Analysis

Lead the troubleshooting of middleware-related incidents (e.g., message backlog, service registration failures) through log analysis, distributed tracing, and monitoring systems.

Develop diagnostic tools using Java/Go to resolve production issues, performance bottlenecks, and compatibility challenges.

Automation & Operational Excellence

Build Python/Go/Shell automation tools to standardize middleware deployment, monitoring, and disaster recovery workflows.

Implement chaos engineering experiments, capacity planning strategies, and failover mechanisms to enhance system resilience.

Strong scripting skills in Shell/Python and experience with Infrastructure as Code (IaC) tools (Terraform preferred).

Job Requirements

Experience: Over 2 years of experience in distributed systems reliability engineering, familiar with high-availability architecture design, and proficient in at least one of Python, Go, or Java.

Messaging: Cluster management, message reliability assurance, and performance optimization for Kafka/RocketMQ.

Hands-on Experience Deploying Middleware On Kubernetes (Helm/Operator Preferred).

Automation: Ability to convert operations experience into automated solutions and familiarity with various message middleware, e.g., Kafka and RocketMQ.

Preferred Qualification

SRE Practices: Familiar with core SRE practices (incident review, error budgeting, chaos engineering) and experienced in building automated risk control systems.

The pay range for this position at commencement of employment is expected to be between $104,400 and $171,000/year. However, base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience.

If hired, employee will be in an “at-will position” and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Messaging SRE: Kubernetes Ops & Resilience
Cloud Messaging SRE: Kubernetes Ops & Resilience

Alibaba Cloud • Sunnyvale (CA)

On-site
USD 104,000 - 171,000
Site Reliability Engineering (SRE) Specialist -Bellevue
Site Reliability Engineering (SRE) Specialist -Bellevue

Alibaba Cloud • Seattle (WA)

On-site
USD 133,200 - 219,600
Medical, dental, and vision insurance
401(k) plan
Paid holidays and vacation days
+1
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New York (NY)

Hybrid
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+1
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New Jersey

On-site
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Cloud SRE Engineer - Mandarin Bilingual
Cloud SRE Engineer - Mandarin Bilingual

Ipro Networks Pte. Ltd. • Palo Alto (CA)

On-site
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Principal Site Reliability Engineer
Principal Site Reliability Engineer

ViziRecruiter,LLC. • Quincy (MA)

Hybrid
USD 146,000 - 221,000
Site Reliability Engineer
Site Reliability Engineer

Stelvio Inc. • Town of Texas (WI)

On-site
USD 125,000 - 145,000
Site Reliability Engineer
Site Reliability Engineer

Knack Solutions • Richmond (VA)

On-site
USD 100,000 - 130,000
Site Reliability Engineer (Senior or Staff), Deployments
Site Reliability Engineer (Senior or Staff), Deployments

MongoDB • New Jersey

On-site
USD 127,000 - 249,000
Employee stock purchase program
Flexible paid time off
20 weeks fully-paid parental leave
+5