Site Reliability Engineer

Alibaba Cloud

Sunnyvale (CA)

On-site

USD 104,000 - 171,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Alibaba Cloud’s Cloud Intelligence Group SRE team ensures stability of production environments, data reliability, and service continuity for cloud customers in Sunnyvale, CA. The role emphasizes incident management and proactive risk mitigation to achieve high availability and strong on-call performance.

The position requires 3+ years in SRE, expertise in Linux/cloud, and hands-on Kubernetes/monitoring experience, with programming in Golang, Python, or Java.

Qualifications

  • A degree in Computer Science or related field.
  • 3+ years of experience as an SRE or above.
  • Proficiency with Linux environments or cloud infrastructure.
  • Strong diagnostic and problem-solving skills.
  • Experience with Kubernetes or monitoring systems.
  • Experience designing distributed systems.

Responsibilities

  • Daily operations and maintenance of applications, databases, and middleware, plus troubleshooting and customer inquiries.
  • Collaborate with R&D to develop critical support plans for peak periods and standby reviews.
  • Respond to incidents and contribute to post-incident reviews.

Skills

Linux environments
Cloud infrastructure
Kubernetes
Monitoring systems
Golang
Python
Java
Distributed systems

Education

Bachelor's degree in Computer Science or related field

Job description

The mission of the Cloud Intelligence Group SRE (Site Reliability Engineering) Team is to ensure the stability of production environments, enterprise-grade cloud data reliability, and service continuity for the Cloud Intelligence Group. Our greatest challenge lies in guaranteeing uninterrupted business operations for cloud-based customers and achieving availability that exceeds 99.99%.

Objectives of the Cloud Intelligence Group SRE Team

Our goal is to establish a systematic stability assurance framework that integrates technology and management, including but not limited to:

  • 1. Developing stability standards and metrics
    • Covering robust architecture, R&D quality, release management, production environment operations, and more.
    • Embedding stability into Alibaba Cloud's technical R&D system.
  • 2. Driving major stability governance campaigns
    • Initiatives such as full-stack disaster recovery, phased change rollout, the 1-5-10 emergency response mechanism (1-minute alerting, 5-minute triage, 10-minute recovery), and financial-loss prevention.
    • Rapidly and continuously mitigating stability risks.
  • 3. Building a stability-focused technical platform
    • Platform capabilities for unattended change management, red/blue team drills, emergency collaboration, risk and vulnerability inspection, and monitoring/alerting.
    • Simplifying stability engineering through automation and tooling.
  • 4. Executing production incident management
    • Emergency response, cross-team coordination, root cause analysis, rapid recovery, and post-incident reviews to drive systemic improvements.
  • 5. Ensuring stability for large-scale customer events
    • Technical and operational support for critical activities such as Olympics and customer business peak periods.
  • 6. On-call responsibilities
    • Responding to customer issues within Service Level Agreement (SLA) timeframes, resolving problems proactively, and enhancing customer experience.
Responsibilities

The objective of the Cloud Intelligence Group's SRE team is to establish a systematic stability assurance framework that integrates technology and management, including but not limited to:

  • 1. Daily operations and maintenance of applications, databases, and middleware, as well as troubleshooting and answering customer inquiries;
  • 2. Collaborating with R&D to develop critical support plans based on customer business requirements during peak periods, including preparation during the standby period, on-duty support during critical periods, and post-standby review;
  • A degree in Computer Science or related field
  • 3 years of experience as a Site Reliability Engineer (SRE) or above.
  • Proficiency with Linux environments or cloud infrastructure
  • Exceptional system diagnostic and problem-solving skills
  • Strong teamwork spirit and ability to work well under pressure
  • In-depth understanding of Kubernetes or monitoring systems
  • Expertise in programming languages such as Golang, Python, or Java
  • Experience in designing and implementing distributed systems

The pay range for this position at commencement of employment is expected to be between $104,400 and $171,000/year. However, base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience.

If hired, employee will be in an "at-will position" and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff SRE
Staff SRE

Alibaba Cloud • Sunnyvale (CA)

On-site
USD 145,000 - 238,000
Site Reliability Engineering (SRE) Specialist -Bellevue
Site Reliability Engineering (SRE) Specialist -Bellevue

Alibaba Cloud • Seattle (WA)

On-site
USD 133,200 - 219,600
Medical, dental, and vision insurance
401(k) plan
Paid holidays and vacation days
+1
Alibaba-Site Reliability Engineer-Bellevue
Alibaba-Site Reliability Engineer-Bellevue

BBG Ventures, LLC • Bellevue (WA)

On-site
USD 133,000 - 220,000
Medical insurance
Dental insurance
Vision insurance
+5
Staff SRE-Sunnyvale
Staff SRE-Sunnyvale

Alibaba Cloud • Sunnyvale (CA)

On-site
USD 145,000 - 238,000
ECS Site Reliability Engineer-Bellevue
ECS Site Reliability Engineer-Bellevue

Alibaba Cloud • Bellevue (NE)

On-site
USD 133,000 - 220,000
Medical insurance
Dental insurance
Vision insurance
+5
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

LexisNexis Risk Solutions • San Jose (CA), Northern (KY)

Hybrid
USD 105,000 - 175,000
401(k) with match
Wellbeing programs
Life Insurance
+1
Site Reliability Engineering (SRE) Consultant
Site Reliability Engineering (SRE) Consultant

TekWissen ® • Charlotte (NC)

On-site
USD 102,286 - 145,681
Site Reliability Engineer
Site Reliability Engineer

Harvey Nash • United States

Remote
USD 120,000 - 150,000
Site Reliability Engineer - Product & Data Security-Sunnyvale
Site Reliability Engineer - Product & Data Security-Sunnyvale

Alibaba Cloud • Sunnyvale (CA)

On-site
USD 104,000 - 171,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Good co India • United States

Remote
USD 120,000 - 160,000