Site Reliability Engineering (SRE) Specialist -Bellevue

Alibaba Cloud

Seattle (WA)

On-site

USD 133,200 - 219,600

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
401(k) plan
Paid holidays and vacation days
Paid sick time

Job summary

Alibaba Cloud is seeking a Site Reliability Engineer (SRE) to join the ECS team in Seattle, Washington. This role involves ensuring the performance and reliability of cloud cloud computing infrastructure by optimizing maintenance operations and enhancing automation processes. The SRE will be responsible for customer support, issue resolution, and contributing to technology advancements.

The ideal candidate will have over 5 years of experience in IT with strong expertise in Linux, container technologies, and operational automation, while working on cross-functional projects.

Qualifications

  • 5+ years of operation and maintenance (O&M) experience in IT, internet, or cloud computing industries.
  • Proficient in Linux operating systems and troubleshooting OS and network issues.
  • Familiar with containerization and orchestration technologies such as Kubernetes.

Responsibilities

  • Responsible for the delivery and operation/maintenance of various clusters.
  • Establish and optimize operation/maintenance service systems for product stability.
  • Develop delivery standards and enhance daily work efficiency through tool platforms.

Skills

Linux operating systems
Troubleshooting OS and network issues
Containerization and orchestration (Kubernetes, Slurm, LSF)
Automation and platform-based solutions
Excellent communication skills

Job description

Elastic Compute Service (ECS) is a core product of Alibaba Cloud. The Elastic Compute team is dedicated to building world‑leading cloud computing infrastructure. As a key component of Alibaba Cloud's self‑developed Apsara operating system, Elastic Compute Service (ECS) provides full‑stack computing resources covering virtual machine instances, container services and heterogeneous computing clusters.

Through technological innovation and product optimization, the Alibaba Cloud Elastic Compute team continuously drives advancements in cloud computing technologies, delivering high‑quality computing services to users worldwide.

  • Our goal is not only to support enterprises in achieving elastic scalability but also to deeply empower infrastructure innovation in the new era. Our mission is to build an intelligent foundation of "Computing as a Service," enabling developers to focus on businesses, to concentrate on breakthroughs, without worrying about the complex engineering implementations from chips to clusters.
SRE Team

The Alibaba Cloud Elastic Compute Service (ECS) SRE (Site Reliability Engineering) team is a critical force in ensuring system stability and reliability. The SRE team focuses on guaranteeing the high availability, high performance, and robust stability of ECS products through technical expertise and innovation.

The Alibaba Cloud ECS SRE team is not only a core technical safeguard but also a driver of technological innovation and continuous optimization. By leveraging technical capabilities and collaborative teamwork, we ensure the stability and reliability of ECS products, safeguarding global customers' businesses. Additionally, we are committed to advancing cloud computing technologies through knowledge sharing and industry collaboration.

Joining the Alibaba Cloud ECS SRE team offers the opportunity to engage in the development and optimization of world‑leading cloud computing technologies, while growing alongside a passionate and creative team.

  • Responsible for the delivery and operation/maintenance of various clusters, and participate in the architecture design and construction of the infrastructure operation platform.
  • Establish and optimize operation/maintenance service systems to achieve product stability and SLA goals.
  • Develop delivery standards, document maintenance specifications, and enhance daily work efficiency through tool platforms.
  • This position involves on‑call responsibilities, requiring timely customer response within Service Level Agreement (SLA) timeframes, driving issue resolution and improving customer experience.
Job Requirements

1.5+ years of operation and maintenance (O&M) experience in IT, internet, or cloud computing industries.

  • Proficient in Linux operating systems and mainstream protocols (e.g., TCP/IP), with solid hands‑on experience in troubleshooting OS and network issues.
  • Familiar with containerization and orchestration technologies such as Kubernetes, Slurm, and LSF.
  • Ability to analyze and document technical issues systematically, develop tools/systems to optimize workflows, and improve operational efficiency through automation and platform‑based solutions.
  • Strong self‑driven learning capabilities, excellent communication skills, and experience leading cross‑team projects. Results‑driven and action‑oriented, with a commitment to excellence.

The pay range for this position at commencement of employment is expected to be between $133,200/year and $219,600/year. However, base pay offered may vary depending on multiple individualized factors, including market location, job‑related knowledge, skills, and experience.

If hired, employee will be in an "at‑will position" and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.

Alibaba U.S. based full‑time regular employees have access to medical, dental, and vision insurance, a 401(k) plan and basic life insurance, and wellbeing benefits like FSA, subject to the terms and conditions of the applicable plans then in effect. U.S. based employees are also eligible to receive up to 12 paid holidays, accrue up to 15 paid vacation days for this position, and receive up to 72 hours paid sick time (front‑loaded) per calendar year.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ECS Site Reliability Engineer-Bellevue
ECS Site Reliability Engineer-Bellevue

Alibaba Cloud • Bellevue (NE)

On-site
USD 133,000 - 220,000
Medical insurance
Dental insurance
Vision insurance
+5
Alibaba-Site Reliability Engineer-Bellevue
Alibaba-Site Reliability Engineer-Bellevue

BBG Ventures, LLC • Bellevue (WA)

On-site
USD 133,000 - 220,000
Medical insurance
Dental insurance
Vision insurance
+5
Staff SRE
Staff SRE

Alibaba Cloud • Sunnyvale (CA)

On-site
USD 145,000 - 238,000
Site Reliability Engineer
Site Reliability Engineer

Alibaba Cloud • Sunnyvale (CA)

On-site
USD 104,000 - 171,000
Staff SRE-Sunnyvale
Staff SRE-Sunnyvale

Alibaba Cloud • Sunnyvale (CA)

On-site
USD 145,000 - 238,000
Site Reliability Engineer-Bellevue
Site Reliability Engineer-Bellevue

Alibaba Cloud • Bellevue (WA)

On-site
USD 145,000 - 238,000
Medical insurance
Dental insurance
Vision insurance
+3
Cloud ECS SRE: Reliability & Performance Engineer
Cloud ECS SRE: Reliability & Performance Engineer

Alibaba Cloud • Bellevue (NE)

On-site
USD 133,000 - 220,000
Medical insurance
Dental insurance
Vision insurance
+5
Software Engineer (KV Storage)-Bellevue
Software Engineer (KV Storage)-Bellevue

Alibaba Cloud • Bellevue (NE)

On-site
USD 142,000 - 234,000
Medical, dental, and vision insurance
401(k) plan
Paid holidays
+1
Cloud SRE Specialist - Reliability & Automation Engineer
Cloud SRE Specialist - Reliability & Automation Engineer

Alibaba Cloud • Seattle (WA)

On-site
USD 133,200 - 219,600
Medical, dental, and vision insurance
401(k) plan
Paid holidays and vacation days
+1
Site Reliability Engineer - Product & Data Security-Sunnyvale
Site Reliability Engineer - Product & Data Security-Sunnyvale

Alibaba Cloud • Sunnyvale (CA)

On-site
USD 104,000 - 171,000