ECS Site Reliability Engineer-Bellevue

Alibaba Cloud

Bellevue (NE)

On-site

USD 133,000 - 220,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
401(k) plan
Wellbeing benefits
Paid holidays
Paid vacation days
Paid sick time

Job summary

Alibaba Cloud is seeking an Elastic Compute Service (ECS) SRE to join our global operations team in Bellevue, NE. You will drive core ECS reliability, participate in performance tuning, and collaborate with engineers to optimize virtualization, containers, and cloud-native components.

You will monitor service health, analyze failures, and implement automation to improve stability for customers worldwide. A strong background in Linux/Windows internals and cloud infrastructure is essential.

Qualifications

  • Bachelor's degree or higher in Computer Science, IT, or related field.
  • At least 3 years of experience in system operations or SRE for cloud services (e.g., ECS, Kubernetes).
  • Solid understanding of Linux or Windows internals; kernel subsystems; perf/eBPF/ftrace for tuning.
  • Strong OS-level troubleshooting for CPU scheduling, memory, I/O, and network stack.
  • Familiarity with cloud resource provisioning and delivering systems; international support preferred.
  • In-depth understanding of ECS product architecture and operations.

Responsibilities

  • Drive core operations of Alibaba Cloud ECS, ensuring service stability for global users.
  • Explore virtualization, containerization, and cloud-native tech to drive innovation.
  • Collaborate with top engineers to solve complex technical challenges and grow within the team.

Skills

SRE experience
Linux/Windows internals
Performance tuning tools
OS-level troubleshooting
Cloud provisioning

Education

Bachelor's degree

Tools

Perf
eBPF
ftrace
Kubernetes

Job description

The Alibaba Cloud ECS SRE (Site Reliability Engineering) team is a critical force in ensuring system stability and reliability. The SRE team focuses on guaranteeing the high availability, high performance, and robust stability of ECS products through technical expertise and innovation.

The Alibaba Cloud ECS SRE team is not only a core technical safeguard but also a driver of technological innovation and continuous optimization. By leveraging deep expertise and close collaboration, we ensure the resilience and reliability of ECS products, safeguarding global customers' businesses. Additionally, we are committed to advancing cloud computing technologies through knowledge sharing and industry collaboration.

Joining the Alibaba Cloud ECS SRE team offers the opportunity to drive the development and optimization of world-leading cloud computing technologies, while growing alongside a passionate and creative team.

As an SRE for Elastic Compute, you will have the opportunity to:

  • Drive the core operations of Alibaba Cloud's Elastic Compute product line, ensuring service stability for global users.
  • Explore cutting-edge technologies in virtualization, containerization, and cloud-native, and drive technological innovation.
  • Grow within an open and innovative team, collaborating with top engineers to solve complex technical challenges.

If you are passionate about technology, strive for excellence, and wish to leverage your expertise in the cloud computing domain, we welcome you to join us!

Elastic Compute Service (ECS) is a core product of Alibaba Cloud. The Elastic Compute team is dedicated to building world-leading cloud computing infrastructure. As a key component of Alibaba Cloud's self-developed Apsara operating system, ECS provides full-stack computing resources covering virtual machine instances, container services , and heterogeneous computing clusters.

Through technological innovation and product optimization, the Alibaba Cloud Elastic Compute team continuously drives advancements in cloud computing technologies, delivering high-quality computing services to users worldwide. Our goal is not only to support enterprises in achieving elastic scalability but also to deeply empower infrastructure innovation in the new era. Our mission is to build an intelligent foundation of "Computing as a Service," enabling developers to focus on their core business and concentrate on innovation breakthroughs, free from the complexity of engineering implementations from chips to clusters.

Job Requirements
  • Bachelor's degree or higher in Computer Science, Information Technology, or a related field.
  • At least 3 years of experience in system operations or SRE, with familiarity in cloud computing services and core products (e.g., ECS, Kubernetes, Heterogeneous Computing).
  • Solid understanding of Linux or Windows operating system internals, including kernel subsystems (scheduling, memory management, I/O stack), and proficiency in system-level performance tuning using tools such as perf, eBPF, and ftrace.
  • Strong troubleshooting skills at the OS level, with the ability to diagnose and resolve complex issues related to CPU scheduling, memory allocation, disk I/O, and network stack.
  • Familiarity with the design and optimization of cloud resource provisioning and delivery systems; experience in supporting international customers is preferred.
  • In-depth understanding of the overall architecture and operational mechanisms of the Elastic Compute product line, with the ability to quickly identify and resolve complex issues.
Preferred Qualification
  • Cloud-related certifications (e.g., ACP, ACE, or other major cloud vendor certifications).
  • Participation in the architectural design or performance optimization projects of large cloud platforms.
  • Outstanding contributions in system stability assurance, automation tool development, or cloud-native domains are highly valued.

The pay range for this position at commencement of employment is expected to be between $133,200/year and $219,600/year. However, base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience.

If hired, employee will be in an “at-will position” and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.

  • Alibaba U.S. based full time regular employees have access to medical, dental, and vision insurance, a 401(k) plan and basic life insurance, and wellbeing benefits like FSA, subject to the terms and conditions of the applicable plans then in effect.
  • U.S. based employees are also eligible to receive up to 12 paid holidays, accrue up to 15 paid vacation days for this position, and receive up to 72 hours paid sick time (front-loaded) per calendar year.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineering (SRE) Specialist -Bellevue
Site Reliability Engineering (SRE) Specialist -Bellevue

Alibaba Cloud • Seattle (WA)

On-site
USD 133,200 - 219,600
Medical, dental, and vision insurance
401(k) plan
Paid holidays and vacation days
+1
Alibaba-Site Reliability Engineer-Bellevue
Alibaba-Site Reliability Engineer-Bellevue

BBG Ventures, LLC • Bellevue (WA)

On-site
USD 133,000 - 220,000
Medical insurance
Dental insurance
Vision insurance
+5
Site Reliability Engineer-Bellevue
Site Reliability Engineer-Bellevue

Alibaba Cloud • Bellevue (WA)

On-site
USD 145,000 - 238,000
Medical insurance
Dental insurance
Vision insurance
+3
Cloud ECS SRE: Reliability & Performance Engineer
Cloud ECS SRE: Reliability & Performance Engineer

Alibaba Cloud • Bellevue (NE)

On-site
USD 133,000 - 220,000
Medical insurance
Dental insurance
Vision insurance
+5
Staff SRE-Sunnyvale
Staff SRE-Sunnyvale

Alibaba Cloud • Sunnyvale (CA)

On-site
USD 145,000 - 238,000
Software Engineer (KV Storage)-Bellevue
Software Engineer (KV Storage)-Bellevue

Alibaba Cloud • Bellevue (NE)

On-site
USD 142,000 - 234,000
Medical, dental, and vision insurance
401(k) plan
Paid holidays
+1
Alibaba-Site Reliability Engineer-Bellevue
Alibaba-Site Reliability Engineer-Bellevue

Alibaba Group • Bellevue (WA)

On-site
USD 133,000 - 220,000
Medical Insurance
Dental Insurance
Vision Insurance
+6
Site Reliability Engineer - Product & Data Security-Sunnyvale
Site Reliability Engineer - Product & Data Security-Sunnyvale

Alibaba Cloud • Sunnyvale (CA)

On-site
USD 104,000 - 171,000
Software Engineer (Cloud Storage Services)
Software Engineer (Cloud Storage Services)

Alibaba Cloud • Seattle (WA)

On-site
USD 156,000 - 257,000
Medical insurance
Dental insurance
Vision insurance
+5
Computing Products SRE Engineer
Computing Products SRE Engineer

Lightspeed Studios • Palo Alto (CA)

On-site
USD 106,000 - 199,000
Sign-on bonus possible
Relocation package
Restricted stock units (RSUs)
+2