Cloud ECS SRE: Reliability & Performance Engineer

Alibaba Cloud

Bellevue (NE)

On-site

USD 133,000 - 220,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
401(k) plan
Wellbeing benefits
Paid holidays
Paid vacation days
Paid sick time

Job summary

Alibaba Cloud is seeking an Elastic Compute Service (ECS) SRE to join our global operations team in Bellevue, NE. You will drive core ECS reliability, participate in performance tuning, and collaborate with engineers to optimize virtualization, containers, and cloud-native components.

You will monitor service health, analyze failures, and implement automation to improve stability for customers worldwide. A strong background in Linux/Windows internals and cloud infrastructure is essential.

Qualifications

  • Bachelor's degree or higher in Computer Science, IT, or related field.
  • At least 3 years of experience in system operations or SRE for cloud services (e.g., ECS, Kubernetes).
  • Solid understanding of Linux or Windows internals; kernel subsystems; perf/eBPF/ftrace for tuning.
  • Strong OS-level troubleshooting for CPU scheduling, memory, I/O, and network stack.
  • Familiarity with cloud resource provisioning and delivering systems; international support preferred.
  • In-depth understanding of ECS product architecture and operations.

Responsibilities

  • Drive core operations of Alibaba Cloud ECS, ensuring service stability for global users.
  • Explore virtualization, containerization, and cloud-native tech to drive innovation.
  • Collaborate with top engineers to solve complex technical challenges and grow within the team.

Skills

SRE experience
Linux/Windows internals
Performance tuning tools
OS-level troubleshooting
Cloud provisioning

Education

Bachelor's degree

Tools

Perf
eBPF
ftrace
Kubernetes

Job description

Alibaba Cloud is seeking an Elastic Compute Service (ECS) SRE to join our global operations team in Bellevue, NE. You will drive core ECS reliability, participate in performance tuning, and collaborate with engineers to optimize virtualization, containers, and cloud-native components.

You will monitor service health, analyze failures, and implement automation to improve stability for customers worldwide. A strong background in Linux/Windows internals and cloud infrastructure is essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cloud SRE Specialist - Reliability & Automation Engineer
Cloud SRE Specialist - Reliability & Automation Engineer

Alibaba Cloud • Seattle (WA)

On-site
USD 133,000 - 220,000
Medical, dental, and vision insurance
401(k) plan
Paid holidays and vacation days
+1
ECS Site Reliability Engineer-Bellevue
ECS Site Reliability Engineer-Bellevue

Alibaba Cloud • Bellevue (NE)

On-site
USD 133,000 - 220,000
Medical insurance
Dental insurance
Vision insurance
+5
Site Reliability Engineering (SRE) Specialist -Bellevue
Site Reliability Engineering (SRE) Specialist -Bellevue

Alibaba Cloud • Seattle (WA)

On-site
USD 133,200 - 219,600
Medical, dental, and vision insurance
401(k) plan
Paid holidays and vacation days
+1
Cloud Computing SRE Engineer – North America
Cloud Computing SRE Engineer – North America

Lightspeed Studios • Palo Alto (CA)

On-site
USD 106,000 - 199,000
Sign-on bonus possible
Relocation package
Restricted stock units (RSUs)
+2
Cloud Messaging SRE: Kubernetes Ops & Resilience
Cloud Messaging SRE: Kubernetes Ops & Resilience

Alibaba Cloud • Sunnyvale (CA)

On-site
USD 104,000 - 171,000
Platform Reliability Engineer — Kubernetes & Infra Ops
Platform Reliability Engineer — Kubernetes & Infra Ops

Alibaba Cloud • Sunnyvale (CA)

On-site
USD 145,000 - 238,000
SRE for AI Inference Platform — Reliability & Automation
SRE for AI Inference Platform — Reliability & Automation

Alibaba Group • Bellevue (WA)

On-site
USD 133,000 - 220,000
Medical Insurance
Dental Insurance
Vision Insurance
+6
SRE, Cloud & Kubernetes Platform Engineer
SRE, Cloud & Kubernetes Platform Engineer

Algolia • United States

Hybrid
USD 65,000 - 90,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000
Cloud SRE: Reliability & Performance Engineer (Onsite)
Cloud SRE: Reliability & Performance Engineer (Onsite)

Cognizant • Plano (TX)

On-site
USD 74,000 - 105,000
Medical/Dental/Vision/Life Insurance
Paid holidays plus Paid Time Off
401(k) plan
+3