Site Reliability Engineer - Big Data Cloud Operations

Alibaba Cloud

Bellevue (WA)

On-site

USD 145,000 - 238,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
401(k) plan
Paid holidays
Paid vacation

Job summary

Alibaba Cloud is seeking a senior operations professional to ensure the stability and performance of its Big Data and PAI products in the US region. You will manage service delivery, deployments, monitoring, and incident response while coordinating cost budgets and resource planning.

The role requires extensive experience with large-scale distributed systems, Kubernetes, and scripting in Python and Shell for automation and troubleshooting.

Qualifications

  • Bachelor's degree in Computer Science or related field with solid fundamentals.
  • Expert-level Linux system administration skills.
  • Experience with Alibaba Cloud proprietary Big Data and PAI products is preferred.
  • 5+ years of experience in development or operations of large-scale distributed systems.
  • Cloud-native competency with hands-on Kubernetes experience, including architecture understanding and change releases.
  • Strong scripting skills in Python and Shell for automated troubleshooting and monitoring.

Responsibilities

  • Ensure stability of Alibaba Cloud Big Data and PAI products in the US region.
  • Handle service delivery and deployment, monitoring configuration, and emergency incident response.
  • Manage changes, releases, and troubleshooting of complex customer issues.
  • Oversee cloud platform costs, budgeting, forecasting, and server procurement coordination.
  • Scale services, including expansion and reduction, deployment and decommissioning.
  • Design and develop intelligent systems agents for cloud operations and automation.

Skills

Linux administration
Kubernetes
Python scripting
Shell scripting
Distributed systems
Cloud operations

Education

Bachelor's degree in Computer Science

Tools

Open-source big data platforms

Job description

Alibaba Cloud is seeking a senior operations professional to ensure the stability and performance of its Big Data and PAI products in the US region. You will manage service delivery, deployments, monitoring, and incident response while coordinating cost budgets and resource planning.

The role requires extensive experience with large-scale distributed systems, Kubernetes, and scripting in Python and Shell for automation and troubleshooting.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer-Bellevue
Site Reliability Engineer-Bellevue

Alibaba Cloud • Bellevue (WA)

On-site
USD 145,000 - 238,000
Medical insurance
Dental insurance
Vision insurance
+3
Cloud SRE Specialist - Reliability & Automation Engineer
Cloud SRE Specialist - Reliability & Automation Engineer

Alibaba Cloud • Seattle (WA)

On-site
USD 133,000 - 220,000
Medical, dental, and vision insurance
401(k) plan
Paid holidays and vacation days
+1
SRE for AI Inference Platform — Reliability & Automation
SRE for AI Inference Platform — Reliability & Automation

Alibaba Group • Bellevue (WA)

On-site
USD 133,000 - 220,000
Medical Insurance
Dental Insurance
Vision Insurance
+6
Infra Controls & Automation Architect
Infra Controls & Automation Architect

Alibaba Cloud • Washington

On-site
USD 142,000 - 234,000
Alibaba Cloud DevOps Engineer
Alibaba Cloud DevOps Engineer

Eitacies Inc • Austin (TX)

On-site
USD 110,000 - 170,000
Senior AI Cloud SRE — HPC & GPU Infra
Senior AI Cloud SRE — HPC & GPU Infra

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+4
Alibaba Cloud DevOps Engineer
Alibaba Cloud DevOps Engineer

EITACIES Inc. • Austin (TX)

On-site
USD 110,000 - 170,000
401(k)
Regional Infrastructure Program Lead – DC
Regional Infrastructure Program Lead – DC

Alibaba Cloud • Washington

On-site
USD 142,000 - 234,000
Site Reliability Engineer - PaaS
Site Reliability Engineer - PaaS

Tencent • Palo Alto (CA)

On-site
USD 115,000 - 160,000
Sign-on payment
Relocation package
Restricted stock units
+1
Site Reliability Engineer - Big Data Platform & Multi-Cloud
Site Reliability Engineer - Big Data Platform & Multi-Cloud

Apple Inc. • Austin (TX), Northern (KY)

Hybrid
USD 120,000 - 170,000