Alibaba-Site Reliability Engineer-Bellevue

Alibaba Group

Bellevue (WA)

On-site

USD 133,000 - 220,000

Full time

31 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical Insurance
Dental Insurance
Vision Insurance
401(k) Plan
Basic Life Insurance
Wellbeing benefits like FSA
Up to 12 paid holidays
Up to 15 paid vacation days
Up to 72 hours paid sick time

Job summary

Alibaba Group's AI Inference Platform seeks a Site Reliability Engineer to deploy, operate, and maintain a high-availability model service platform, including initial construction and ongoing changes. You will optimize monitoring, respond to incidents, and build automation to ensure stability at scale.

Requirements: 3+ years in SRE/DevOps or backend, strong Linux/networking, cloud experience (Alibaba Cloud a plus), and programming in Python/Go/Java/C++.

Qualifications

  • 3+ years of experience in SRE, DevOps, or backend development.
  • Strong knowledge of distributed system operations.
  • Fluency in Chinese and English for daily communication.
  • Experience with cloud computing and Alibaba Cloud is a plus.
  • Familiarity with MaaS or related concepts.

Responsibilities

  • Oversee deployment, operation, and maintenance of the platform.
  • Design monitoring metrics, log collection, and alerting strategies.
  • Participate in emergency response and root cause analysis (RCA).
  • Develop tools and scripts (Python/Go) to automate deployment, scaling, and fault recovery.
  • Build automated diagnostic toolchains to accelerate issue resolution.

Skills

Python
Golang
Java
C++

Tools

Kubernetes
Prometheus
Istio
Calico

Job description

We are the AI Inference Platform at Alibaba Group, committed to delivering a cutting-edge MaaS platform and toolkits for application development through technological innovation and engineering practices. Our team focuses on the fundamental R&D in model services, while also providing full-stack development that ranges from architecture design to model applications. Our goal is to build the industry's largest model service platform with excellent cost-efficiency, high performance, and enterprise-level reliability. By doing so, we aim to empower numerous enterprise clients to accelerate the development of model applications.

We are seeking a passionate and technically skilled Site Reliability Engineer (SRE) to join the our team. You will play a critical role in building and maintaining highly available, high-performance model service platform. Your responsibilities will include optimizing monitoring and alerting, incident response, troubleshooting customer issues, and developing automation systems to ensure the stability and reliability of the AI Inference Platform and system aplications.

Key Responsibilities
  • 1. Oversee the deployment, operation, maintenance, and continuous improvement of the standalone website and platform, including its initial construction and subsequent operational changes.
System Reliability
  • * Oversee the monitoring and alerting of our platform's and system aplications, rapidly diagnosing and resolving network, service, and hardware-level failures to meet SLA targets.
  • * Design and optimize monitoring metrics, log collection, and alerting strategies to enhance system observability.
  • * Participate in the emergency response and handling of online incidents, conduct root cause analysis (RCA), and drive long-term solutions to prevent recurrence.
Customer Issue Resolution
  • * Investigate and resolve customer-reported issues related to QoS of API service(e.g., latency, performance, optimization), collaborating with development teams to identify flaws in application clusters, edge networks, or infrastructure.
  • * Develop tools and scripts (Python/Go) to automate deployment, scaling, fault recovery, and other operational workflows.
  • * Build automated diagnostic toolchains to accelerate issue resolution and improve customer satisfaction.
Job Requirement
  • - 3+ years of experience in SRE, DevOps, or backend development, with expertise in distributed system operations. Experience in cloud computing, AI infrastructure, Alibaba Cloud is a plus.
  • - Experience programming with at least one modern language such as Python, Golang, Java, C++.
  • - Strong ability to work under pressure, manage critical incidents, and participate in an on-call rotation.
  • - Fluency in both Chinese and English for daily communication.
  • - Familiarity with MaaS or related knowledge.
  • - Deep knowledge of Linux systems, network protocols (TCP/HTTP), and databases, have deep understanding of cloud-native architecture design.
  • - Experience with large-scale containers, kubernetes cluster operation and maintenance, have strong professional knowledge of Cloud Native related components (e.g., Prometheus, Istio, Calico, etc.).
  • - Extensive experience in building large-scale monitoring systems and utilizing them for in-depth analysis and operations.

The pay range for this position at commencement of employment is expected to be between $133,200/year and $219,600/year. However, base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience.

If hired, employee will be in an “at-will position” and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.

  • medical insurance
  • dental insurance
  • vision insurance
  • 401(k) plan
  • basic life insurance
  • wellbeing benefits like FSA
  • up to 12 paid holidays
  • up to 15 paid vacation days for this position
  • up to 72 hours paid sick time (front-loaded) per calendar year
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineering (SRE) Specialist -Bellevue
Site Reliability Engineering (SRE) Specialist -Bellevue

Alibaba Cloud • Seattle (WA)

On-site
USD 133,200 - 219,600
Medical, dental, and vision insurance
401(k) plan
Paid holidays and vacation days
+1
SRE for AI Inference Platform — Reliability & Automation
SRE for AI Inference Platform — Reliability & Automation

Alibaba Group • Bellevue (WA)

On-site
USD 133,000 - 220,000
Medical Insurance
Dental Insurance
Vision Insurance
+6
Site Reliability Engineer-Bellevue
Site Reliability Engineer-Bellevue

Alibaba Cloud • Bellevue (WA)

On-site
USD 145,000 - 238,000
Medical insurance
Dental insurance
Vision insurance
+3
Site Reliability Engineer
Site Reliability Engineer

Quality Ai • Northern (KY)

Hybrid
USD 110,000 - 130,000
Competitive pay
Global opportunities
Technical training & certification
Site Reliability Engineer II
Site Reliability Engineer II

Akamai Technologies • Cambridge (MA)

On-site
USD 95,000 - 171,000
Flexible working options
Healthcare benefits
401K savings plan
+2
Staff Software Engineer, AI Reliability
Staff Software Engineer, AI Reliability

Anthropic • San Francisco (CA)

Hybrid
USD 325,000 - 485,000
Visa sponsorship available
Opportunity to work in dynamic teams
Flexible hybrid work policy
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Tiger Analytics, LLC • Washington

Hybrid
USD 120,000 - 160,000
Career development opportunities
Site Reliability Engineer
Site Reliability Engineer

Longbridge Singapore • New York (NY)

On-site
USD 140,000 - 190,000
Competitive compensation
Growth opportunities
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Tiger Analytics • Washington

Hybrid
USD 100,000 - 140,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

Remote
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3