Infra Operations SRE -Sunnyvale

Alibaba Cloud

Sunnyvale (CA)

On-site

USD 145,000 - 238,000

Full time

27 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Alibaba Cloud is seeking an Infrastructure Operations SRE in Sunnyvale, CA to own infra operations and the stability platform engineering. You will manage Kubernetes clusters, databases, middleware, and application releases to keep critical platforms running in a dedicated cloud environment.

You will monitor, troubleshoot, and scale systems; perform backups, DR drills, and security hardening; and drive incident response and postmortems.

Qualifications

  • Bachelor's degree in CS or a related field and 3+ years in infra operations or SRE.
  • Solid Linux administration and scripting skills with Kubernetes and cloud-native monitoring.
  • Experience with MySQL/Redis/Kafka/RocketMQ/Nacos and database reliability.
  • Strong development skills in Python, Go, or Java and automation tooling.
  • Familiar with release, gateway config, and change/rollback practices.
  • Security awareness, incident coordination, and ownership mindset.

Responsibilities

  • Monitor Kubernetes clusters, databases, and middleware; manage capacity and scaling.
  • Troubleshoot failures, perform fault isolation, and ensure rapid recovery.
  • Maintain databases and middleware with backups, DR drills, and DR plans.
  • Handle application initialization, versions, and configuration changes.
  • Manage permissions, platform consoles, and account synchronization.
  • Drive security and incident response, postmortems, and disaster recovery readiness.
  • Oversee data onboarding for infra stability platforms and ensure data accuracy.

Skills

Linux administration
Shell scripting
Python
Kubernetes
Cloud-native monitoring
Database/Middleware
Go/Java development
Change management

Education

Bachelor's degree in Computer Science or related field

Tools

MySQL
Redis
Kafka
RocketMQ
Nacos
MongoDB

Job description

We are looking for a Infrastructure Operations SRE to own Infra infrastructure operations and stability platform engineering. The role covers day-to-day operations of Kubernetes clusters, databases and middleware, and application releases and changes, as well as data onboarding and operations for Infra stability platforms, ensuring business platforms run stably in an independent cloud environment.

Responsibilities
  • Perform daily monitoring, health checks, and capacity management of Kubernetes clusters; track node CPU/memory/disk utilization, Pod status, and database/middleware metrics; handle alerts, collect logs to identify risks, and execute scaling or disk cleanup as needed;
  • Respond to monitoring alerts and business feedback; troubleshoot failures of K8s nodes, middleware components, and application instances; perform fault isolation, hardware decommission and repair, workload migration, and capacity expansion to ensure fast recovery;
  • Maintain databases and middleware including MySQL/MongoDB/PostgreSQL/Kafka/RocketMQ/Redis/Nacos; perform backup/restore and disaster recovery drills; manage gateway configuration including domains, routing rules, and certificates;
  • Own application initialization, version releases and configuration changes, elastic scaling for traffic fluctuations, and application migration during node failures; maintain a high change success rate;
  • Manage permission provisioning for middleware and databases, servers/network/Infra consoles, and account synchronization to keep daily operations requests handled efficiently;
  • Drive production safety and security: vulnerability remediation and hardening, incident coordination, severity assessment and postmortems, emergency response, and day-to-day operations support;
  • Own data onboarding and operations for Infra stability platforms, covering device management, energy monitoring, emergency drills, and key process control; support platform iteration and ensure data completeness and accuracy;
  • Ensure the stability of operations platforms deployed in an independent cloud environment; build platform monitoring, emergency playbooks, and incident response mechanisms; drive high-availability architecture and disaster recovery capability.
Job Requirements

Minimum qualifications:

  • Bachelor's degree or above in Computer Science or a related field, with 3+ years of experience in infrastructure operations, SRE, or Infra operations;
  • Solid Linux administration and Shell/Python scripting skills; familiar with Kubernetes orchestration and common cloud-native monitoring; hands-on experience with cluster health checks, troubleshooting, and capacity management;
  • Proficient in at least one mainstream database or middleware such as MySQL, Redis, or Kafka, with an understanding of high-availability architecture, backup/restore, and common performance issue handling;
  • Solid development skills in at least one of Python/Go/Java, able to independently build operations automation tools and platform features;
  • Familiar with application release/change processes and gateway configuration management, with disciplined change and rollback practices;
  • Strong security awareness; familiar with permission management, vulnerability remediation, and incident response fundamentals;
  • Strong sense of ownership, excellent cross-team collaboration and communication skills, and the ability to respond effectively to emergencies.
Preferred Qualifications
  • Experience in Infra / Infra infrastructure operations;
  • Experience building or developing stability platforms or operations platforms;
  • Experience operating large-scale multi-cluster or multi-site environments.

The pay range for this position at commencement of employment is expected to be between $145,200/year and $238,000/year. However, base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience.

If hired, employee will be in an “at-will position” and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Infrastructure – Site Reliability Engineer (SRE)-Sunnyvale
Cloud Infrastructure – Site Reliability Engineer (SRE)-Sunnyvale

Alibaba Cloud • Sunnyvale (CA)

On-site
USD 104,000 - 171,000
SRE/Platform Engineer
SRE/Platform Engineer

Stash Talent Services • Virginia (MN)

Remote
Devops SRE
Devops SRE

Tata Consultancy Services • Austin (TX)

On-site
USD 110,000 - 130,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Charlotte (NC)

On-site
USD 125,000 - 168,000
Discretionary incentive eligible
Annual discretionary plan
Senior Software Engineer - SRE
Senior Software Engineer - SRE

Socure • City of Albany (NY)

On-site
USD 160,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Veritas Search Group • Tustin (CA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Hobbsnews • Chandler (AZ), Northern (KY)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Koitecc Solutions • Plano (TX)

On-site
USD 125,000 - 168,000
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Bank of America • Plano (TX)

On-site
USD 125,000 - 168,000
Industry-leading benefits
Discretionary incentive plan