Cloud SRE – Kubernetes, Observability & Reliability

TP-Link Systems Inc.

Irvine (CA)

On-site

USD 100,000 - 140,000

Full time

5 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Free snacks and drinks
Fridays lunch provided
Fully paid medical, dental, and vision
401k contributions
Bi-annual reviews and pay raises
Wellness benefits and gym membership

Job summary

TP-Link Systems Inc. in Irvine, CA is seeking a Site Reliability Engineer to help scale our cloud platform with robust reliability and security. You will work with Cloud and DevOps teams to operate microservices, implement observability, and drive incident response.

You will influence architecture decisions, develop automation scripts, and ensure compliance with security standards while mentoring junior engineers and contributing to on-call coverage.

Qualifications

  • Bachelor's degree in CS/IT or a related field.
  • 1–3 years of experience as a Site Reliability Engineer or in a related role.
  • Proficiency in Java, Python, Bash, or PowerShell.
  • Hands-on experience in SRE, DevOps, cloud operations, and cloud security best practices.
  • Basic knowledge of IAM, network security, application security, and data protection.
  • Ability to write technical documentation and implement compliance requirements.

Responsibilities

  • Assist in implementing and operating Microservices on Kubernetes cloud-based platforms.
  • Collaborate with Cloud Technical Development and DevOps teams to deploy services to the Multi-Cloud Platform.
  • Conduct Load Tests and Chaos Tests to ensure scalability and reliability of microservices.
  • Build observability for Microservices and cloud platforms like AWS, OCI, Azure, and GCP.
  • Contribute to writing and executing disaster recovery plans with Development and DevOps teams.
  • Help analyze and resolve production risks caused by resources like CPU/memory/HPA scheduling.
  • Write and maintain scripts for automation using Python, Go, or Bash.
  • Issue KPIs (SLA/SLO/SLI) with development teams to better understand business impact.
  • Create and maintain architecture diagrams, design docs, and SOPs.
  • Ensure adherence to security/compliance standards including ISO27001, SOC2, GDPR.
  • Participate in incident response and post-incident analysis.
  • Contribute to product/technology selection and POCs.
  • Mentor junior staff and participate in on-call rotations.

Skills

Java
Python
Bash
PowerShell
SRE
DevOps
Cloud operations
Kubernetes
Security best practices
Documentation

Education

Bachelor's degree in Computer Science, Information Technology, or a related field

Tools

Kubernetes
AWS
Azure
GCP
OCI
CI/CD tools

Job description

TP-Link Systems Inc. in Irvine, CA is seeking a Site Reliability Engineer to help scale our cloud platform with robust reliability and security. You will work with Cloud and DevOps teams to operate microservices, implement observability, and drive incident response.

You will influence architecture decisions, develop automation scripts, and ensure compliance with security standards while mentoring junior engineers and contributing to on-call coverage.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE - Observability & Cloud Reliability
Senior SRE - Observability & Cloud Reliability

Cisco Systems, Inc. • San Francisco (CA)

On-site
USD 168,000 - 245,000
Medical benefits
401(k) matching
Parental leave
+1
Cloud SRE/DevOps Engineer — Automation & Reliability
Cloud SRE/DevOps Engineer — Automation & Reliability

Cisco Systems, Inc. • Milpitas (CA)

On-site
USD 155,000 - 223,000
Senior Observability SRE – Kubernetes & Cloud
Senior Observability SRE – Kubernetes & Cloud

Cisco • San Francisco (CA)

On-site
USD 168,000 - 245,000
Site Reliability Engineer
Site Reliability Engineer

TP-Link Corporation Limited • Irvine (CA)

On-site
USD 100,000 - 140,000
Free snacks and drinks
Fully paid medical insurance
401k contributions
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

TP-Link Corporation Limited • Irvine (CA)

On-site
USD 140,000 - 180,000
Free snacks and drinks
Fully paid medical, dental, and vision insurance
Contributions to 401k funds
+2
Cloud SRE & Automation Engineer
Cloud SRE & Automation Engineer

Cisco Systems, Inc • Milpitas (CA)

On-site
USD 155,000 - 223,000
Medical, dental and vision insurance
401(k) with Cisco matching
Paid parental leave
+1
Cloud SRE & Resiliency Engineer - Observability & Chaos
Cloud SRE & Resiliency Engineer - Observability & Chaos

United States Digital Space LLC • United States

Remote
USD 140,000 - 180,000
Attractive remuneration package and "_
Senior SRE: Observability & Cloud Platform Engineer
Senior SRE: Observability & Cloud Platform Engineer

Cisco • San Francisco (CA)

On-site
USD 168,000 - 245,000
Medical, dental, and vision insurance
401(k) with Cisco matching
Paid parental leave
+3
Senior Site Reliability Engineer – Observability
Senior Site Reliability Engineer – Observability

Cisco • Boise (ID)

On-site
USD 120,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

TP-Link Systems Inc. • Irvine (CA)

On-site
USD 100,000 - 140,000
Free snacks and drinks
Fridays lunch provided
Fully paid medical, dental, and vision
+3