Site Reliability Engineer (SRE)

PT. HTC Global Software Services

Jakarta Utara

On-site

IDR 550,000,000 - 900,000,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

HTC Global Services is seeking an experienced Site Reliability Engineer in Jakarta to maintain and optimize Tencent Cloud production services. The role focuses on ensuring 99.99%+ availability, running highly distributed systems and collaborating with Tencent product teams.

You will define SLIs/SLOs, build monitoring and alerting, lead incident responses and implement self-healing automation, while participating in on-call rotations for critical cloud services.

Qualifications

  • Minimum 5 years experience as Site Reliability Engineer (SRE).
  • Core Skills: Linux, Kubernetes, containers, Python/Go/C++, distributed systems, observability, automation, CI/CD, cloud infrastructure, incident management and performance engineering.

Responsibilities

  • Maintain the reliability, availability, scalability and performance of Tencent Cloud's production services and infrastructure.
  • Engineer Tencent Cloud services to achieve 99.99%+ availability targets, depending on the service SLA.
  • Operate and improve highly distributed production systems supporting large numbers of customers.
  • Define and monitor SLIs, SLOs, error budgets and service health indicators.
  • Build monitoring, logging, tracing and automated alerting capabilities.
  • Detect and resolve live production incidents across compute, network, storage and database environments.
  • Lead incident response, root-cause analysis and post-mortem activities.
  • Develop self-healing and automated remediation mechanisms.
  • Automate repetitive operational activities and reduce engineering toil.
  • Perform capacity planning, performance optimization and resilience testing.
  • Design failover and disaster-recovery mechanisms.
  • Participate in production/on-call rotations for critical cloud services.
  • Work directly with Tencent Cloud product engineering teams to improve service reliability

Skills

Linux
Kubernetes
containers
Python
Go
C++
distributed systems
observability
automation
CI/CD
cloud infrastructure
incident management
performance engineering

Job description

Maintain the reliability, availability, scalability and performance of Tencent Cloud's production services and infrastructure.

Engineer Tencent Cloud services to achieve 99.99%+ availability targets, depending on the service SLA.

Operate and improve highly distributed production systems supporting large numbers of customers.

Define and monitor SLIs, SLOs, error budgets and service health indicators.

Build monitoring, logging, tracing and automated alerting capabilities.

Detect and resolve live production incidents across compute, network, storage and database environments.

Lead incident response, root-cause analysis and post-mortem activities.

Develop self-healing and automated remediation mechanisms.

Automate repetitive operational activities and reduce engineering toil.

Perform capacity planning, performance optimization and resilience testing.

Design failover and disaster-recovery mechanisms.

Participate in production/on-call rotations for critical cloud services.

Work directly with Tencent Cloud product engineering teams to improve service reliability

Requirements:

Minimum 5 years experience as Site Reliability Engineer (SRE)

Core Skills: Linux, Kubernetes, containers, Python/Go/C++, distributed systems, observability, automation, CI/CD, cloud infrastructure, incident management and performance engineering.

Ready to join ASAP is preferably
Willing work in shifting (8 hours x 5 days/week)

Information Technology Services 1,001-5,000 employees

HTC Global Services

Established in 1990, HTC Global Services is an Inc. 500 Hall of Fame company and one of the fastest growing Asian American companies in the US with headquarters in Troy, Michigan. A global provider of IT Solutions and Business Process Outsourcing services, HTCs client base spans several Global 2000 organizations. HTC is committed to providing solutions that translate into tangible business outcomes for our customers. HTC manages IT environments, IT applications, and business processes of customers, focusing on providing transformational benefits.

Mission :

We are a global IT solutions provider adding value to our clients and people through emerging technologies. We are dedicated to the success of our clients, employees, business partners, suppliers, community, and stakeholders.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer — 99.99% Cloud Availability
Senior Site Reliability Engineer — 99.99% Cloud Availability

PT. HTC Global Software Services • Jakarta Utara

On-site
IDR 550,000,000 - 900,000,000
Cloud Infrastructure Engineer
Cloud Infrastructure Engineer

PT. HTC Global Software Services • Jakarta Utara

On-site
IDR 200,000,000 - 500,000,000
Cloud Operations ( GCP & Kubernetes )
Cloud Operations ( GCP & Kubernetes )

PT. HTC Global Software Services • Jakarta Utara

On-site
IDR 279,000,000 - 502,200,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

AccelByte • Sleman

On-site
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

AccelByte • Sleman

On-site
Site Reliability Engineer
Site Reliability Engineer

Pengiklan Anonim • Jakarta Utara

On-site
IDR 446,400,000 - 781,200,000
Hyperscale Cloud Infra Engineer - SDN, K8s & Automation
Hyperscale Cloud Infra Engineer - SDN, K8s & Automation

PT. HTC Global Software Services • Jakarta Utara

On-site
IDR 200,000,000 - 500,000,000
Site Reliability Engineer (Junior)
Site Reliability Engineer (Junior)

CloudMile • Jakarta Pusat

On-site
IDR 178,560,000 - 245,520,000
SRE: Architect Reliable, Scalable Infrastructure
SRE: Architect Reliable, Scalable Infrastructure

StraitsX Group • Jakarta Pusat

On-site
IDR 200,000,000 - 300,000,000
Site Reliability Engineer New Jakarta, Jakarta, Indonesia
Site Reliability Engineer New Jakarta, Jakarta, Indonesia

StraitsX Group • Jakarta Pusat

On-site
IDR 200,000,000 - 300,000,000