Senior Site Reliability Engineer

BP PLC

Kuala Lumpur

Hybrid

MYR 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Life and health insurance
Learning opportunities

Job summary

BP plc is seeking an experienced Site Reliability Engineer in Kuala Lumpur to enhance the reliability and scalability of our cloud-based platforms. You will collaborate with product and engineering teams to implement automation, observability, and secure deployment practices in production.

Lead incident investigations, drive root cause analysis, and propagate reliable patterns across services. A strong emphasis on resilience, automation, and mentoring will help lift engineering standards

Qualifications

  • Proficiency in Python, Ruby or Go for automation and tooling.
  • Experience with CI/CD pipelines and IaC practices.
  • Strong troubleshooting across cloud infrastructure and services.

Responsibilities

  • Improve reliability, availability, performance and scalability of cloud-based apps.
  • Build and enhance automation to reduce manual operational work.
  • Develop and maintain observability with monitoring, logging, alerts and tracing.
  • Investigate production incidents and implement preventive improvements.
  • Strengthen security and resilience of cloud infrastructure.
  • Mentor engineers and promote best practices in reliability engineering.

Skills

Python
Ruby
Go
CI/CD
Monitoring

Education

Bachelor's degree in Computer Science, Engineering or related

Tools

Terraform
CloudFormation
Kubernetes
Linux

Job description

**Role Summary**As a Site Reliability Engineer, you will be responsible for improving the reliability, resilience and operational effectiveness of our technology platforms and services. You will work closely with engineering and product teams to ensure systems are highly available, scalable, secure and supportable in production. You will use software engineering, automation and modern cloud practices to reduce manual effort, improve performance and strengthen production reliability. **Key Responsibilities*** Improve the reliability, availability, performance and scalability of cloud-based applications and services.* Design and implement automation to reduce manual operational activities and improve engineering efficiency.* Build and improve monitoring, logging, alerting and observability across production systems.* Investigate complex production issues and drive improvements to prevent recurring failures.* Improve system resilience, recovery and operational readiness.* Build and improve CI/CD pipelines to enable reliable and repeatable software delivery.* Develop and maintain infrastructure using Infrastructure as Code and automation.* Identify reliability risks, operational gaps and technical debt and drive appropriate improvements.* Improve cloud infrastructure security and operational practices.* Develop reusable engineering patterns and mentor engineers across teams. **Required Experience and Qualifications*** 7+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, Cloud Engineering, Software Engineering or related technical disciplines, with strong experience operating production systems.* Strong understanding of cloud infrastructure security, including identity and access management, least privilege, network security, secrets management and secure configuration.* Understanding of security practices within CI/CD pipelines and Infrastructure as Code.* Able to independently investigate and resolve complex technical and production problems.* Strong communication and collaboration skills across engineering, product, security and operational teams.* Able to influence engineering practices, drive technical improvements and mentor other engineers.* Degree in Computer Science, Engineering or a related discipline, or equivalent professional experience.* Relevant cloud or engineering certifications are beneficial but not essential. *Technical Skills** Strong experience operating and improving production systems in cloud-based environments.* Strong troubleshooting skills across applications, infrastructure, networking and cloud services.* Experience managing system reliability, scalability, availability and performance.* Strong knowledge of monitoring, logging, alerting and production diagnostics.* Experience with incident investigation, root cause analysis and operational improvement.* Good understanding of distributed systems, resilience and recovery practices. *Software Engineering** Strong programming and scripting skills using Python, Ruby, Go or equivalent technologies.* Strong understanding of software engineering practices including source control, code review, automated testing and software delivery.* Strong experience designing, building and maintaining CI/CD pipelines.* Strong experience with deployment automation, release management and rollback or recovery practices.* Experience building automation, tooling and reusable engineering solutions. *Cloud Infrastructure** Strong hands-on experience with AWS, Microsoft Azure or equivalent cloud platforms.* Strong experience with Infrastructure as Code, using technologies such as Terraform, CloudFormation or equivalent.* Strong knowledge of Linux/Unix systems, networking and infrastructure troubleshooting.* Experience with containers and modern cloud application infrastructure.* Experience with observability technologies such as Prometheus, Grafana, OpenTelemetry or cloud-native equivalents.* Good understanding of cloud services including compute, networking, storage, databases, identity and messaging. **Skills That Set You Apart*** Experience improving reliability and operational practices across multiple services or engineering teams.* Experience with automated recovery, resilience engineering or self-service platform capabilities.* Experience operating large-scale or highly available distributed systems.* Strong understanding of cloud-native engineering and modern operational practices. **About bp**At bp, we provide the following environment and benefits to you:* A company culture where we respect our diverse and unified teams, where we are proud of our achievements and where fun and the attitude of giving back to our environment are highly valued.* Possibility to join our social communities and networks* Learning opportunities and other development opportunities to craft your career path* Life and health insurance, medical care packageAnd many other benefits. We are an equal opportunity employer and value diversity at our company. We do not discriminate based on race, religion, colour, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, perform crucial job functions, and receive other benefits and privileges of employment.**Travel Requirement**No travel is expected with this role**Relocation Assistance:**This role is not eligible for relocation**Remote Type:**This position is a hybrid of office/remote working**Skills:**Agility core practices, Agility core practices, Analytics, API and platform design, Business Analysis, Cloud Platforms, Coaching, Communication, Configuration management and release, Continuous deployment and release, Data Structures and Algorithms (Inactive), Digital Project Management, Documentation and knowledge sharing, Facilitation, Information Security, iOS and Android development, Mentoring, Metrics definition and instrumentation, NoSql data modelling, Relational Data Modeling, Risk Management, Scripting, Service operations and resiliency, Software Design and Development, Source control and code management {+ 4 more}
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Praxy • Kuala Lumpur

Hybrid
MYR 240,000 - 360,000
Life and health insurance
Learning opportunities
Social communities and networks
Site Reliability Engineer
Site Reliability Engineer

Experian Group • Cyberjaya

Hybrid
MYR 180,000 - 240,000
Great compensation
Discretionary bonus
Hybrid work arrangement
+1
System Reliability Engineer, Consultant
System Reliability Engineer, Consultant

AIA Hong Kong and Macau • Kuala Lumpur

On-site
MYR 70,000 - 90,000
Senior Cloud Reliability Engineer | Observability & Automation
Senior Cloud Reliability Engineer | Observability & Automation

Praxy • Kuala Lumpur

Hybrid
MYR 240,000 - 360,000
Life and health insurance
Learning opportunities
Social communities and networks
Senior Platform Enablement Engineer
Senior Platform Enablement Engineer

bp • Kuala Lumpur

Hybrid
MYR 150,000 - 210,000
Senior Platform Enablement Engineer
Senior Platform Enablement Engineer

Praxy • Kuala Lumpur

Hybrid
MYR 180,000 - 240,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Ryt Bank • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Experian Asia Pacific • Cyberjaya

On-site
MYR 180,000 - 240,000
Great compensation package
Hybrid work arrangement
Equal opportunities employer
Senior Devops Engineer
Senior Devops Engineer

ABACUS STRATEGIC ADVISORY SDN BHD • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Senior Platform Enablement Engineer
Senior Platform Enablement Engineer

BP PLC • Kuala Lumpur

Hybrid
MYR 180,000 - 240,000
Hybrid work model