Site Reliability Engineer - Cloud & Platform Resilience

Nexus Recruitment Group

Pasig

Vor Ort

PHP 900.000 - 1.300.000

Vollzeit

Vor 2 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Nexus Recruitment Group is seeking an experienced Site Reliability Engineer to ensure the reliability, scalability, and performance of cloud-native platforms in a regulated fintech context. You will blend software engineering with systems administration to automate ops, improve observability, and manage incident response.

The ideal candidate has hands-on experience with Kubernetes, cloud platforms, and modern observability stacks, and will collaborate with software, infrastructure, and security

Qualifikationen

  • Bachelor's degree in Computer Science, Information Systems, or related field.
  • 3-5 years of experience in SRE, DevOps, or platform engineering roles in enterprise or cloud-based environments.
  • Experience with infrastructure automation tools (Terraform, Ansible, Helm) and scripting languages (Python, Bash, Go).
  • Hands-on experience with Kubernetes, Docker, and container orchestration platforms.
  • Strong familiarity with observability stacks (Prometheus, Grafana, ELK, Datadog).
  • Experience with incident management, monitoring strategies, and system performance tuning.
  • Understanding of CI/CD pipelines (e.g., Jenkins, GitLab CI, ArgoCD) and GitOps principles.
  • Knowledge of system security, compliance, and availability in regulated industries is a plus.

Aufgaben

  • Maintain and improve platform reliability through automation, monitoring, and performance tuning.
  • Build tools and services that reduce manual operations and improve developer productivity.
  • Define and implement Service Level Indicators (SLIs), Objectives (SLOs), and Agreements (SLAs) for key services.
  • Design and maintain CI/CD pipelines, infrastructure as code (IaC), and container orchestration (e.g., Kubernetes).
  • Monitor application health, conduct incident response and postmortems, and drive root cause analysis.
  • Collaborate with software engineers, infrastructure teams, and cybersecurity to implement secure, resilient system designs.
  • Optimize usage of cloud resources (e.g., AWS, Azure, GCP) and enhance cost efficiency and fault tolerance.
  • Develop and maintain runbooks, operational playbooks, and documentation for ongoing support.

Kenntnisse

SRE experience
DevOps
Platform engineering
Automation

Ausbildung

Bachelor's degree in Computer Science / Information Systems

Tools

Terraform
Ansible
Helm
Python
Bash
Go
Kubernetes
Docker
Prometheus
Grafana
ELK
Datadog
Jenkins
GitLab CI
ArgoCD

Jobbeschreibung

Job Openings Site Reliability Engineer - Cloud & Platform Resilience

About the job Site Reliability Engineer - Cloud & Platform Resilience
Job Summary:

We are seeking an experienced and proactive Site Reliability Engineer (SRE) to ensure the reliability, scalability, and performance of critical digital platforms in a high-availability environment. This role combines software engineering and systems administration principles to automate operations, optimize monitoring, and manage incident response within cloud-native banking and fintech ecosystems.

The ideal candidate has a deep understanding of distributed systems, automation frameworks, and modern observability stacks, with experience operating production-grade environments in regulated industries.

Key Responsibilities:
  • Maintain and improve platform reliability through automation, monitoring, and performance tuning.
  • Build tools and services that reduce manual operations and improve developer productivity.
  • Define and implement Service Level Indicators (SLIs), Objectives (SLOs), and Agreements (SLAs) for key services.
  • Design and maintain CI/CD pipelines, infrastructure as code (IaC), and container orchestration (e.g., Kubernetes).
  • Monitor application health, conduct incident response and postmortems, and drive root cause analysis.
  • Collaborate with software engineers, infrastructure teams, and cybersecurity to implement secure, resilient system designs.
  • Optimize usage of cloud resources (e.g., AWS, Azure, GCP) and enhance cost efficiency and fault tolerance.
  • Develop and maintain runbooks, operational playbooks, and documentation for ongoing support.
Qualifications:
  • Bachelors degree in Computer Science, Information Systems, or related field.
  • 3-5 years of experience in SRE, DevOps, or platform engineering roles in enterprise or cloud-based environments.
  • Proficiency in infrastructure automation tools (Terraform, Ansible, Helm) and scripting languages (Python, Bash, Go).
  • Hands-on experience with Kubernetes, Docker, and container orchestration platforms.
  • Strong familiarity with observability stacks (Prometheus, Grafana, ELK, Datadog, etc.).
  • Experience with incident management, monitoring strategies, and system performance tuning.
  • Understanding of CI/CD pipelines (e.g., Jenkins, GitLab CI, ArgoCD) and GitOps principles.
  • Knowledge of system security, compliance, and availability in regulated industries is a plus.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Blackfort Consulting, Inc.. • Pateros

Vor Ort
PHP 1.200.000 - 2.000.000
Site Reliability Engineer (SRE) - AWS & Kurbernetes
Site Reliability Engineer (SRE) - AWS & Kurbernetes

Lewis Personnel Management • Metro Manila

Remote
PHP 900.000 - 1.500.000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Astek • Santo Niño 1st

Vor Ort
PHP 1.100.000 - 1.900.000
SRE (Site Reliability Engineer)
SRE (Site Reliability Engineer)

GCash • Manila

Vor Ort
PHP 892.800 - 1.116.000
Opportunity for career growth and development
Dynamic collaborative team environment
Highly competitive compensation and benefits package
SRE: Cloud Reliability & Automation for Fintech Platforms
SRE: Cloud Reliability & Automation for Fintech Platforms

Nexus Recruitment Group • Pasig

Vor Ort
PHP 900.000 - 1.300.000
Technical Lead - Site Reliability Engineering
Technical Lead - Site Reliability Engineering

LSEG • Taguig

Vor Ort
PHP 4.914.004 - 7.371.007
Healthcare
Retirement planning
Paid volunteering days
+1
Site Reliability Engineer
Site Reliability Engineer

TymblHub • Hinoba-an

Vor Ort
PHP 900.000 - 1.500.000
Staff Site Reliability Engineer – Cloud Efficiency
Staff Site Reliability Engineer – Cloud Efficiency

Super • España

Vor Ort
PHP 1.200.000 - 1.600.000
Medical / Health Insurance
Employee Assistance Programme
Lead Site Reliability Engineer
Lead Site Reliability Engineer

iScale Solutions • Philippinen

Remote
PHP 7.538.000 - 11.307.000
Competitive salary
Health coverage
Vacation & sick leave
+7
Senior Platform Engineer
Senior Platform Engineer

Our Clients • Pasay

Vor Ort
PHP 1.200.000 - 1.800.000