Site Reliability Engineer (SRE) - AWS & Kurbernetes

Lewis Personnel Management

Metro Manila

Remote

PHP 900,000 - 1,500,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Lewis Personnel Management seeks a Site Reliability Engineer (SRE) to maintain and evolve production infrastructure across AWS and Kubernetes. The role emphasizes reliability, scalability, automation, and incident management, with responsibilities for ML infrastructure.

You will operate a 24/7 Kubernetes-based production stack, lead incident responses, implement IaC with Terraform, Helm, and ArgoCD, and collaborate with data scientists to operationalize ML models at scale.

Qualifications

  • 2+ years of AWS experience with production use.
  • Hands-on experience operating AWS EKS in production.
  • Software engineering experience applying DevOps/SRE practices, including 1+ year in a technical leadership capacity.
  • Experience operating production systems at scale with incident response.
  • Experience with ML workflows, MLOps, or ML infrastructure.

Responsibilities

  • Maintain and improve a 24/7 production environment on Kubernetes.
  • Apply SRE and DevOps practices to improve reliability and automation.
  • Monitor systems, manage configuration, performance, and scalability.
  • Lead incident response, troubleshooting, and postmortems.
  • Manage AWS infrastructure including EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3.
  • Build CI/CD pipelines and infrastructure as code using Terraform, Helm, and ArgoCD.
  • Identify reliability risks and implement scalable solutions.

Skills

AWS & Kubernetes
SRE/DevOps
Programming & ML Infra

Tools

Terraform
Helm
ArgoCD
AWS EKS

Job description

Location: Remote — Dayshift
Employment Type: Full-Time

Job Summary

We are seeking a Site Reliability Engineer (SRE) to maintain and evolve production infrastructure across AWS and Kubernetes. The role is primarily focused on SRE, reliability, scalability, automation, and incident management, with additional responsibility for applying SRE practices to machine learning infrastructure, model deployment, and MLOps workflows.

Key Responsibilities
SRE & Cloud Infrastructure
  • Maintain and improve a 24/7 production environment running on Kubernetes.

  • Apply SRE and DevOps practices to improve reliability, automation, and engineering efficiency.

  • Monitor systems proactively and manage configuration, performance, and scalability.

  • Lead incident response, troubleshooting, root-cause analysis, and postmortems.

  • Manage and evolve AWS infrastructure including EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3.

  • Build and maintain CI/CD pipelines and infrastructure as code using Terraform, Helm, and ArgoCD.

  • Identify reliability risks and implement solutions for large-scale production systems.

Machine Learning Infrastructure
  • Apply SRE principles to ML infrastructure, including model serving, training pipelines, and data systems.

  • Support ML model deployment pipelines and MLOps practices.

  • Monitor ML model performance and implement observability and alerting for ML systems.

  • Collaborate with data scientists and product teams to operationalize ML models at scale.

  • Support ML workloads running on Kubernetes and AWS.

Required Experience
  • 2+ years of AWS experience, including production use of EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3.

  • Hands-on experience operating AWS EKS in production.

  • Software engineering experience applying DevOps/SRE practices, including 1+ year in a technical leadership capacity.

  • Experience operating production systems at scale, including incident response and troubleshooting.

  • Experience with machine learning workflows, MLOps, or ML infrastructure.

Nice-to-Have Experience
  • Startup experience.

  • Experience deploying and managing ML workloads on Kubernetes.

  • Experience with ML model serving, inference, or ML pipeline infrastructure.

Required Skills
  1. AWS & Kubernetes: Advanced hands-on experience with AWS infrastructure and Kubernetes, particularly EKS.

  2. SRE/DevOps: Production reliability engineering, incident management, CI/CD, infrastructure as code, and automation.

  3. Programming & ML Infrastructure: Proficiency in at least one of Python, Ruby, Elixir, Go, JavaScript, or Rust, plus Python experience with ML-adjacent tooling.

Preferred Skills
  • Terraform or Pulumi

  • Helm and ArgoCD

  • Prometheus and Grafana

  • YAML and Bash

  • Container and hypervisor fundamentals

  • PostgreSQL and MongoDB

  • Kafka and event-driven architecture

  • Security, PCI-DSS, GDPR, and digital forensics

  • TensorFlow Serving, TorchServe, or Triton

  • Feature stores, experiment tracking, and model registry tools

  • ML model deployment and inference serving

  • MLOps and ML observability

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Full-Time: Senior Site Reliability Engineer (SRE) - URGENT
Full-Time: Senior Site Reliability Engineer (SRE) - URGENT

Lewis Personnel Management • Manila

On-site
PHP 1,800,000 - 2,400,000
Site Reliability Engineer (Machine Learning) - REMOTE
Site Reliability Engineer (Machine Learning) - REMOTE

Lewis Personnel Management • Metro Manila

Remote
PHP 1,200,000 - 1,700,000
Remote SRE: AWS & Kubernetes, ML Infra Focus
Remote SRE: AWS & Kubernetes, ML Infra Focus

Lewis Personnel Management • Metro Manila

Remote
PHP 900,000 - 1,500,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Blackfort Consulting, Inc.. • Pateros

On-site
PHP 1,200,000 - 2,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Astek • Santo Niño 1st

On-site
PHP 1,100,000 - 1,900,000
Site Reliability Engineer
Site Reliability Engineer

TymblHub • Hinoba-an

On-site
PHP 900,000 - 1,500,000
Site Reliability Application Engineers (AI Platform)
Site Reliability Application Engineers (AI Platform)

Astek • Santo Niño 1st

On-site
PHP 600,000 - 1,000,000
Senior Site Reliability Engineer (SRE) – Kubernetes
Senior Site Reliability Engineer (SRE) – Kubernetes

Accenture • Cebu City

On-site
PHP 700,000 - 1,100,000
Staff Site Reliability Engineer – Cloud Efficiency
Staff Site Reliability Engineer – Cloud Efficiency

Super • España

On-site
PHP 1,200,000 - 1,600,000
Medical / Health Insurance
Employee Assistance Programme
Senior Devops Engineer
Senior Devops Engineer

TymblHub • Hinoba-an

On-site
PHP 1,200,000 - 2,400,000