Stand out for this role — generate a tailored resume and cover letter in about a minute.
Lewis Personnel Management is seeking an experienced Site Reliability Engineer (SRE) with ML expertise for a fully remote role in Metro Manila, Philippines. You will maintain and scale a 24/7 production environment on AWS and Kubernetes for a Japan-based restaurant platform.
Responsibilities include designing IaC with Terraform, Helm, and ArgoCD; managing AWS services including EKS, EC2, RDS, and S3; driving incident response and observability; and enabling ML model deployment pipelines on
Location: Metro Manila, Philippines | Remote
Employment type: Full-time | Day Shift
This is a fully remote Site Reliability Engineer (SRE) position with Machine Learning Expertise, partnering with Japan's leading restaurant reservation management platform. You will maintain and scale a 24/7 high-traffic production environment built on AWS and Kubernetes, while bringing SRE and MLOps discipline to machine learning pipelines and inference services.
Maintain and scale a 24/7 high-traffic production environment built on AWS and Kubernetes
Design, build, and evolve Infrastructure as Code (IaC) using Terraform, Helm, and ArgoCD
Manage AWS core services, including EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3
Drive incident response, proactive monitoring, configuration management, and postmortem processes
Implement DevOps best practices to boost developer productivity, platform scalability, and system reliability
Apply SRE discipline to ML infrastructure, ensuring high availability, observability, and performance for model training and inference pipelines
Support, streamline, and automate ML model deployment pipelines on Kubernetes
Build monitoring, logging, and alerting systems tailored to production ML models
Partner with data science and engineering teams to operationalize and scale intelligent features across the platform
Minimum 2 years of hands-on experience with AWS, with significant depth in AWS EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3
Deep experience deploying, managing, and scaling production workloads on Kubernetes
Strong familiarity with MLOps practices and hands-on Python experience with ML-adjacent tooling (e.g., inference serving, model deployment, ML pipelines)
Experience in direct software engineering under DevOps/SRE practices, including at least 1 year in a Technical Lead capacity
Current proficiency in Python (preferred) or Ruby, Elixir, Go, JavaScript, or Rust
Strong skills in configuration management (YAML, Bash) along with containerization and hypervisor fundamentals
Track record of operating large-scale production systems with a deep understanding of fault-tolerance and distributed systems failure modes
Experience in fast-paced startup environments
Infrastructure as Code & GitOps tools: Terraform, Pulumi, ArgoCD
Observability stacks: Prometheus, Grafana
Job Ref: RYAN - TT