Senior AI Platform Engineer

OpsWerks

Mandaluyong

On-site

PHP 800,000 - 1,100,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

OpsWerks in Manila is seeking a Senior AI Platform Engineer to operate, maintain, and improve AI platforms on Kubernetes (on-premise and cloud). You will deploy releases via GitOps, monitor health with observability tools, and drive reliability across production workloads.

You will mentor junior engineers in Python and lead adoption of MLOps best practices to shape the team roadmap.

Qualifications

  • 3+ years supporting production workloads on ML platforms.
  • 5+ years hands-on AI/ML lifecycle including DevOps/MLOps.
  • 5+ years of Python development for AI/ML workflows and data pipelines.
  • Practical Kubernetes experience in cloud flavors (AWS-EKS, GCP-GKE).
  • Knowledge of microservice architectures and service communication.
  • Strong troubleshooting for crashes, resource contentions, latency and scaling.
  • Well-rounded logging, metrics and observability skills.
  • Mentorship in Python and MLOps practices.

Responsibilities

  • Deploy releases and configuration changes via GitOps/DevOps.
  • Monitor platform health using logs, metrics and observability tools.
  • Improve observability, tooling, automation, and reliability practices.
  • Participate in incident response, RCA and 24x7 rotations.
  • Lead adoption of MLOps and mentor junior engineers.

Skills

Python
MLOps
Kubernetes
Troubleshooting
Observability

Tools

Kubeflow AI Tools
Ray.IO
Jupyter Notebooks
AWS SageMaker
AWS EKS
GCP GKE

Job description

Job Description:

Your Role

As a Senior AI Platform Engineer, you will be responsible for operating, maintaining, and continuously improving the company's AI platforms running on Kubernetes (On-premise and/or on AWS/GCP) - similar on the AIoEKS (AI on EKS) deployment frameworks and Kubeflow's Machine Learning Toolkit

Platform Ownership & Operations
  • Deploy new releases and configuration changes through GitOps/DevOps
  • Monitor platform and service health using logs, metrics, and observability tools
  • Improve platform observability, operational tooling/automations, self-service capabilities and reliability practices to reduce recurring issues
  • Participate in incident response, root cause analysis and 24x7 operational rotations
User & Developer Experience
  • Investigate & troubleshoot user concerns by either correlating them to system-related issues, breaking integrations and/or user-specific errors/misconfigurations up to recommending/executing resolutions
  • Advocate for platform standards, security best practices, and operational excellence
Collaboration and Leadership
  • Provide structured Python mentorship to junior engineers, focusing on strong fundamentals and bridge foundational Python knowledge toward MLOps competencies
  • Lead the adoption of MLOps best practices for the team
  • Influence the team roadmap by identifying gaps in tooling, skills, and processes required to support production-grade AI systems
Your Qualifications
  • 3+ years of experience supporting production workloads/platforms (Ray.IO, Jupyter Notebooks, AWS SageMaker, Kubeflow AI Tools or an AI-related equivalent)
  • 5+ years of hands-on experience AI/ML lifecycle (development/deployment, DevOps/MLOps)
  • 5+ years of Python experience in development & support on AI/ML workflows and data engineering pipelines
  • Practically skilled in Kubernetes environments including Cloud-provider managed Kubernetes flavors (AWS-EKS/GCP-GKE)
  • Knowledge on microservice architectures and service communication patterns
  • Strong troubleshooting fundamentals such as application crashes, resource contentions, service latency, and scaling behavior
  • Well-rounded competency in analyzing logs, metrics, monitoring systems, and service KPIs
Plus points if you have:
  • Exposure in other Data/AI platforms such as Flyte, HuggingFace & AI Agent Platforms (Vertex AI, Claude Code, LangChain, etc...)
  • Hands-on experience with automation or scripting (Bash, Python)
  • Kubernetes or cloud certifications (CKAD, AWS)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Platform Reliability Engineer
Senior Data Platform Reliability Engineer

OpsWerks • Philippines

On-site
PHP 1,200,000 - 2,400,000
Senior Data Platform Reliability Engineer
Senior Data Platform Reliability Engineer

Hammerjack Pty Ltd • Philippines

On-site
PHP 1,200,000 - 2,400,000
Senior Manager
Senior Manager

V2 Solutions • Hinoba-an

On-site
PHP 2,000,000 - 4,200,000
Senior Data Platform Reliability Engineer
Senior Data Platform Reliability Engineer

OpsWerks • Mandaluyong

On-site
PHP 1,200,000 - 2,400,000
Senior Data Platform Reliability Engineer (Kubernetes)
Senior Data Platform Reliability Engineer (Kubernetes)

Hammerjack Pty Ltd • Philippines

On-site
PHP 1,800,000 - 2,400,000
AI Engineer
AI Engineer

OpsWerks • Mandaluyong

On-site
PHP 1,200,000 - 1,800,000
AI Engineer
AI Engineer

Hammerjack Pty Ltd • Philippines

On-site
PHP 800,000 - 1,100,000
Senior Engineer, AI & Data Platform
Senior Engineer, AI & Data Platform

Poet Technologies Pte Ltd • Santo Niño 1st

On-site
PHP 7,542,000 - 11,314,000
Site Reliability Application Engineers (AI Platform)
Site Reliability Application Engineers (AI Platform)

Astek • Santo Niño 1st

On-site
PHP 600,000 - 1,000,000
Devops Engineer
Devops Engineer

Acquire Intelligence • Metro Manila

On-site
PHP 1,200,000 - 1,800,000