Senior AI Platform Engineer

Opswerks

Manila

On-site

PHP 1,200,000 - 1,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Opswerks is seeking a Senior AI Platform Engineer to operate and improve AI platforms on Kubernetes, enhancing service health and operational processes.

The ideal candidate will have substantial experience in AI/ML lifecycles, strong Python skills, and a commitment to mentoring junior engineers in MLOps best practices. Candidates should have at least 5 years of experience with Kubernetes environments and have a strong analytical mindset.

Qualifications

  • 3+ years of experience supporting production workloads/platforms.
  • 5+ years of hands-on experience in AI/ML lifecycle.
  • 5+ years of Python experience in development & support.

Responsibilities

  • Operate and maintain the company's AI platforms running on Kubernetes.
  • Monitor platform health using logs, metrics, and observability tools.
  • Improve operational tooling and reliability practices.

Skills

Kubernetes environments
Python
AI/ML lifecycle
DevOps/MLOps
Troubleshooting fundamentals

Tools

Ray.IO
Jupyter Notebooks
AWS SageMaker
Kubeflow AI Tools

Job description

Senior AI Platform Engineer

OpsWerks is a technical consulting company specializing in operational services for the high-tech industry. We help platform and infrastructure teams operate multi‑cloud environments, execute complex migrations, and enable seamless app deployments.

Your Role

As a Senior AI Platform Engineer, you will be responsible for operating, maintaining, and continuously improving the company’s AI platforms running on Kubernetes (On‑premise and/or on AWS/GCP) – similar to the AIoEKS (AI on EKS) deployment frameworks and Kubeflow’s Machine Learning Toolkit.

Platform Ownership & Operations
  • Deploy new releases and configuration changes through GitOps/DevOps.
  • Monitor platform and service health using logs, metrics, and observability tools.
  • Improve platform observability, operational tooling/automations, self‑service capabilities and reliability practices to reduce recurring issues.
  • Participate in incident response, root cause analysis and 24x7 operational rotations.
User & Developer Experience
  • Investigate & troubleshoot user concerns by correlating them to system‑related issues, breaking integrations and/or user‑specific errors/misconfigurations, and recommending/executing resolutions.
  • Advocate for platform standards, security best practices, and operational excellence.
Collaboration and Leadership
  • Provide structured Python mentorship to junior engineers, focusing on strong fundamentals and bridging foundational Python knowledge toward MLOps competencies.
  • Lead the adoption of MLOps best practices for the team.
  • Influence the team roadmap by identifying gaps in tooling, skills, and processes required to support production‑grade AI systems.
Your Qualifications
  • 3+ years of experience supporting production workloads/platforms (Ray.IO, Jupyter Notebooks, AWS SageMaker, Kubeflow AI Tools or an AI‑related equivalent).
  • 5+ years of hands‑on experience in AI/ML lifecycle (development/deployment, DevOps/MLOps).
  • 5+ years of Python experience in development & support on AI/ML workflows and data engineering pipelines.
  • Practically skilled in Kubernetes environments including cloud‑provider managed Kubernetes flavors (AWS‑EKS/GCP‑GKE).
  • Knowledge of microservice architectures and service communication patterns.
  • Strong troubleshooting fundamentals such as application crashes, resource contentions, service latency, and scaling behavior.
  • Well‑rounded competency in analyzing logs, metrics, monitoring systems, and service KPIs.
Plus points if you have:
  • Exposure in other Data/AI platforms such as Flyte, HuggingFace & AI Agent Platforms (Vertex AI, Claude Code, LangChain, etc.).
  • Hands‑on experience with automation or scripting (Bash, Python).
  • Kubernetes or cloud certifications (CKAD, AWS).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Platform Engineer: MLOps Lead on Kubernetes
Senior AI Platform Engineer: MLOps Lead on Kubernetes

Opswerks • Manila

On-site
PHP 1,200,000 - 1,800,000
AI Engineer
AI Engineer

OpsWerks • Mandaluyong

On-site
PHP 1,200,000 - 1,800,000
AI Engineer
AI Engineer

Opswerks • Philippines

Hybrid
PHP 900,000 - 1,200,000
Senior Data Platform Reliability Engineer
Senior Data Platform Reliability Engineer

TerraBarn Inc • Mandaluyong

On-site
PHP 1,200,000 - 2,400,000
Principal AI Engineer
Principal AI Engineer

Cotiviti • Mexico

Hybrid
Senior Software Engineer - AI Platform
Senior Software Engineer - AI Platform

EtonHouse International Education Group • Manila

On-site
PHP 80,000 - 120,000
Senior Data Platform Reliability Engineer | Onsite
Senior Data Platform Reliability Engineer | Onsite

TerraBarn Inc • Cebu City

On-site
PHP 1,200,000 - 2,400,000
Senior AI Engineer
Senior AI Engineer

Adventure Consultancy Solutions Philippines Inc. • Metro Manila

On-site
PHP 2,200,000 - 4,000,000
SR. Data Platform Reliability Engineer/Data SRE (Permanent)- Onsite
SR. Data Platform Reliability Engineer/Data SRE (Permanent)- Onsite

ATS CONSULTING SERVICES PH INC. • Mandaluyong

On-site
PHP 1,000,000 - 2,000,000
AI Engineer
AI Engineer

OpsWerks • Manila

On-site
PHP 600,000 - 1,000,000