Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Private Advertiser is seeking a Senior AI Platform Engineer to operate, maintain, and improve production AI platforms running on Kubernetes across on-premises and cloud environments, including AWS and GCP.
You will be responsible for platform reliability, deployment automation, observability, incident response, and developer experience while driving MLOps best practices and mentoring junior engineers.
We are looking for a Senior AI Platform Engineer to operate, maintain, and continuously improve production AI platforms running on Kubernetes across on-premises and cloud environments, including AWS and GCP.
You will be responsible for platform reliability, deployment automation, observability, incident response, and developer experience while driving MLOps best practices and mentoring junior engineers.
The role involves working with AI platform technologies and frameworks such as Ray.io, Kubeflow, and AWS SageMaker.
Deploy new releases and configuration changes through GitOps and DevOps practices.
Monitor platform and service health using logs, metrics, and observability tools.
Improve platform observability, operational tooling, automation, self-service capabilities, and reliability practices to reduce recurring issues.
Participate in incident response, root cause analysis, and 24x7 operational rotations.
Investigate and troubleshoot user concerns by identifying system-related issues, broken integrations, or user-specific errors and misconfigurations.
Recommend and execute resolutions to platform and integration issues.
Advocate for platform standards, security best practices, and operational excellence.
Provide structured Python mentorship to junior engineers, strengthening their fundamentals and bridging foundational Python knowledge toward MLOps competencies.
Lead the adoption of MLOps best practices across the team.
Influence the team roadmap by identifying gaps in tooling, skills, and processes required to support production-grade AI systems.
3+ years of experience supporting production workloads or platforms such as Ray.io, Jupyter Notebooks, AWS SageMaker, Kubeflow AI tools, or equivalent AI platforms.
5+ years of hands-on experience across the AI/ML lifecycle, including development, deployment, DevOps, and MLOps.
5+ years of Python development and support experience involving AI/ML workflows and data engineering pipelines.
Practical experience with Kubernetes environments, including cloud-managed Kubernetes services such as AWS EKS and GCP GKE.
Knowledge of microservice architectures and service communication patterns.
Strong troubleshooting skills, including application crashes, resource contention, service latency, and scaling behavior.
Competency in analyzing logs, metrics, monitoring systems, and service KPIs.
Exposure to additional data and AI platforms such as Flyte, Hugging Face, and AI agent platforms, including Vertex AI, Claude Code, and LangChain.
Hands-on experience with automation and scripting using Bash and Python.
Kubernetes or cloud certifications, such as CKAD or AWS certifications.
Onsite Work Arrangement
All roles are onsite, with automatic work-from-home arrangements when shifts fall on weekends or holidays.
Shifting Schedules
Available shift schedules: 6:00 AM, 10:00 AM, or 2:00 PM.
HMO on Day 1
Health maintenance organization coverage starts on your first day.
Transportation Allowance
Transportation allowances are provided.