Stand out for this role — generate a tailored resume and cover letter in about a minute.
ATS CONSULTING SERVICES PH INC. is seeking a Senior Data Platform Reliability Engineer to operate, maintain, and continuously improve data platforms on Kubernetes (on‑prem and cloud). You will contribute to DoEKS/AIoEKS deployment frameworks and drive reliability across the stack.
You will deploy releases through GitOps, monitor health with logs and metrics, participate in incident response, mentor junior engineers, and advocate for platform standards, security, and operational excellence.
As a Senior Data Platform Reliability Engineer, you will be responsible for operating, maintaining, and continuously improving the company's data platforms running on Kubernetes (on-premises and/or on AWS/GCP) - similar to the DoEKS (Data on EKS) / AIoEKS (AI on EKS) deployment frameworks.
Deploy new releases and configuration changes through GitOps/DevOps
Monitor platform and service health using logs, metrics, and observability tools
Participate in incident response, root cause analysis, and 24x7 operational rotations
Improve platform observability, operational tooling/automations, self-service capabilities, and reliability practices to reduce recurring issues
Investigate & troubleshoot user concerns by either correlating them to system-related issues, breaking integrations, and/or user-specific errors/misconfigurations up to recommending/executing resolutions
Provide technical mentorship to junior engineers
Advocate for platform standards, security best practices, and operational excellence
3+ years of solid experience supporting production data workloads/platforms (Spark/Airflow/Jupyter)
5+ years of hands-on experience on ETL/ELT pipeline development & data transformations (Python/Java & SQL)
Practical proficiency in Kubernetes environments including Cloud-provider managed Kubernetes flavors (AWS-EKS/GCP-GKE)
Comprehensive knowledge on Linux environments, microservice architectures and service communication patterns
Strong troubleshooting fundamentals such as application crashes, resource contentions, service latency, and scaling behaviour
Well-rounded competency in analysing logs, metrics, monitoring systems, and service KPIs
Exposure to other Data/AI platforms such as Flink, Trino, Druid, and Ray
Hands-on experience with automation or scripting (Bash, Python)
Kubernetes or Data certifications (CKAD, AWS Certified Data Engineer)