A leading tech solutions provider in the Philippines is seeking a skilled operations engineer to manage multi-tenant data and AI platforms. You will design efficient workflows to enhance customer experience, lead incident management, and mentor junior engineers. The ideal candidate will have over 3 years of experience with ETL and production pipelines, strong skills in tools like Spark and Airflow, and proficiency in programming languages like Python or Java. Join a dynamic team to optimize performance in the cloud.
Qualifications
3+ years of experience supporting production pipelines/workflows.
Strong experience in Spark, Airflow, Flink, or Jupyter.
Solid knowledge in Python, Java, or Scala for automations and data manipulations.
Responsibilities
Operate multi-tenant data/AI platforms with clear SLAs/SLIs/SLOs.
Lead incident management and customer communications during events.
Design processes to improve customer experience and reduce job failures.
Skills
ETL/ELT support
SQL
Spark
Airflow
Flink
Jupyter
Python
Java
Scala
Tools
Terraform
Helm
Argo CD
Prometheus
Grafana
Datadog
Job description
Responsibilities
Run managed services, not just systems. Operate multi-tenant data/AI platforms (Spark, Airflow, Flink, Jupyter) with clear SLAs/SLIs/SLOs, cost guardrails, and capacity plans across AWS/GCP + Kubernetes.
Be the face of reliability. Lead incidents end-to-end, own customer comms and post-incident reviews (RCA with actions customers can see and feel).
Design for Customer experience. Help Data scientists and customers reduce failed/slow jobs, improve time-to-data, and optimize costs—so customers notice faster pipelines and fewer surprises.
Standardize & scale. Build service runbooks, golden paths, and automation that make onboarding and daily ops predictable across customers.
Automate the toil away. Ship tooling (Bash/Python, GitOps, CI/CD) for backups, DR drills, upgrades, access, and environment bootstrapping.
Make signals meaningful. Instrument platforms with metrics/logs/traces; tune alerting to cut noise and improve detection and response times.
Govern change. Plan and execute upgrades/migrations within change windows; champion safe deploys and rollback strategies.
Partner & mentor. Guide junior engineers; collaborate with customer dev/data teams to unblock delivery and raise the reliability bar.
Participate in on-call. Join a 24x7 rotation with crisp handoffs and playbooks.
Your Qualifications
Hands-on support for ETL/ELT, SQL, and production pipelines/workflows.
Strong experience in at least one of Spark, Airflow, Flink, or Jupyter (plus the ecosystem around it).
Solid working knowledge in at least one (1) language - Python, Java or Scala (Automations, Data Manipulations & Orchestrations)
Real-world AWS or GCP and production environment usage as a User or Administrator.
Incident management, post-incident reviews, change management, and service reporting.
3+ years across the domains above, with depth in at least 1–2 tools per domain.
Plus points if you have:
Certifications: CKA/CKAD, AWS (Associate/Professional), or equivalent.
IaC & DevOps: Terraform, Helm, Argo CD/GitOps, CI/CD for data platforms.
Observability & ITSM: Prometheus/Grafana/Datadog; Jira Service Management/ServiceNow, StatusPage.