Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Get past ATS filters
Benefits offered by this job
Competitive compensation
Great benefits package
Endless opportunities for career advancement
Hybrid working opportunities
Job summary
A leading global marketing technology company is seeking a System and Platform Operations Director to provide technical leadership for the support and stability of production systems. The role focuses on ensuring a premium customer support experience and requires managing the Platform Operations Team while collaborating closely with engineering and product teams. Ideal candidates should have strong expertise in site reliability engineering, incident management, and cloud services, with excellent communication skills.
Qualifications
5+ years of hands-on experience in Site Reliability focused positions.
Strong knowledge of containerization technologies like Docker and Kubernetes.
Expertise in ITSM principles and Service Delivery strategies.
Responsibilities
Manage the Platform Operation Team for production systems reliability and stability.
Implement proactive solutions for incident detection, response, and remediation.
Ensure high-level support and system availability through IT Service Management processes.
Skills
Site Reliability Engineering
Containerization (Docker, Kubernetes)
Infrastructure as Code (Terraform)
Scripting Languages (Java, Golang, Python)
Monitoring Tools (DataDog, Prometheus)
Database Management (PostgreSQL, Bigtable)
API and Microservices Architecture
ITSM Principles
Cloud Services
Education
Bachelor’s degree or equivalent
Job description
A leading global marketing technology company is seeking a System and Platform Operations Director to provide technical leadership for the support and stability of production systems. The role focuses on ensuring a premium customer support experience and requires managing the Platform Operations Team while collaborating closely with engineering and product teams. Ideal candidates should have strong expertise in site reliability engineering, incident management, and cloud services, with excellent communication skills.