Who we are:
At Evam, we're reshaping how enterprises engage with their customers, enabling real-time, data-driven interactions at scale. From our HQ in London and offices in Amsterdam,Istanbul and Sofia we help leading Telcos, banks, and global brands connect with over 500 million people every month, delivering the right message at the right moment. Our AI-powered event processing engine and real-time machine learning capabilities turn complex customer data into instant, personalized experiences that drive measurable business outcomes across every channel.
Recognized as a Forbes Türkiye Top 50 Startup, an Endeavor High-Impact Venture, and a Mar-Tech Awards winner, Evam is also proudly a Happy Place to Work.
We're building a platform that runs mission-critical production workloads at scale. If you enjoy owning the technical direction of our Kubernetes platform, building and growing a strong team, and turning incident learnings into systemic fixes, you'll fit here.
This is a platform-ownership and people-management role, not a ticket queue, so we're looking for a talented Lead DevOps Engineer to build, lead, and grow our platform team!
Job requirements
- BSc/MSc in Computer Science
- 7+ years of hands-on experience in DevOps, SRE, Platform Engineering, or Infrastructure Engineering in production environments, with at least 2 years directly managing engineers (performance reviews, career development, hiring)
- Strong Kubernetes and container orchestration experience (cluster lifecycle, networking, storage, performance, troubleshooting) at a level where you can set standards and review others' designs
- Experience operating cloud environments (AWS and/or Azure), ideally multi-cloud and OnPrem environments
- Proficiency in Infrastructure as Code (Terraform, Ansible) and automated platform management
- Experience designing and operating CI/CD pipelines (Jenkins, GitHub Actions, or similar)
- Strong Linux and scripting skills with confidence in distributed systems troubleshooting
- Experience with observability stacks (metrics, logs, traces) and production monitoring practices
- Proven track record of making and owning architectural decisions for infrastructure/platform systems
- Experience building and growing a technical team from hiring and onboarding to performance management and career planning
- Strong communication and stakeholder management skills — able to represent the platform team to other engineering leads and to senior leadership
Nice to Have:
- Experience with event-driven architectures and Kafka at scale
- Observability tooling: Prometheus, Grafana, SigNoz, OpenTelemetry, Mimir, OneUptime
- Database operations (PostgreSQL, MongoDB, Redis, Elasticsearch)
- Experience in fintech, banking, or regulated industries
- GitOps tooling (ArgoCD, Flux) and DevSecOps practices
- Experience supporting AI/ML workloads on Kubernetes (Kubeflow, KServe, model serving, GPU scheduling)
- Familiarity with MLOps lifecycle (model deployment, monitoring, versioning)
- JVM-based containerized applications
- Relevant certifications (CKA, CKS, AWS, Azure)
Job responsibilities
- Set the technical direction and roadmap for Kubernetes platforms across AWS, Azure, and bare-metal environments, supporting 50+ microservices
- Own people management for the DevOps/platform team: hiring, onboarding, 1:1s, performance reviews, and career development
- Coach and grow the team's engineers, reviewing designs and raising the technical bar
- Set team goals and workload priorities, balancing platform roadmap with individual growth areas
- Own and improve CI/CD and GitOps-based delivery pipelines, enabling safe, zero-downtime releases
- Build and evolve observability (metrics, logs, traces) to ensure deep visibility and rapid incident detection across distributed systems
- Implement autoscaling, self healing, and resilience patterns across services and infrastructure
- Collaborate with data and ML teams to operate and scale real-time ML and inference workloads on Kubernetes
- Enhance monitoring and incident detection using AI-assisted analysis and anomaly detection techniques
- Integrate security controls and scanning into pipelines and platform layers (DevSecOps)
- Lead incident response and drive postmortems into systemic reliability improvements, tracking follow-through across teams
- Own platform strategy conversations with engineering leadership, balancing reliability, cost, and delivery speed
- Build self-service platform tooling and documentation that enable developers to ship safely and fast
Our Stack
Orchestration: Kubernetes, Docker, Docker Swarm
Cloud: AWS, Azure
Cloud Services: EKS, ECR, ALB, WAF
CI/CD: Jenkins, Github, Nexus
IaC: Terraform, Ansible
Observability: Prometheus, Grafana, Loki
Messaging: Kafka
Databases: PostgreSQL, MongoDB, Redis, Elasticsearch
Networking: Istio, reverse proxy, TLS, load balancing
Security: Trivy, Grype, Snyk, Fortify
EVAM is committed to diversity, equity, inclusion, and equal opportunity. We welcome applications from individuals of all backgrounds and are dedicated to creating an inclusive workplace where everyone is treated with dignity, fairness, and respect.