Company : EPAM
Description
Senior Engineering Manager for a new site reliability and platform automation team being built in India for a global SaaS customer. The platform runs on EKS and RDS across four production realms. The team is built from zero: nine engineers who replace manual operational work with production-quality software and move operational ownership back to the product teams that build the services.
This is a hands‑on engineering manager role. You hire and build the team, set its charter, and stay close enough to the code to review design and hold an engineering bar with senior engineers. Hiring and delivery run in parallel from the start.
Dedicated team built for a client, with a defined path for the team to transition to the client's own organisation over the course of the engagement. Candidates must be told this at first contact.
WHAT YOU WILL DO
- Recruit, onboard, coach and develop a team of nine engineers
- Own the team charter, roadmap, operating model and delivery priorities across reliability measurement, toil automation, resilience testing and self‑service enablement
- Establish the reliability practice: SLIs, SLOs, error budgets, burn‑rate alerting, observability, incident response, change‑safety governance and capacity tuning
- Reduce recurring operational work by automating CVE remediation, EKS and RDS upgrades, third‑party provider testing and routine infrastructure workflows; set a measured baseline first and report reduction against it
- Build golden paths, Terraform modules and policy‑as‑code guardrails so product teams provision and operate services themselves
- Shift ownership of deployments, operational readiness and SLOs toward the service owners
- Run intake and prioritisation across several engineering teams, communicate trade‑offs, and report progress against committed dates
- Introduce AI‑assisted engineering and operational tooling with explicit guardrails for validating output
- Manage across US and India time zones
MUST HAVE
- 15+ years in software, infrastructure, platform, DevOps or SRE engineering, including 3+ years managing engineering teams
- Has built a team from zero: hired, onboarded and delivered at the same time
- Hands‑on engineering foundation. Can review code, evaluate architecture and make technical calls
- Has established or operated a reliability practice: SLIs, SLOs, error budgets, observability, incident response, capacity planning or disaster recovery
- Has reduced operational toil with a measured baseline and reported outcomes, not anecdotes
- Strong AWS or GCP, Kubernetes, Terraform, CI/CD, and automation in Python, Go or comparable
- Has built internal platforms or self‑service capability that product teams actually adopted
- Practical use of AI‑assisted engineering tools, and the judgement to validate their output
- Bachelor's degree in computer science, computer engineering, IT or a related technical field; master's a plus
NICE TO HAVE
- Kubernetes and cloud platforms at multi‑tenant, multi‑region or organisational scale
- Both AWS and GCP; regulated environments such as FedRAMP, SOC 2 or ISO
- Prometheus, Grafana, OpenTelemetry, ELK or Datadog
- Policy as code: Kyverno, OPA or Gatekeeper
- Jenkins, GitHub Actions, GitOps or Argo
- LLM‑based or agentic tooling applied to operational workflows
- Incident command on high‑severity production incidents
- Managing distributed teams across US and India time zones