Role Summary
We are looking for an experienced Azure DevOps SRE Engineer to be the senior engineer on a dedicated weekday support engagement for a large enterprise customer. You will work embedded alongside the customer's project engineering team, owning CI/CD pipeline engineering, release support, and the infrastructure and configuration changes needed to move releases from the lowest environment through to production. You will also be the senior technical responder when production issues are escalated past the customer's front-line support.
This is a hands-on senior individual contributor role, not a people-management position. You will set the technical standard on the engagement and act as the day-to-day technical anchor for two junior engineers who cover the later shift reviewing their work, handing over context at shift boundaries, and unblocking them but you will spend most of your time engineering, not managing.
Key Responsibilities
CI/CD Pipeline Engineering & Ownership
- Own the health and evolution of the customer's CI/CD pipelines design new pipelines for new projects, refactor fragile ones, and build reusable templates and patterns.
- Own build-server and agent infrastructure: upkeep, security patching, upgrades, and tooling migrations as the toolchain evolves.
- Manage build/agent cost and capacity — right-sizing, and shutting down or restarting resources as needed.
- Integrate automated testing, quality gates, and rollback paths into pipelines.
Release & Environment Engineering
- Support the customer's full release cadence — daily, weekly, biweekly, monthly, and quarterly — end to end.
- Create and update infrastructure, application configuration, and related changes needed to promote a release from the lowest environment through to production.
- Implement and maintain Infrastructure as Code with Terraform — reusable modules, remote state, and environment-specific configuration across Development, QA, Stage, and Production.
- Manage configuration drift between environments and keep promotion paths predictable and repeatable.
- Assess release readiness and own rollback strategy for the changes you deploy.
Reliability & Escalation-Tier Production Support
- Act as the senior technical responder on production issues escalated beyond the customer's front-line support — working alongside their engineering team and, where needed, Microsoft support.
- Acknowledge and begin analysis within shift hours against severity-based targets: Severity 1 within 30 minutes, Severity 2 within 2 hours, Severity 3 by next business day.
- Drive structured root-cause analysis using logs, metrics, and traces; deliver written issue analysis notes or RCA inputs within five business days of closure for Severity 1 and Severity 2 issues.
- Push for permanent fixes over repeat workarounds — feed recurring failure patterns back into pipelines, configuration, and automation.
- Improve the reliability of what we own: reduce pipeline flakiness and deployment failure rates, and strengthen monitoring and alerting signal where it falls within engagement scope.
- Operate containerized workloads on AKS — deployments, scaling, rolling updates, ConfigMaps, Secrets, ingress, and cluster troubleshooting.
Platform Security & Hygiene
- Enforce secure configuration and pipeline controls, with secrets managed through Azure Key Vault.
- Apply Azure networking, Entra ID, and RBAC fundamentals correctly in the changes you make.
- Keep runbooks, deployment playbooks, and pipeline documentation current and accurate.
Technical Guidance & Handover
- Review the junior engineers' pipeline and configuration changes before they reach higher environments.
- Run a clean handover at the shift boundary — open issues, in-flight releases, and context that would otherwise be lost overnight.
- Unblock the junior engineers on technical problems and coach them through escalation judgement — when to dig in and when to raise it.
- Raise tooling and process improvement recommendations to the customer's engineering team as you spot them.
Delivery & Reporting
- Execute tasks assigned by the customer's project engineering team; the engagement targets 90%+ completion of assigned tasks within the committed sprint or cycle.
- Own weekly status reporting and contribute to the monthly delivery summary.
Required Skills & Experience
- 4–5+ years of hands-on Azure cloud and DevOps engineering with production ownership of a non-trivial estate.
- Strong CI/CD pipeline design using YAML — GitHub Actions and/or Azure Pipelines, including multi-stage builds, quality gates, and rollback.
- Strong Terraform skills: reusable modules, remote state, workspaces, and multi-environment deployment promotion.
- Hands-on Kubernetes, AKS preferred — deployments, scaling, rolling updates, ConfigMaps, Secrets, ingress, and troubleshooting under pressure.
- Docker containerization for real applications.
- Strong scripting in PowerShell and/or Bash for deployment automation and operational tooling. Python is an added advantage.
- Git workflow depth — branching strategies, pull requests, and code review as a habit.
- Secrets and configuration management using Azure Key Vault.
- Working knowledge of Azure networking, Entra ID (Azure AD), RBAC, and cloud security fundamentals.
- Strong production incident response and troubleshooting — can lead analysis on a live issue and drive it to a structured root cause.
- Monitoring and log analysis competence — reading metrics, logs, and traces to diagnose rather than guess.
- Comfortable guiding one or two junior engineers day to day — reviewing their work and coaching them — without this becoming a full-time management load.
- Clear written and verbal English — you work directly with the customer's engineering team and author issue analysis notes, status reports, and documentation.
- Willingness and ability to work the 5:00 PM – 2:00 AM IST shift from our Chennai office on a sustained basis, Monday to Friday.
Good to Have
Familiarity with the tooling used on this engagement is an advantage. Equivalent experience on comparable platforms is acceptable — we care that you have operated a stack of this shape, not that you have used these exact products.
- GitHub Actions as the delivery pipeline — or equivalent (Azure Pipelines, GitLab CI, Jenkins).
- Azure Monitor and Application Insights for platform and application telemetry — or equivalent monitoring/APM tooling (Dynatrace, New Relic, Datadog).
- ELK stack (Elasticsearch, Logstash, Kibana) for log aggregation and analysis — or equivalent (Splunk, Grafana Loki, Azure Log Analytics).
- JIRA for task, incident, and RCA tracking — or equivalent (Azure Boards, ServiceNow).
- Confluence for runbooks and documentation — or equivalent.
- Prometheus and Grafana for metrics and dashboarding.
- GitOps practices using Argo CD or FluxCD.
- ARM Templates or Bicep alongside Terraform.
- Azure Policy, tags, and budgets for cost governance and resource hygiene.
- Azure Container Registry (ACR) and Azure API Management (APIM).
- Azure DevOps services (Boards, Repos, Artifacts) beyond Pipelines.
- Prior experience working directly with an overseas customer engineering team or on a shift-based support engagement.
Preferred Certifications
- AZ-400: Designing and Implementing Microsoft DevOps Solutions — preferred for this role.
- AZ-104 (Azure Administrator Associate) or AZ-305 (Solutions Architect Expert) — acceptable alternatives.
- HashiCorp Terraform Associate — good to have.
- CKA (Certified Kubernetes Administrator) — good to have.
What This Role Offers
- Senior technical ownership of a live enterprise Azure estate without the overhead of a people-management track.
- Direct daily working relationship with a large enterprise engineering team — the kind of exposure that compounds.
- A defined 12-month progression from knowledge transfer to full steady-state ownership.
- Weekday-only schedule with weekends and India-observed holidays off.
- On-site collaboration from our Chennai office, with KnackForge's wider DevOps practice around you.