Role: Senior SRE
Location: Bengaluru
We’re MiQ, a global programmatic media partner for marketers and agencies. Our people are at the heart of everything we do, so you will be too. No matter the role or the location, we’re all united in the vision to lead the programmatic industry and make it better.
Job Description
As an Senior SRE in our Engineering department, you'll have the chance to own the health of Sigma, MiQ’s enterprise platform, by building world‑class observability across its services and infrastructure. Your responsibilities will include designing and implementing monitoring, alerting, and synthetic checks that surface application errors, performance regressions, and infrastructure issues before our users notice them, working with product engineering teams to drive resolution, and building dashboards and runbooks while strengthening incident response practices and on‑call readiness.
Responsibilities
- Own the health of Sigma, MiQ’s enterprise platform, by building world‑class observability across its services and infrastructure.
- Design and implement monitoring, alerting, and synthetic checks that proactively surface application errors, performance regressions, and infrastructure issues before our users notice them and partner with product engineering teams to drive them to resolution.
- Work hands‑on with our Grafana‑based observability stack (metrics, logs, traces, and synthetic monitoring) and our AWS/Kubernetes (EKS) environment to define meaningful SLIs and SLOs, reduce alert noise, build actionable dashboards and runbooks, and strengthen incident response practices including on‑call readiness.
- Improve release safety and platform stability by designing and optimizing release pipelines, enabling progressive delivery, and implementing health gates, automated deployment analysis, and feature‑flag rollbacks to detect issues early and minimize user impact.
- Bring a performance engineering mindset – load testing, capacity analysis, latency and resource profiling – and automate away toil through scripting and infrastructure‑as‑code.
- Accelerate SRE maturity by applying AIOps capabilities to improve signal quality, speed up detection and diagnosis, and reduce manual operational effort.
- Over time, grow into owning the observability charter for the platform, setting standards and evangelizing best practices across engineering teams.
Qualifications
- 4 – 8 years of experience in Site Reliability Engineering, DevOps, or platform/production engineering roles supporting customer‑facing systems.
- Hands‑on experience with observability tooling – Grafana, Prometheus, Datadog, and log/trace aggregation (e.g., Loki, Tempo, OpenTelemetry) – covering metrics, logs, traces, and events.
- Experience setting up synthetic monitoring (API and browser checks) to validate critical user journeys and catch failures proactively.
- Deep knowledge of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning and noise reduction, and blameless post‑incident reviews.
- Solid experience operating workloads on Kubernetes (ideally EKS) and AWS – comfortable debugging issues across the application, container, and infrastructure layers.
- Strong scripting and automation skills in Python, Bash, or Go, with exposure to infrastructure‑as‑code (e.g., Terraform) and CI/CD pipelines.
- A performance engineering orientation: load/stress testing (e.g., k6, JMeter, Locust), capacity planning, and profiling latency, throughput, and resource bottlenecks.
- Familiarity with leveraging AIOps capabilities to advance SRE maturity and drive innovation.
- Experience contributing to incident management – triage, escalation, communication, and post‑mortems – and helping set up or improve on‑call processes and rotations.
- A proactive, ownership‑driven mindset: identify problems from telemetry before they’re reported and follow through with the teams who need to fix them.
- Clear written and verbal communication – turn noisy signals into crisp findings, runbooks, and recommendations for engineering teams.
- Strong collaboration abilities to work effectively alongside peer Platform teams, such as Cloud, DevOps, and Security.
Stakeholders
Key stakeholders include the DevOps & Cloud team, Sigma product engineering teams, Tech Leads across Engineering, and engineering leadership who rely on reliability and performance insights to make informed decisions.
Impact
You will be the engineer who makes reliability visible and actionable for MiQ’s flagship platform. You’ll build the observability foundations that let teams detect and resolve errors, performance degradations, and infrastructure issues before they impact users, reduce mean time to detection and resolution, cut alert fatigue, and raise the bar on incident response and on‑call readiness. As you grow, you’ll own the observability charter: defining standards, introducing new tooling and practices where appropriate, and evangelizing a proactive reliability culture across engineering.
Benefits
- A hybrid work environment
- New hire orientation with job‑specific onboarding and training
- Internal and global mobility opportunities
- Competitive healthcare benefits
- Bonus and performance incentives
- Generous annual PTO, paid parental leave, and two additional paid days to acknowledge holidays, cultural events, or inclusion initiatives
- Employee resource groups designed to connect people across all MiQ regions, drive action, and support our communities
Values
- We do what we love – Passion
- We figure it out – Determination
- We anticipate the unexpected – Agility
- We always unite – Unite
- We dare to be unconventional – Courage
Apply today!
Equal Opportunity Employer
Equal Opportunity Employer