Site Reliability Engineer (SRE)

TTEC Digital

Hyderabad

On-site

INR 3,500,000 - 7,000,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

TTEC Digital is seeking a seasoned Site Reliability Engineer to own reliability for a real-time platform across voice, desktop, intelligence and AI capabilities. You will define SLOs, manage incident response, and drive scalable production systems and observability with a bias for uptime.

In a startup-like environment with weekly deploys and 1-week sprints, you will champion canary analyses, automatic rollbacks, and AI-assisted reliability.

Qualifications

  • 8+ years operating production systems at scale.
  • Strong Go or Python; automate reliability and minimize toil.
  • Deep knowledge of event-driven and real-time systems and associated failure modes.

Responsibilities

  • Own SLOs and error budgets per tenant/service.
  • Lead incident response and blameless postmortems.
  • Scale production capacity and manage observability depth.

Skills

SRE
Go
Python
Observability
Incident response
Chaos engineering
GCP
Multi-cloud
NATS
WebSocket
Load testing
Automation

Tools

Canary analysis
Runbooks
Telemetry tooling

Job description

At TTEC Digital, we coach clients to ensure their employees feel valued, and fully supported, because an amazing customer experience is an employee first process. Our vision is the same, a place where employees know they can thrive.

The role:
  • Own production reliability for a real-time platform where uptime and latency ARE the product — voice, desktop, intelligence, and AI combined; an agent mid-call can't wait for a retry.
  • First SRE hired immediately (Day 0–14) for production scaling and SLO ownership; a second joins at the start of Phase 3 for 24/7 coverage.
  • Pairs with C1 Platform Foundation on observability and tenancy isolation.
  • Startup environment: weekly deploys, 1-week sprints, fail fast, move forward — reliability engineering at that speed, not against it.
What you'll own:
  • SLOs and error budgets per tenant/service
  • Incident response and blameless postmortems
  • Production scaling and capacity
  • Observability depth (p50/p95/p99 per event hop)
  • Uptime as a personal mission
  • On-call rotation with DevOps
  • Your committed timelines.
Who you are:
  • Self-starter, grit, show-me mentality — you prove reliability with dashboards and drills, not assertions.
  • A ways-to-YES engineer: weekly deploys are the heartbeat and your job is making them safe, never slowing them.
  • You love new technology, adapt fast when the stack changes under you, use AI tools daily to multiply velocity, and consider yourself exceptional.
  • Calm in an incident, relentless after it.
  • Team player who likes winning.
  • 8+ years operating production systems at scale; owns SLOs, error budgets, incident command.
  • Strong Go or Python — you automate reliability, you don't toil at it. Everything you build is code: runbooks execute, remediation is automatic, toil trends to zero.
  • Deep on event-driven and real-time systems reliability — NATS-class buses, WebSocket fleets, streaming pipelines — and the failure physics underneath: state, race conditions, locking, ordering, back-pressure, cascading load. You've debugged these in production.
  • Strong monitoring and uptime mindset — metrics, logs, traces wired to alerting that catches it before the customer does; you know the difference between a noisy alert and a real signal.
  • Good networking understanding — protocols and how they work (TCP/UDP, TLS, WebSocket, DNS, load balancing); RTP/SIP a strong plus for our media paths.
  • GCP at scale; multi-cloud literacy a plus. Multi-tenancy isolation experience a strong plus.
  • Capacity modeling and load testing partnership with QA — find the knee of the curve before customers do.
  • Chaos engineering — failure injection as routine practice; prove graceful degradation, don't assume it.
  • Deploy-safety partnership with DevOps — canary analysis, automatic rollback triggers, error-budget-driven release gates.
  • AI-aware reliability — monitoring model latency, drift, and cost as production signals, not just CPU and memory.
  • Incident communication craft — clear, fast, blameless; execs and customers get truth at the right altitude.
  • A master debugger of production — reads the trace, the metric, the flame graph, and sees it; narrows an incident to the service, the deploy, the event.
What You Will Bring:
  • 8+yearsoperating production systems at scale; owns SLOs, error budgets, incident command.
  • Strong Go or Python — you automatereliability,youdon'ttoil at it. Everything you build is code: runbooksexecute,remediation is automatic, toil trends to zero.
  • Deep on event-driven and real-time systems reliability — NATS-class buses, WebSocket fleets, streaming pipelines — and thefailurephysics underneath: state, race conditions, locking, ordering, back-pressure, cascading load.You'vedebuggedtheseinproduction.
  • Strong monitoring and uptime mindset — metrics, logs, traces wired to alerting that catches it before the customer does; you know the difference between a noisy alert and a real signal.
  • Good networking understanding — protocols and how they work (TCP/UDP, TLS, WebSocket, DNS, load balancing); RTP/SIP a strong plus for our media paths.
  • GCP at scale; multi-cloud literacya plus. Multi-tenancy isolationexperiencea strongplus.
  • Capacity modeling and load testing partnership with QA — find the knee of the curve before customers do.
  • Chaos engineering — failure injection as routine practice;provegraceful degradation,don'tassume it.
  • Deploy-safety partnership with DevOps — canary analysis, automatic rollback triggers, error-budget-driven release gates.
  • AI-aware reliability — monitoring model latency, drift, and cost as production signals, not just CPU and memory.
  • Incident communication craft — clear, fast, blameless; execs and customers get truth at the right altitude.
  • A master debugger of production — reads the trace, the metric, the flame graph, and sees it; narrows an incident to the service, the deploy, the event.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer/ Expert
Lead Site Reliability Engineer/ Expert

SITA Group • Delhi

On-site
INR 1,200,000 - 2,400,000
Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC Acceleration Centers • Hyderabad

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineering Lead (Application SRE Lead)
Site Reliability Engineering Lead (Application SRE Lead)

Hirexa Solutions • Bengaluru

Hybrid
INR 3,500,000 - 7,000,000
Site Reliability Engineer(SRE)
Site Reliability Engineer(SRE)

Techdome • Hyderabad

On-site
INR 150,000 - 210,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Infosys • Hyderabad

On-site
INR 1,400,000 - 2,200,000
SRE Architect
SRE Architect

Prodapt • Chennai District

On-site
INR 6,000,000 - 9,000,000
Site Reliability Engineering Lead_Truist
Site Reliability Engineering Lead_Truist

Infosys • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Infinx • Bengaluru

On-site
INR 2,500,000 - 5,000,000
SRE Engineer @ Investment Banking | Mumbai
SRE Engineer @ Investment Banking | Mumbai

Net Connect Global • Bengaluru, Mumbai

Hybrid
INR 1,800,000 - 2,400,000