Senior Director- Infrastructure, Operations & App Support

Randstad

Hyderabad

On-site

INR 3,500,000 - 5,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Randstad is seeking an experienced senior leader to drive cloud platform operations, SRE maturity, and MSP governance across enterprise environments in Hyderabad. The role requires strong expertise in multi-cloud architectures, observability, and ITSM aligned with ITIL v4, delivering reliable, scalable services.

You'll lead end-to-end IT support, incident management, and vendor relationships while mentoring teams to foster a culture of reliability and continuous improvement.

Qualifications

  • Deep expertise in cloud platform operations across AWS, Azure and GCP.
  • Strong grounding in AI/ML operations (MLOps) and AIOps tooling.
  • SRE mastery including observability stacks and incident management.

Responsibilities

  • Own the operational health, availability, and performance of enterprise cloud platforms and AI/ML infrastructure.
  • Establish MLOps/AIOps standards, model monitoring, drift detection, and automated retraining pipelines.
  • Lead the enterprise SRE function with reliability engineering principles and post-incident reviews.
  • Govern MSP and vendor relationships, contracts, SLAs, and quarterly business reviews.
  • Oversee ITSM processes (Incident, Problem, Change, Release) aligned to ITIL v4 and risk controls.
  • Provide executive visibility, reporting, and mentorship to the team.

Skills

Cloud platform operations (AWS, Azure,

Tools

Datadog
Dynatrace
Prometheus
Grafana
ServiceNow

Job description

Role & responsibilities
Cloud & Data & AI/ML Platform Operations
  • Own the operational health, availability, and performance of enterprise cloud platforms (AWS, Azure, GCP) and AI/ML infrastructure ensuring production-grade reliability for models, pipelines, data & analytics platforms, and cloud-native applications.
  • Establish and enforce MLOps and AIOps operational standards model monitoring, drift detection, automated retraining pipelines, inference infrastructure management, and incident response for AI/ML workloads.
Site Reliability Engineering (SRE)
  • Build, lead, and mature the enterprise SRE function embedding reliability engineering principles (SLOs, SLIs, error budgets, chaos engineering) across critical digital platforms and services.
  • Lead the post-incident review (PIR) and blameless retrospective culture, ensuring every significant incident drives lasting systemic improvements rather than short-term fixes.
Managed Service Provider (MSP) Governance
  • Serve as the executive owner of all MSP and third-party operational vendor relationships governing contracts, SLAs, performance metrics, and strategic alignment across managed infrastructure, cloud, support, and security services.
  • Lead structured QBRs, performance reviews, and executive-level escalations with MSP partners holding providers accountable to contractual commitments while fostering collaborative, long-term partnerships.
Enterprise IT Support & Service Management
  • Oversee the enterprise IT support function Tier 1/2/3 support, service desk operations, and application support ensuring exceptional end-user experience and first-contact resolution metrics.
  • Lead the continuous maturation of ITSM processes (Incident, Problem, Change, Release, and Configuration Management) in alignment with ITIL best practices and enterprise risk controls.
Executive Visibility, Reporting & Stakeholder Engagement
People Leadership, Mentoring & Team Culture
Technical Competencies
  • Deep expertise in cloud platform operations (AWS, Azure, GCP) architecture patterns, operational tooling, FinOps, and multi-cloud governance.
  • Strong grounding in AI/ML operations MLOps pipelines, model monitoring, inference infrastructure, and AIOps platform tooling.
  • SRE mastery - observability stacks (Datadog, Dynatrace, Prometheus/Grafana), chaos engineering, SLO frameworks, and incident management platforms.
  • ITSM fluency - ServiceNow or equivalent, ITIL v4 processes, and enterprise support operations at scale.
  • MSP governance and vendor management - SLA construction, performance metrics, contract lifecycle, and strategic sourcing principles.
  • Security and compliance operations awareness - vulnerability management, cloud security posture, and regulatory compliance frameworks.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform SRE
Platform SRE

YASH Technologies • Bengaluru

On-site
INR 4,000,000 - 7,000,000
DevOps Specialist Engineer SRE, Cloud & Applied AI
DevOps Specialist Engineer SRE, Cloud & Applied AI

Clarus Advisers • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Clarus Advisers • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Site Reliability Engineering Lead (Application SRE Lead)
Site Reliability Engineering Lead (Application SRE Lead)

Hirexa Solutions • Bengaluru

Hybrid
INR 3,500,000 - 7,000,000
Engagement Manager - Support Lead
Engagement Manager - Support Lead

Quantiphi Analytics Solutions • Thiruvananthapuram

Hybrid
INR 2,500,000 - 3,500,000
Site Reliability Engineer
Site Reliability Engineer

Cybage Software • Pune District

On-site
INR 2,500,000 - 4,000,000
Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC Acceleration Centers • Hyderabad

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer
Site Reliability Engineer

Nexcess • India

On-site
INR 2,000,000 - 4,200,000
Production Support Lead
Production Support Lead

Cloudxtreme • Hyderabad

On-site
INR 3,200,000 - 6,000,000
Engineering Manager
Engineering Manager

WaferWire Cloud Technologies • Hyderabad

On-site
INR 4,000,000 - 7,000,000