Senior Data Platform Reliability Engineer

Scientific Games Technologies

Bengaluru

On-site

INR 500,000 - 900,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Scientific Games is seeking an experienced Senior Data Platform Reliability Engineer to ensure the reliability, resilience, observability, performance, and operational excellence of our enterprise data platform built on Databricks and AWS.

This is a highly hands-on role partnering with engineering leaders to design, automate, monitor, troubleshoot, and continuously improve production platform services while leading reliability initiatives across the Data Platform.

Qualifications

  • 8+ years in Site Reliability Engineering, Platform Engineering, Cloud Operations, or DevOps.
  • 4+ years operating Databricks platforms in production.
  • Strong AWS operational experience including IAM, S3, CloudWatch, CloudTrail, VPC, KMS, Secrets Manager, and networking.
  • Strong Python and scripting experience; active incident management.

Responsibilities

  • Engineer highly available and resilient platform services and improve SLOs/SLAs.
  • Develop platform health dashboards, monitoring, logging, and alerting.
  • Lead Data Platform incident response and recovery procedures.
  • Optimize Databricks clusters, Spark workloads, and cloud cost.
  • Operate Databricks workspaces, jobs, and Unity Catalog; manage AWS services.
  • Drive FinOps initiatives with autoscaling and cost dashboards.
  • Design multi-region disaster recovery and business continuity.
  • Provide technical leadership and mentor engineers on reliability.

Skills

SRE/Platform Engineering
Cloud Operations
Python scripting
Spark & Databricks
Monitoring & Observability

Education

Bachelor's in CS/Engineering

Tools

Terraform
IaC (Infrastructure as Code)
AWS (multi-region)
Docker/Kubernetes

Job description

Job Summary

Scientific Games is seeking an experienced Senior Data Platform Reliability Engineer to ensure the reliability, resilience, observability, performance, and operational excellence of our enterprise data platform built on Databricks and AWS.

Working closely with the Principal Analytics Engineer, Principal Enterprise Data Architect, and engineering teams, this role is responsible for engineering a highly available, secure, and resilient platform that supports business-critical data services and analytics. The role drives operational excellence through automation, proactive monitoring, performance optimization, incident engineering, and continuous improvement.

This is a highly hands-on engineering role. The successful candidate is expected to design, automate, monitor, troubleshoot, and continuously improve production platform services while leading reliability engineering initiatives across the Data Platform.

Success requires deep expertise in Databricks, AWS, cloud operations, observability, disaster recovery, and platform reliability engineering, along with the ability to influence engineering practices through technical leadership.

Mission

Engineer operational excellence by ensuring the Databricks and AWS Data Platform is reliable, observable, resilient, secure, performant, scalable, and cost-efficient.

Scope

Owns the operational reliability of the enterprise Data Platform built on Databricks and AWS.

Success is measured by:
  • Platform availability and SLO/SLA compliance
    MTTR, MTBF, RTO and RPO achievement
    Platform health and reliability
    Operational automation
    Platform cost efficiency
    Reduced incidents and improved developer experience
Job Duties / Key Accountabilities
Reliability Engineering
  • Engineer highly available and resilient platform services.
    Define and improve SLOs, SLAs, and error budgets.
    Eliminate single points of failure.
    Continuously improve resilience through automation.
Observability Engineering
  • Engineer monitoring, logging, metrics, traces, dashboards, and alerting.
    Develop platform health dashboards and proactive alerting.
    Improve diagnostics and operational visibility.
Incident Engineering
  • Lead Data Platform incident response.
    Develop operational runbooks and recovery procedures.
    Perform root cause analysis and automate recovery where practical.
Performance & Capacity Engineering
  • Optimize Databricks clusters, Spark workloads, SQL Warehouses, storage, and compute.
    Engineer capacity planning and growth forecasting.
    Continuously optimize performance and cloud cost.
Databricks & AWS Platform Operations
  • Operate Databricks for workspaces, clusters, jobs, workflows, SQL Warehouses, and Unity Catalog operational components.
    Operate AWS services including IAM, S3, CloudWatch, CloudTrail, VPC, KMS, Secrets Manager, networking, and storage.
    Engineer automation that reduces manual operational effort.
FinOps
  • Optimize cloud consumption and cluster utilization.
    Engineer autoscaling, cluster policies, and cost optimization.
    Build operational cost dashboards and reporting.
Disaster Recovery & Business Continuity
  • Design and maintain multi-region disaster recovery capabilities.
    Engineer backup, replication, failover, and recovery automation.
    Define and validate RTO and RPO objectives.
    Conduct regular disaster recovery exercises and maintain recovery runbooks.
Technical Leadership
  • Lead reliability engineering reviews.
    Mentor engineers on SRE and operational excellence.
    Partner with the Principal Data Engineer and Principal Enterprise Data Architect.
    Promote automation and continuous improvement.
Qualifications / Skills / Knowledge
Required
  • Bachelor's degree in Computer Science, Engineering, Information Systems, or related discipline.
    8+ years in Site Reliability Engineering, Platform Engineering, Cloud Operations, or DevOps.
    4+ years operating Databricks platforms in production.
    Strong AWS operational experience including IAM, S3, CloudWatch, CloudTrail, VPC, KMS, Secrets Manager, and networking.
    Strong Python and scripting experience.
    Strong Spark and Databricks operational knowledge.
    Experience with monitoring and observability platforms.
    Experience with Terraform and Infrastructure as Code.
    Experience operating highly available multi-region AWS platforms.
    Experience with disaster recovery, business continuity, backup, failover, and operational automation.
    Strong troubleshooting and incident management skills.
Desired
  • Databricks Certification.
    AWS Certified DevOps Engineer or Solutions Architect.
    Kubernetes experience.
    Experience in regulated industries.
Authority / Decision Making
Authority To
  • Define operational engineering standards.
    Recommend reliability improvements.
    Define monitoring, alerting, and automation standards.
    Drive platform operational excellence.
Requires Approval For
  • Architecture changes outside approved standards.
    Technology adoption outside the approved platform strategy.
Key Contacts
  • Head of AI, Data & Infrastructure
    Principal Analytics Engineer
    Principal Enterprise Data Architect
    Platform Engineering
    Governance Engineering
    Product Leadership
    Databricks
    AWS
Language Skills

Required: English

Desired: Additional languages considered an asset.

Job Conditions
  • Remote position based in India.
    Regular collaboration across India and North America.
    Flexible schedule with recurring North America overlaps.
    Occasional international travel (up to 10%).
    Participation in operational reviews, architecture reviews, disaster recovery testing, incident response, sprint planning, and technical planning.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Databricks Infrastructure Engineer
Senior Databricks Infrastructure Engineer

Scientific Games Technologies • Bengaluru

Hybrid
INR 3,500,000 - 5,500,000
Databricks Platform Engineer
Databricks Platform Engineer

Scientific Games Technologies • Bengaluru

Hybrid
INR 4,200,000 - 7,200,000
Senior Databricks Governance Engineer
Senior Databricks Governance Engineer

Scientific Games Technoligies • Bengaluru

Hybrid
INR 3,500,000 - 7,000,000
Senior Data Engineer
Senior Data Engineer

Amgen • Hyderabad

On-site
INR 2,500,000 - 4,500,000
AWS Databricks Platform Engineer
AWS Databricks Platform Engineer

Qualcomm • Hyderabad

On-site
INR 2,300,000 - 3,000,000
Lead, Data Engineering
Lead, Data Engineering

S&P Global Market Intelligence • Chennai District, Hyderabad

On-site
INR 4,000,000 - 6,500,000
Databricks Platform Engineer/ Admin
Databricks Platform Engineer/ Admin

Cirruslabs • Pune District, Bengaluru, Hyderabad

Hybrid
INR 2,600,000 - 3,800,000
Senior Databricks Engineer
Senior Databricks Engineer

DataBeat • Hyderabad

On-site
INR 2,500,000 - 4,200,000
Lead, Data Engineering
Lead, Data Engineering

S&P Global, Inc. • Gowlidoddi

On-site
INR 1,500,000 - 2,100,000
Health insurance
Flexible downtime
Continuous learning
+1
Senior Platform Data Engineer (Databricks)
Senior Platform Data Engineer (Databricks)

Anblicks • Hyderabad

On-site
INR 4,500,000 - 7,000,000
Databricks certification
Cloudera certification
AWS certification