Manager, Site Reliability Engineer - Data Platforms

Amgen Inc. (IR)

Hyderabad

On-site

INR 4,000,000 - 7,000,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Amgen in Hyderabad seeks an experienced Manager, Site Reliability Engineer - Data Platforms to lead secure, scalable cloud platforms and guide a pod of engineers.

You will architect and operate production workloads on AWS, implement IaC, observability, and CI/CD pipelines, and drive reliability, security, and cost efficiency.

This hands-on leadership role requires strong collaboration with global teams and a passion for scalable data platforms that support Amgen's patient-focused mission.

Qualifications

  • Hands-on experience designing, implementing and operating production workloads on AWS and data platforms.
  • Proficiency with EKS, Sagemaker, Bedrock, VPC, S3, EC2, KMS, CloudWatch, CloudTrail, Secrets Manager, Lambda, RDS.
  • Experience with Infrastructure as Code using Terraform, CloudFormation, AWS CDK, or equivalent.
  • Programming skills in Python, Go, Java, TypeScript, Bash, or PowerShell.
  • Strong knowledge of SRE concepts including SLIs, SLOs, incident response, post-incident reviews, and toil reduction.
  • Experience with Docker and Kubernetes, CI/CD, GitOps, and observability tooling.
  • Ability to mentor engineers and communicate architectural decisions to diverse stakeholders.
  • Experience with cost optimization and cloud governance practices.

Responsibilities

  • Define reusable platform services, IaC modules, and guardrails for scalable AWS architectures.
  • Lead on-call and incident response for complex production incidents with blameless reviews.
  • Advance CI/CD, GitOps, and automated testing for safe production changes.
  • Build observable platforms with standardized metrics, logs, and dashboards.
  • Coach engineers, set priorities, and drive cross-functional delivery across global teams.

Skills

AWS
Terraform
Python
Kubernetes
CI/CD
GitOps
Observability
Linux
Networking
SRE principles

Education

Master’s or Bachelor’s degree in computer science or engineering

Tools

Docker
OpenTelemetry
CloudWatch
Prometheus
Grafana
Jira

Job description

Career Category Engineering Job Description

Join Amgen's Mission of Serving Patients At Amgen, if you feel like you're part of something bigger, it's because you are. Our shared mission-to serve patients living with serious illnesses-drives all that we do. Since 1980, we've helped pioneer the world of biotech in our fight against the world's toughest diseases. With our focus on four therapeutic areas -Oncology, Inflammation, General Medicine, and Rare Disease- we reach millions of patients each year. As a member of the Amgen team, you'll help make a lasting impact on the lives of patients as we research, manufacture, and deliver innovative medicines to help people live longer, fuller happier lives. Our award-winning culture is collaborative, innovative, and science based. If you have a passion for challenges and the opportunities that lay within them, you'll thrive as part of the Amgen team. Join us and transform the lives of patients while transforming your career.

Manager, Site Reliability Engineer - Data Platforms

About Amgen Amgen harnesses the best of biology and technology to fight the world’s toughest diseases, and make people’s lives easier, fuller and longer. We discover, develop, manufacture and deliver innovative medicines to help millions of patients. Amgen helped establish the biotechnology industry more than 40 years ago and remains on the cutting-edge of innovation, using technology and human genetic data to push beyond what’s known today. About the Role AWS Site Reliability Engineer with strong cloud architecture expertise to design, build, automate, and operate secure, resilient, scalable, and cost-efficient AWS platforms. You will combine AWS architecture leadership with practical Site Reliability Engineering-writing Infrastructure as Code, developing automation, improving observability, solving complex production problems, and strengthening operational excellence. This role also includes responsibility for leading an SRE pod and managing assigned engineers. However, it is primarily a hands-on technical role. The successful candidate will remain a key technical contributor who leads through architecture, implementation, production ownership, incident response, coaching, and example.

Roles & Responsibilities

AWS Architecture and Platform Engineering Design secure, highly available, scalable, and cost-efficient AWS architectures for EDSE applications and data platforms. Define reusable platform services, Infrastructure as Code modules, reference architectures, and engineering guardrails. Make and document architectural decisions across networking, identity, compute, containers, storage, security, observability, and resilience. Lead architecture and production-readiness reviews, addressing reliability, security, performance, operability, and cost risks. Reliability and Production Operations Establish and improve SRE practices, including SLIs, SLOs, error budgets, availability targets, capacity planning, and operational-readiness criteria. Use operational data to identify recurring failures, performance bottlenecks, capacity risks, and opportunities to reduce manual toil. Participate in the on-call rotation and lead the technical response to complex or high-severity production incidents. Conduct blameless post-incident reviews and ensure corrective actions produce durable engineering improvements. Automation, Delivery, and Observability Build reusable, tested infrastructure and operational automation using Terraform and programming languages such as Python. Automate provisioning, configuration, validation, deployment, recovery, compliance checks, and routine operational activities. Strengthen CI/CD and GitOps practices through automated testing, deployment controls, progressive delivery, and reliable rollback mechanisms. Standardize metrics, logs, traces, dashboards, and actionable alerting to improve issue detection, diagnosis, and recovery. Resilience, Security, and Cost Efficiency Design and validate backup, high-availability, and disaster-recovery strategies aligned with defined business and recovery objectives. Conduct restore tests, failover exercises, resilience reviews, and controlled game-day scenarios. Embed least-privilege access, encryption, network segmentation, secrets management, vulnerability remediation, and policy-based controls into platform designs. Improve AWS cost efficiency through right-sizing, resource lifecycle management, tagging, usage analysis, and architecture optimization without compromising reliability or security. Pod and People Leadership Lead the SRE pod by setting technical direction, prioritizing work, managing operational commitments, and driving delivery to completion. Serve as a senior AWS and SRE advisor, partnering with application engineering, data engineering, cybersecurity, architecture, product, and FinOps teams. Manage and mentor assigned engineers through regular feedback, one-on-one discussions, technical coaching, design and code reviews, and career-development support.

Basic Qualifications and Experience

Master’s or Bachelor’s degree in computer science or engineering field and 9 to 12 years of relevant experience, including substantial hands-on experience in AWS cloud engineering, platform engineering, infrastructure engineering, DevOps, or Site Reliability Engineering. Prior experience leading a technical pod, engineering squad, or small team while continuing to contribute hands-on. Must-Have Skills Significant hands-on experience designing, implementing, and operating business-critical production workloads on AWS, including experience with EKS, Sagemaker, Bedrock, VPC, PrivateLink, S3, EC2, KMS, CloudWatch, CloudTrail, Secrets Manager, Lambda, and RDS. Deep knowledge of AWS architecture across networking, identity and access management, security, compute, containers, storage, monitoring, and resilience. Strong experience with Infrastructure as Code using Terraform, CloudFormation, AWS CDK, or comparable technologies. Proficiency in at least one programming or scripting language, such as Python, Go, Java, TypeScript, Bash, or PowerShell. Ability to write maintainable, production-quality infrastructure code, automation, and operational tooling. Practical experience applying SRE principles, including SLIs and SLOs, actionable alerting, incident response, root-cause analysis, post-incident improvement, and toil reduction. Experience building or operating containerized platforms using Docker and Kubernetes, preferably Amazon EKS. Experience with CI/CD, GitOps, source control, automated infrastructure testing, deployment controls, and safe production-change practices. Experience implementing observability using AWS CloudWatch, OpenTelemetry, Prometheus, Grafana, Splunk, Datadog, or comparable platforms. Strong knowledge of Linux, networking, distributed systems, performance analysis, high-availability design, and disaster recovery. Demonstrated ability to troubleshoot complex production systems and lead technical decision-making during high-severity incidents. Experience of FinOps practices, cloud cost allocation, unit-cost modeling, Savings Plans, and capacity optimization. Experience leading an engineering pod, squad, or technical workstream and driving cross-functional delivery. Experience managing or mentoring engineers, providing constructive feedback, and supporting their technical and career development. Strong written and verbal communication skills, including the ability to explain technical risks and architecture trade-offs to engineering and non-engineering stakeholders. Ability to collaborate effective with global teams, make risk-based priority decisions, and drive complex work to completion. Experience working in Agile delivery environments and using planning and delivery tools such as Jira or Jira Align. Good-to-Have Skills Experience supporting enterprise data platforms or data-intensive workloads on AWS. Experience with Databricks, Graph platforms, Lakehouse platforms, data lakes, or large-scale analytics environments. Candidates without prior Databricks experience should be willing to develop this capability after joining. Experience with AWS Organizations, Control Tower, multi-account landing zones, service control policies, and cloud-governance frameworks. Experience creating internal developer platforms, self-service infrastructure, golden paths, or paved-road capabilities. Experience with policy as code, automated compliance, or security controls in regulated or highly governed environments. Experience with chaos engineering, automated recovery, resilience testing, or large-scale performance testing. Exposure to Azure, hybrid-cloud, or multi-cloud environments. Relevant certifications are preferred but not required, including: AWS Certified Solutions Architect - Professional AWS Certified DevOps Engineer - Professional SAFe Agilist or another SAFe certification Functional Skills Excellent written and verbal communication, with the ability to explain complex platform concepts, architecture decisions, risks, and trade-offs in clear, business-relevant language. Strong influencing and consensus-building skills, including the ability to establish standards and drive adoption across teams without relying solely on formal authority. A platform-product mindset focused on reusable capabilities, paved roads, developer experience, measurable outcomes, and long-term platform health. Strong systems-thinking and structured problem-solving skills, with the ability to diagnose issues across application, platform, cloud governance, security, and operating-model boundaries. High degree of ownership, initiative, and follow-through, with the ability to move ambiguous topics from exploration through decision, implementation, adoption, and continuous improvement. Collaborative and globally minded, with experience working effectively across architecture, governance, cybersecurity, operations, engineering, product, and business teams. Strong planning, estimation, prioritization, and execution skills, with the ability to manage multiple initiatives while maintaining high standards for security, reliability, quality, and reusability. Ability to balance innovation with enterprise risk, distinguishing between experimentation, limited preview adoption, and production-ready capabilities. Strong coaching and enablement skills, with the ability to create clear documentation, facilitate technical workshops, mentor engineers, and build an active platform community. A growth mindset and commitment to continuous learning, modern engineering practices, constructive challenge, and responsible adoption of AI. Increase the use of reusable, tested automation for infrastructure and operational changes. Build a high-performing pod through clear priorities, strong engineering practices, effective coaching, and shared operational ownership.

What you can expect of us

As we work to develop treatments that take care of others, we also work to care for your professional and personal growth and well-being. From our competitive benefits to our collaborative culture, we’ll support your journey every step of the way. In addition to the base salary, Amgen offers competitive and comprehensive Total Rewards Plans that are aligned with local industry standards.

Equal Opportunity Statement

Amgen is an Equal Opportunity employer and will consider all qualified applicants for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, disability status, or any other basis protected by applicable law. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.

GCF Level GCF Level 05 Career Category Information Systems Position Type Full time . Amgen is committed to unlocking the potential of biology for patients suffering from serious illnesses by discovering, developing, manufacturing and delivering innovative human therapeutics. This approach begins by using tools like advanced human genetics to unravel the complexities of disease and understand the fundamentals of human biology. Amgen focuses on areas of high unmet medical need and leverages its biologics manufacturing expertise to strive for solutions that improve health outcomes and dramatically improve people's lives. A biotechnology pioneer since 1980, Amgen has grown to be one of the world's leading independent biotechnology companies, has reached millions of patients around the world and is developing a pipeline of medicines with breakaway potential. For more information, visit www.amgen.com and follow us on www.twitter.com/amgen

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer – Data Platforms
Site Reliability Engineer – Data Platforms

Amgen Inc. (IR) • Hyderabad

On-site
INR 1,500,000 - 2,100,000
Manager, Site Reliability Engineer - Data Platforms
Manager, Site Reliability Engineer - Data Platforms

Amgen SA • Hyderabad

On-site
INR 3,000,000 - 4,500,000
Site Reliability Engineer – Data Platforms
Site Reliability Engineer – Data Platforms

Amgen SA • Hyderabad

On-site
INR 3,500,000 - 7,000,000
Senior Data Engineer
Senior Data Engineer

Amgen Inc. (IR) • Hyderabad

On-site
INR 3,500,000 - 6,000,000
Sr Associate Software Engineer Cloud Platform DevOps
Sr Associate Software Engineer Cloud Platform DevOps

Amgen • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Senior Data Engineer
Senior Data Engineer

Amgen • Hyderabad

On-site
INR 1,200,000 - 2,000,000
Competitive benefits
Collaborative culture
Professional development opportunities
Specialist IS Engineer
Specialist IS Engineer

Amgen Inc. (IR) • Hyderabad

On-site
INR 3,000,000 - 5,000,000
Specialist Software Engineer (Full Stack)
Specialist Software Engineer (Full Stack)

Amgen Inc. (IR) • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Business Systems Analyst
Business Systems Analyst

Amgen Technology Private Limited • Hyderabad

On-site
INR 3,500,000 - 6,000,000
Total rewards program
Associate Software Engineer - Full Stack
Associate Software Engineer - Full Stack

Amgen Inc. (IR) • Hyderabad

On-site
INR 1,200,000 - 1,600,000