Senior Machine Learning Platform Engineer

Amgen Inc. (IR)

Hyderabad

On-site

INR 2,500,000 - 4,500,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Amgen Inc. (IR) in Hyderabad seeks a Senior Machine Learning Platform Engineer to design, build and scale enterprise-grade ML and GenAI platform capabilities.

You will work at the intersection of platform engineering, MLOps and GenAI, delivering shared APIs, automation and patterns for teams to develop, evaluate and deploy AI solutions at scale. You will collaborate with data scientists, ML engineers, DevOps and security to ensure a secure, reliable developer experience and to define platform

Qualifications

  • 3-5+ years in ML platform engineering, MLOps, backend or enterprise AI systems.
  • Strong software fundamentals; production-grade distributed services and APIs.
  • Experience with REST, microservices and backend platform services.
  • Hands-on with Docker and Kubernetes or equivalent.
  • Knowledge of ML platforms like MLflow, Kubeflow, SageMaker or Databricks.

Responsibilities

  • Design and build reusable ML and GenAI platform capabilities for model development, evaluation, deployment and production operations across teams.
  • Build self-service platform services, APIs and automation to abstract infrastructure complexity.
  • Develop and maintain MLOps capabilities including experiment tracking, model registries and deployment workflows.
  • Build model-serving and inference capabilities for ML models and LLMs via REST, gRPC or event-driven interfaces.
  • Develop GenAI capabilities including prompt management, embeddings, vector search, RAG and agent frameworks.
  • Integrate cloud AI/ML SaaS/PaaS and provide reusable enterprise capabilities.
  • Containerize platform services with Docker and Kubernetes and define deployment patterns.
  • Design platform APIs, SDKs, templates and shared libraries to standardize development patterns.
  • Implement observability with logs, metrics, tracing, health and dashboards.
  • Implement AI evaluation, quality gates and regression testing pipelines.
  • Embed security and governance controls: authentication, RBAC, secrets, data access, auditability.
  • Meter usage, monitor costs and optimize AI service consumption.
  • Build scalable, reliable and resilient platform services with retries and fault tolerance.
  • Advance CI/CD for AI platform services and model workloads with tests and artifact management.
  • Evaluate new AI/ML technologies for reuse vs. application-specific use cases.
  • Collaborate with data scientists to surface recurring challenges into reusable platform patterns.

Skills

Python programming
REST APIs design
Docker and Kubernetes
MLOps concepts
GenAI architectures
Cloud platforms AWS/Azure/GCP

Tools

MLflow
Kubeflow
SageMaker
Databricks
LangChain
LangGraph
Semantic Kernel

Job description

Career Category Engineering Job Description

We are seeking a Senior Machine Learning Platform Engineer to design, build and scale enterprise-grade machine-learning and generative-AI platform capabilities. This role sits at the intersection of platform engineering, software engineering, MLOps and GenAI engineering. Rather than primarily developing individual machine-learning models or business-specific AI applications, you will build the shared platform capabilities, APIs, automation and engineering patterns that enable teams to develop, evaluate, deploy, govern and operate AI solutions at scale. You will work closely with data scientists, ML engineers, application teams, DevOps, Security, Compliance and Product teams to create a secure, reliable and frictionless AI developer experience. The role combines hands-on engineering with technical leadership, helping define platform standards, reusable patterns and architecture for enterprise AI systems.

Roles & Responsibilities
  • Design and build reusable ML and GenAI platform capabilities that support model development, experimentation, evaluation, deployment and production operations across multiple teams and use cases.
  • Build self-service platform services, APIs and automation that abstract infrastructure complexity and enable developers to provision and consume AI capabilities consistently.
  • Develop and maintain MLOps capabilities including experiment tracking, model and prompt registries, evaluation frameworks, deployment workflows and automated promotion across environments.
  • Build model-serving and inference capabilities that support classical ML models, deep-learning models and LLMs through scalable REST, gRPC or event-driven interfaces.
  • Develop platform capabilities for GenAI and agentic systems, including model access, prompt management, embeddings, vector search, Retrieval-Augmented Generation, tool integration and agent frameworks.
  • Engineer integrations with major cloud-based AI and data platforms, using APIs, SDKs and managed services to provide reusable enterprise capabilities.
  • Build and maintain containerized platform services using Docker and Kubernetes, including deployment patterns, scaling strategies, service configuration and lifecycle management.
  • Design and implement platform APIs, SDKs, templates and shared libraries that establish standardized development patterns and reduce duplication across engineering teams.
  • Implement comprehensive observability and operational monitoring, including logs, metrics, distributed tracing, service health, model/LLM usage, latency, errors and operational dashboards.
  • Implement AI evaluation and quality-management capabilities, including automated evaluation pipelines, regression testing, model comparison and release-quality gates.
  • Build security and governance controls into platform capabilities, including authentication, authorization, secrets management, data access controls, auditability, lineage and responsible-AI controls.
  • Design platform mechanisms for usage metering, cost visibility and optimization, enabling teams to understand infrastructure and AI-service consumption.
  • Engineer platform services for scalability, reliability and resilience, including retries, asynchronous processing, concurrency controls, fault tolerance and graceful failure handling.
  • Develop and improve CI/CD pipelines for AI platform services, reusable components and model-based workloads, including automated testing, artifact management and environment promotion.
  • Evaluate new AI, ML and cloud technologies and determine when they should be introduced as shared platform capabilities versus application-specific solutions.
  • Partner with data scientists and ML engineers to identify recurring development and operational challenges and convert them into reusable platform patterns and services.
  • Provide technical guidance on architecture, scalability, performance, security and production-readiness for ML and GenAI workloads.
  • Participate in architecture reviews, code reviews, incident resolution and production troubleshooting across application, platform and infrastructure layers.
  • Create and maintain technical designs, architecture decision records, development standards, operational runbooks and platform documentation.
  • Help define the longer-term technical roadmap and engineering standards for ML and GenAI platform capabilities.
Must-Have Skills
  • 3-5 years of experience in machine learning engineering, ML platform engineering, MLOps, backend engineering, cloud engineering or enterprise AI systems.
  • Strong software-engineering fundamentals with experience building production-grade distributed services, APIs and reusable libraries.
  • Strong programming skills in Python; experience with Java or another enterprise programming language is preferred.
  • Experience designing and developing REST APIs, microservices and backend platform services.
  • Hands-on experience with Docker and Kubernetes or equivalent containerization and orchestration technologies.
  • Strong understanding of MLOps and GenAIOps concepts, including experiment tracking, model lifecycle management, evaluation, deployment, monitoring and reproducibility.
  • Experience with platforms and technologies such as MLflow, Kubeflow, SageMaker, Databricks or equivalent ML platforms.
  • Experience building or integrating model-serving infrastructure for ML models and/or LLM-based applications.
  • Strong understanding of modern GenAI architecture patterns, including LLM APIs, prompt management, embeddings, vector databases, RAG and agent-based systems.
  • Experience with GenAI or agent frameworks such as LangChain, LangGraph, Semantic Kernel or equivalent frameworks.
  • Experience integrating cloud-based AI/ML SaaS and PaaS services and building abstractions around those services for broader developer consumption.
  • Experience with at least one major cloud platform such as AWS, Azure or GCP.
  • Familiarity with CI/CD, Git-based development, automated testing and release-management practices.
  • Understanding of observability and distributed-system concepts, including logging, metrics, tracing, retries, timeouts and failure handling.
  • Understanding of enterprise security patterns, including authentication, authorization, RBAC, OAuth/OIDC, secrets management and encryption.
  • Familiarity with data-governance and responsible-AI concepts, including lineage, explainability, access controls, model evaluation and bias monitoring.
  • Experience working with relational and NoSQL databases and structured or unstructured data systems.
  • Ability to design reusable platform abstractions rather than solving requirements through one-off application implementations.
  • Ability to evaluate architectural options and communicate trade-offs across performance, scalability, reliability, security and cost.
  • Strong collaboration and communication skills, with the ability to work across data science, engineering, infrastructure, security and product teams.
Preferred Skills
  • Experience building internal developer platforms, ML platforms or self-service engineering platforms.
  • Experience with Infrastructure as Code such as Terraform.
  • Experience designing multi-tenant platform services with resource isolation, quotas, governance and usage attribution.
  • Experience with feature stores, model registries, evaluation platforms, AI gateways or centralized model-serving architectures.
  • Familiarity with event-driven architectures, message queues and asynchronous processing.
  • Experience defining technical standards, architecture patterns and reusable engineering frameworks across multiple teams.

Amgen is committed to unlocking the potential of biology for patients suffering from serious illnesses by discovering, developing, manufacturing and delivering innovative human therapeutics. This approach begins by using tools like advanced human genetics to unravel the complexities of disease and understand the fundamentals of human biology. Amgen focuses on areas of high unmet medical need and leverages its biologics manufacturing expertise to strive for solutions that improve health outcomes and dramatically improve people's lives. A biotechnology pioneer since 1980, Amgen has grown to be one of the world's leading independent biotechnology companies, has reached millions of patients around the world and is developing a pipeline of medicines with breakaway potential. For more information, visit www.amgen.com and follow us on www.twitter.com/amgen

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Machine Learning Engineer
Principal Machine Learning Engineer

Amgen Inc. (IR) • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Sr. Associate - Machine Learning Engineer
Sr. Associate - Machine Learning Engineer

Amgen Inc. (IR) • Hyderabad

On-site
INR 1,500,000 - 2,200,000
Specialist Software Engineer - AI Systems
Specialist Software Engineer - AI Systems

Amgen Inc. (IR) • Hyderabad

On-site
INR 3,500,000 - 7,000,000
Senior Associate, Agents/Chatbots IC
Senior Associate, Agents/Chatbots IC

Amgen Inc. (IR) • Hyderabad

On-site
INR 1,800,000 - 3,200,000
Machine Learning Engineer
Machine Learning Engineer

Amgen • Hyderabad

On-site
INR 1,500,000 - 2,400,000
Specialist Software Engineer – Full Stack AI/ML/GenAI
Specialist Software Engineer – Full Stack AI/ML/GenAI

Amgen Inc. (IR) • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Principal Software Engineer
Principal Software Engineer

Amgen Inc. (IR) • Hyderabad

On-site
INR 3,500,000 - 7,000,000
Total Rewards Plans
Principal Machine Learning Engineer - Forecasting
Principal Machine Learning Engineer - Forecasting

Amgen Inc. (IR) • Hyderabad

On-site
INR 3,000,000 - 6,000,000
Sr Machine Learning Engineer
Sr Machine Learning Engineer

Amgen • Hyderabad

On-site
INR 2,500,000 - 5,000,000
Sr Associate Data Scientist
Sr Associate Data Scientist

Amgen Inc. (IR) • Hyderabad

On-site
INR 3,000,000 - 6,000,000