Machine Learning Ops Engineer

ResMed Inc

Bengaluru

On-site

INR 3,500,000 - 5,200,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

ResMed Inc is seeking an AI/ML Platform Engineer to design, build, and operate a Kubernetes-based AI/ML platform on AWS. You will own infrastructure, CI/CD for data pipelines and models, and drive observability with Prometheus, Grafana, and related tools.

Collaboration with data scientists and GenAI teams is essential to scale workloads and ensure reliable production systems. Ideal candidates have hands-on Kubernetes in production, Terraform expertise, and strong Python skills, plus experience

Qualifications

  • 3+ years of engineering experience in a complex, technical environment.
  • Hands-on Kubernetes in production.
  • Hands-on AWS with multiple services (EKS, IAM, S3, etc.).
  • Terraform experience with modules, state, reviews, drift.
  • Proficiency in Python and SQL for data work.
  • Experience building CI/CD pipelines and APIs end-to-end.

Responsibilities

  • Design, build, and operate the AI/ML platform on AWS + Kubernetes.
  • Provision infrastructure with Terraform and manage it as code.
  • Own CI/CD for data pipelines, ML models, and AI apps.
  • Stand up and evolve the platform observability stack (Prometheus, Loki, Grafana).
  • Automate environment provisioning and self-serve tooling for data scientists.
  • Collaborate with product, data science, and GenAI teams for workload readiness.
  • Run POCs to evaluate new tech and raise engineering bar.

Skills

Kubernetes
AWS
Terraform
Python
CI/CD pipelines
Platform observability

Tools

Prometheus
Grafana
Loki
Datadog
Terraform (state/review)

Job description

Global Technology Solutions (GTS) at ResMed is a division dedicated to creating innovative, scalable, and secure platforms and services for patients, providers, and people across ResMed. The primary goal of GTS is to accelerate well‑being and growth by transforming the core, enabling patient, people, and partner outcomes, and building future‑ready operations. The strategy of GTS focuses on aligning goals and promoting collaboration across all organizational areas. This includes fostering shared ownership, developing flexible platforms that can easily scale to meet global demands, and implementing global standards for key processes to ensure efficiency and consistency.

About the role ResMed’s AI platform powers dozens of data scientists and a growing set of GenAI / Agentic AI products that touch patients, clinicians, and providers worldwide. We run on AWS and Kubernetes, provisioned with Terraform, and shipped through modern CI/CD. We are looking for AI/ML Platform Engineer whose core is Kubernetes, AWS, Terraform, AI and platform observability — someone who can design, build, and operate the platform end‑to‑end and instrument it so nothing is a mystery in production. You should also bring an AI working mindset: curious about how ML and agentic workloads run on the platform, comfortable partnering with data scientists and GenAI teams, and eager to grow the platform toward LLMOps and Agentic AI as those workloads scale.

What you’ll do
  • Design, build, and operate the AI/ML platform on AWS + Kubernetes — clusters, networking, IAM, storage, cost, and reliability.
  • Provision and evolve infrastructure with Terraform; treat infra as code with real review and rollback.
  • Own CI/CD for data pipelines, ML models, and AI applications — from repo to production with confidence.
  • Stand up and evolve the platform observability stack — Prometheus, Loki, Grafana / Datadog — for metrics, logs, traces, dashboards, alerting, and SLOs.
  • Automate what shouldn’t be manual: environment provisioning, golden‑path pipelines, self‑serve tooling for data scientists.
  • Partner with product, data science, and GenAI teams to make their workloads first‑class on the platform — model serving, evaluation, cost/latency controls, and safe rollout.
  • Run POCs to pull promising tech into the platform without accumulating debt.
  • Participate in code review, mentoring, and process improvement; raise the engineering bar.
What we’re looking for
  • Must‑have 3+ years of engineering experience in a complex, technical environment.
  • Deep, hands‑on Kubernetes in production.
  • Hands‑on AWS — comfortable with 3+ of: EKS, Lambda, EC2, S3, IAM, Networking (VPC, ALB/NLB), RDS, EMR, Glue, Athena, Batch, SageMaker, MWAA/Airflow.
  • Working command of Terraform — modules, state, reviews, drift.
  • Platform observability experience: Prometheus, Loki, Grafana and/or Datadog — metrics, logs, dashboards, alerting, SLOs.
  • Strong production Python (and SQL for data work).
  • Experience building CI/CD pipelines and APIs end‑to‑end — GitHub / GitHub Actions, CodePipeline or Jenkins.
  • Hands‑on working experience with an AI/ML platform in production — data science tooling, model‑lifecycle, feature / inference infrastructure, and self‑serve enablement for DS and GenAI teams.
  • Deploying AI agents / LLM workloads on Kubernetes — containerizing agent workloads, autoscaling (HPA/KEDA), GPU scheduling where needed, secure egress for tool calls, secrets and rate‑limit management, and running long‑lived / stateful sessions safely.
  • Exposure to the modern AI / Agentic AI stack is required — working familiarity with at least a few of: an agent framework (LangChain / LangGraph / CrewAI / AutoGen / Strands / Semantic Kernel / PydanticAI), LLM serving (vLLM, KServe, Ray Serve, TGI), a RAG / vector‑store setup (OpenSearch, pgvector, Pinecone, Weaviate), LLM observability (Langfuse, LangSmith, Arize Phoenix, OpenTelemetry GenAI), and MCP (Model Context Protocol) for tool integration.
  • Nice‑to‑have: AI / Agentic AI skills AI / Agent frameworks: LangChain, LangGraph, Strands, or similar. Running AI agents on Kubernetes: containerizing agent workloads, autoscaling (HPA/KEDA), stateful sessions, long‑running tasks/jobs, secure egress for tool calls, secrets and rate‑limit management. Managed agent platforms: AWS Bedrock AgentCore, Bedrock Agents / Knowledge Bases, SageMaker. MCP (Model Context Protocol): authoring or hosting MCP servers/clients, exposing internal tools/data safely to agents. LLM/agent observability: Langfuse, LangSmith, Arize or OpenTelemetry GenAI — traces, evaluations, token / cost / latency tracking. LLM serving on Kubernetes: vLLM, KServe, Ray Serve, TGI; GPU node pools and scheduling. RAG stack: vector stores (OpenSearch, pgvector, Pinecone), embeddings pipelines, retrieval evaluation. Guardrails & safety: Bedrock Guardrails, prompt‑injection defenses, PII redaction. ML platform tooling: Kubeflow, MLflow, or comparable. Snowflake and modern data stack experience.
Why join
  • A supportive, senior team with real problems and real users.
  • Freedom to design and influence.
  • Global collaboration and open exchange of ideas.
  • And the chance to build a platform whose output shows up — directly — in better sleep, better breathing, and better health for millions of people.

Resmed (NYSE:RMD, ASX: RMD) creates life‑changing health technologies that people love. We’re relentlessly committed to pioneering innovative technology to empower millions of people in more than 140 countries to live happier, healthier lives. Our AI‑powered digital health solutions, cloud‑connected devices and intelligent software make home healthcare more personalized, accessible and effective. Ultimately, Resmed envisions a world where every person can achieve their full potential through better sleep and breathing, with care delivered in their own home.

Learn more about how we’re redefining sleep health at Resmed.com and follow @Resmed.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Ops Engineer
Machine Learning Ops Engineer

Worklane GmbH • Bengaluru

On-site
INR 3,500,000 - 7,500,000
Machine Learning Ops Engineer
Machine Learning Ops Engineer

resmed • Bengaluru

On-site
INR 450,000 - 900,000
Senior Engineer, AIML Platform Architect
Senior Engineer, AIML Platform Architect

Boston Scientific • Gurugram District

On-site
INR 4,000,000 - 7,500,000
Senior Specialist, Platform Analytics AI ( Python and AWS)
Senior Specialist, Platform Analytics AI ( Python and AWS)

IN10 (FCRS = IN010) Novartis Healthcare Private Limited • Hyderabad

On-site
INR 4,500,000 - 6,500,000
MLR Content Library Specialist
MLR Content Library Specialist

ResMed Inc • Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Senior Decision Scientist
Senior Decision Scientist

Astrazeneca India Private Limited • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior AI and Cloud Solution Architect
Senior AI and Cloud Solution Architect

GE Healthcare Private Limited • Bengaluru

On-site
INR 4,500,000 - 7,500,000
Account Manager
Account Manager

ResMed Inc • Mumbai

On-site
INR 1,500,000 - 2,600,000
AWS DevOps And AI Engineer
AWS DevOps And AI Engineer

AstraZeneca • Chennai District

On-site
INR 1,800,000 - 3,000,000
AI/MLOps SRE Lead Engineer
AI/MLOps SRE Lead Engineer

Regeneron India Private Limited • Hyderabad

Hybrid
INR 4,000,000 - 7,500,000
Hybrid work model
Comprehensive benefits