Machine Learning Ops Engineer

resmed

Bengaluru

On-site

INR 450,000 - 900,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

ResMed's Global Technology Solutions (GTS) division is building and operating an AI/ML platform on AWS and Kubernetes. The team focuses on scalable, secure platforms and has strong emphasis on observability and cost controls.

You will own infrastructure, CI/CD, and platform supervision for data pipelines and AI workloads. Join a cross-functional environment with data scientists and GenAI teams to drive production-ready capabilities and safe rollouts.

Qualifications

  • Must have 3+ years engineering in a complex tech environment.
  • Hands-on production Kubernetes experience.
  • Extensive AWS exposure including multiple services (EC2, S3, IAM, VPC, SageMaker, etc.).
  • Proficient with Terraform and CI/CD tooling—GitHub Actions, CodePipeline or Jenkins.
  • Strong Python and SQL skills for data work.
  • Experience with platform observability stacks (Prometheus, Loki, Grafana, Datadog).

Responsibilities

  • Design, build, and operate the AI/ML platform on AWS + Kubernetes - clusters, networking, IAM, storage, cost, and reliability.
  • Provision and evolve infrastructure with Terraform; treat infra as code with real review and rollback.
  • Own CI/CD for data pipelines, ML models, and AI applications - from repo to production with confidence.
  • Stand up and evolve the platform observability stack - Prometheus, Loki, Grafana / Datadog - for metrics, logs, traces, dashboards, alerting, and SLOs.
  • Automate environment provisioning and self-serve tooling for data scientists.
  • Partner with product and GenAI teams to optimize workloads on the platform.
  • Run POCs to pull promising tech into the platform; participate in code review and process improvement.

Skills

Kubernetes
AWS
Terraform
Observability tooling
Python
SQL
CI/CD pipelines
GitHub Actions

Tools

Terraform
Prometheus
Loki
Grafana
Datadog
OpenTelemetry GenAI

Job description

Global Technology Solutions (GTS) at ResMed is a division dedicated to creating innovative, scalable, and secure platforms and services for patients, providers, and people across ResMed. The primary goal of GTS is to accelerate well-being and growth by transforming the core, enabling patient, people, and partner outcomes, and building future-ready operations.

The strategy of GTS focuses on aligning goals and promoting collaboration across all organizational areas. This includes fostering shared ownership, developing flexible platforms that can easily scale to meet global demands, and implementing global standards for key processes to ensure efficiency and consistency.

About the role

ResMed's AI platform powers dozens of data scientists and a growing set of GenAI / Agentic AI products that touch patients, clinicians, and providers worldwide. We run on AWS and Kubernetes , provisioned with Terraform , and shipped through modern CI/CD.

What you'll do
  • Design, build, and operate the AI/ML platform on AWS + Kubernetes - clusters, networking, IAM, storage, cost, and reliability.
  • Provision and evolve infrastructure with Terraform ; treat infra as code with real review and rollback.
  • Own CI/CD for data pipelines, ML models, and AI applications - from repo to production with confidence.
  • Stand up and evolve the platform observability stack - Prometheus, Loki, Grafana / Datadog - for metrics, logs, traces, dashboards, alerting, and SLOs.
  • Automate what shouldn't be manual: environment provisioning, golden-path pipelines, self-serve tooling for data scientists.
  • Partner with product, data science, and GenAI teams to make their workloads first‑class on the platform - model serving, evaluation, cost/latency controls, and safe rollout.
  • Run POCs to pull promising tech into the platform without accumulating debt.
  • Participate in code review, mentoring, and process improvement; raise the engineering bar.
What we're looking for
Must-have
  • 3 + years of engineering experience in a complex, technical environment.
  • Deep, hands‑on Kubernetes in production.
  • Hands-on AWS - comfortable with 3+ of: EKS, Lambda, EC2, S3, IAM, Networking (VPC, ALB/NLB), RDS, EMR, Glue, Athena, Batch, SageMaker, MWAA/Airflow.
  • Working command of Terraform - modules, state, reviews, drift.
  • Platform observability experience: Prometheus, Loki, Grafana and/or Datadog - metrics, logs, dashboards, alerting, SLOs.
  • Strong production Python (and SQL for data work).
  • Experience building CI/CD pipelines and APIs end-to-end - GitHub / GitHub Actions, CodePipeline or Jenkins.
  • Hands‑on working experience with an AI/ML platform in production - data science tooling, model lifecycle, feature / inference infrastructure, and self‑serve enablement for DS and GenAI teams.
  • Deploying AI agents / LLM workloads on Kubernetes - containerizing agent workloads, autoscaling (HPA/KEDA), GPU scheduling where needed, secure egress for tool calls, secrets and rate‑limit management, and running long‑lived / stateful sessions safely.
  • Exposure to the modern AI / Agentic AI stack is required - working familiarity with at least a few of: an agent framework ( LangChain / LangGraph / CrewAI / AutoGen / Strands / Semantic Kernel / PydanticAI ), LLM serving ( vLLM , KServe , Ray Serve, TGI), a RAG / vector-store setup (OpenSearch, pgvector , Pinecone, Weaviate ), LLM observability ( Langfuse , LangSmith , Arize Phoenix, OpenTelemetry GenAI), and MCP (Model Context Protocol) for tool integration.
Nice-to-have - AI / Agentic AI skills
  • AI / Agent frameworks: LangChain , LangGraph , Strands, or similar.
  • Running AI agents on Kubernetes: containerizing agent workloads, autoscaling (HPA/KEDA), stateful sessions, long-running tasks/jobs, secure egress for tool calls, secrets and rate-limit management.
  • Managed agent platforms: AWS Bedrock AgentCore , Bedrock Agents / Knowledge Bases, SageMake r.
  • MCP (Model Context Protocol): authoring or hosting MCP servers/clients, exposing internal tools/data safely to agents.
  • LLM/agent observability: Langfuse , LangSmith , Arize or OpenTelemetry GenAI - traces, evaluations, token / cost / latency tracking.
  • LLM serving on Kubernetes: vLLM , KServe , Ray Serve, TGI; GPU node pools and scheduling.
  • RAG stack: vector stores (OpenSearch, pgvector , Pinecono e ), embeddings pipelines, retrieval evaluation.
  • Guardrails & safety: Bedrock Guardrails, prompt-injection defenses, PII redaction.
  • ML platform tooling: Kubeflow, MLflow , or comparable.
  • Snowflake and modern
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Ops Engineer
Machine Learning Ops Engineer

Worklane GmbH • Bengaluru

On-site
INR 3,500,000 - 7,500,000
AWS DevOps & AI Engineer
AWS DevOps & AI Engineer

AstraZeneca • India

On-site
INR 1,500,000 - 2,300,000
AI Engineer
AI Engineer

Andpayments • India

On-site
INR 1,800,000 - 3,000,000
Data Scientist
Data Scientist

Aligned Automation • Maharashtra

On-site
INR 400,000 - 700,000
AI Architect
AI Architect

MathCo • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Machine Learning Ops. Engineer
Machine Learning Ops. Engineer

Navlakha Management Services • Mumbai

On-site
INR 2,400,000 - 4,200,000
Agentic AI Product developer
Agentic AI Product developer

Tredence • Bengaluru

On-site
INR 3,000,000 - 5,500,000
Principal Architect AI Data Engineer
Principal Architect AI Data Engineer

EXL • Maharashtra

On-site
INR 3,500,000 - 5,200,000
Backend AI Engineer
Backend AI Engineer

Cirruslabs • Hyderabad, Bengaluru

Hybrid
INR 4,000,000 - 8,000,000
Principal Architect AI Data Engineer
Principal Architect AI Data Engineer

EXL • Gurugram District

On-site
INR 4,000,000 - 9,000,000