Distributed AI Infra Engineer (Kubernetes & Ray)

anyscale

United States

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anyscale is hiring a Software Engineer for the Infrastructure team to help build the control and data planes for scalable distributed AI applications. You will work on Kubernetes-based orchestration, cloud-native infrastructure, and Ray integration to power our cloud platform.

You will contribute to high-performance, secure systems and collaborate with a team of experts to push the boundaries of AI infrastructure, with on-call support and cross-functional engagement with customers.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
  • 3+ years of experience writing high-quality production code.
  • Hands-on experience building highly available, scalable distributed systems.
  • Expertise in cloud-native technologies (AWS, Azure, GCP) and Kubernetes-based deployments.
  • Deep understanding of networking, security, and authentication in cloud environments.
  • Familiarity with observability stacks (Prometheus, Grafana, etc.).
  • Proficiency in Go and Python.
  • Knowledge of Linux, containers, and low-level OS foundations.

Responsibilities

  • Design, build, and scale services that orchestrate Ray clusters across cloud and on-prem environments.
  • Optimize control plane components for large-scale distributed AI/ML workloads.
  • Develop intelligent scheduling and resource management systems for heterogeneous compute clusters.
  • Improve reliability, performance, and observability of Anyscale-managed Ray workloads.
  • Collaborate with teams to integrate Ray with the infinite laptop product and customer environments.
  • Provide on-call support and troubleshoot infrastructure issues with customer teams.

Skills

Go language
Python
Kubernetes
Distributed systems

Education

Bachelor's degree in CS or related

Tools

Prometheus
Grafana

Job description

Anyscale is hiring a Software Engineer for the Infrastructure team to help build the control and data planes for scalable distributed AI applications. You will work on Kubernetes-based orchestration, cloud-native infrastructure, and Ray integration to power our cloud platform.

You will contribute to high-performance, secure systems and collaborate with a team of experts to push the boundaries of AI infrastructure, with on-call support and cross-functional engagement with customers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed AI Infra Engineer (Go/Python, Kubernetes)
Distributed AI Infra Engineer (Go/Python, Kubernetes)

Anyscale • Palo Alto (CA), San Francisco (CA)

On-site
USD 150,000 - 210,000
Staff Platform Engineer, Distributed Infra & Kubernetes
Staff Platform Engineer, Distributed Infra & Kubernetes

Anyscale, Inc. • San Francisco (CA)

On-site
USD 180,000 - 280,000
Senior Site Reliability Engineer — AI Cloud Infra
Senior Site Reliability Engineer — AI Cloud Infra

Anyscale • San Francisco (CA)

On-site
USD 130,000 - 180,000
Software Engineer, Infrastructure
Software Engineer, Infrastructure

anyscale • United States

On-site
USD 140,000 - 210,000
Software Engineer, Infrastructure
Software Engineer, Infrastructure

Socket.dev • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Forward Deployed Engineer - AI/ML Platforms
Forward Deployed Engineer - AI/ML Platforms

anyscale • San Francisco (CA)

On-site
USD 120,000 - 160,000
Software Engineer, Ray Data & AI Pipelines
Software Engineer, Ray Data & AI Pipelines

Anyscale, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Staff Software Engineer, Platform Infrastructure (Foundations)
Staff Software Engineer, Platform Infrastructure (Foundations)

Anyscale, Inc. • San Francisco (CA)

On-site
USD 180,000 - 280,000
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)

Cerebras • Palo Alto (CA)

On-site
USD 130,000 - 170,000
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)

Anyscale • San Francisco (CA)

On-site
USD 130,000 - 180,000