Staff Software Engineer, Ray Data

Anyscale

San Francisco, Northern (CA, KY)

On-site

USD 170,000 - 260,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Anyscale is seeking exceptional engineers to design, build, and scale Ray Data, a Python-native data processing engine powering modern AI workloads. You will work on core systems, optimize distributed execution, and collaborate with customers to solve hard scalability challenges.

The role requires 6+ years of production-grade software experience with Python, deep knowledge of distributed systems, and a track record of architectural decision making in data or compute engines.

Qualifications

  • 6+ years of production-grade software or infrastructure experience.
  • Strong Python engineering experience.
  • Experience owning architectural decisions in distributed data or compute engines.

Responsibilities

  • Design, build, and optimize core Ray Data systems for performance and scalability.
  • Work on distributed execution across data pipelines.
  • Develop data loading and processing for training and inference workloads.
  • Address fault tolerance, partitioning, and resource management.

Skills

Python
Distributed systems
Performance optimization
Scheduling

Tools

Ray
Linux
Docker

Job description

About Anyscale:

At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.

With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.

Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.

About Ray Data Team:

Ray Data is a Python-native data processing engine and a one-stop shop for all AI data processing needs. Ray Data provides performant, first-class integration with cutting-edge AI frameworks using both multimodal and structured data.

The Ray Data team develops and maintains Ray Data, building the underlying distributed data processing infrastructure that powers modern AI workloads. We are a team of engineers passionate about solving challenging problems in distributed systems, data processing, and performance at scale. We are looking for exceptional engineers to build, optimize, and scale Ray Data for increasingly complex AI workloads, including multimodal data processing and large-scale batch inference.

Learn more about how Ray Data enables scalable multimodal AI workloads in Architecting Multimodal Data Pipelines That Scale with Ray.

As part of this role, you will:

  • Design, build, and improve the core systems that power Ray Data, with a focus on performance, scalability, and reliability.

  • Design and optimize distributed execution across different stages of data pipelines in heterogeneous environments.

  • Build data loading and processing solutions for production training and inference workloads.

  • Solve challenging problems in distributed execution, scheduling, resource management, data partitioning, fault tolerance, and performance optimization.

  • Make system-level architectural decisions and reason through tradeoffs in areas such as resource allocation, execution models, batch vs. streaming workloads, and consistency and availability.

  • Work with customers and new-age AI-native companies to understand and solve challenges in scaling their AI workloads.

We'd love to hear from you if you have:

  • 6+ years of experience building production-grade software, infrastructure, or developer-facing systems, with strong Python engineering experience.

  • 6+ years of experience personally owning core architectural decisions within a distributed data or compute engine, rather than primarily operating or using a platform someone else designed.

  • Deep experience with distributed systems internals, such as scheduling, fault tolerance, data partitioning, distributed execution, performance optimization, or database and query engine internals.

  • A track record of reasoning through system-level tradeoffs and defending architectural decisions, such as batch vs. streaming, static vs. dynamic resource allocation, or consistency vs. availability.

  • Passion for solving the unsolved problems in large-scale AI infrastructure and building systems that enable the next generation of AI applications.

Why Anyscale
  • We're on a mission to make scalable computing effortless. Ray is the AI Compute Engine at the center of some of the world's most powerful AI platforms

  • Our tech is in production at companies like OpenAI, Uber, Spotify, Instacart, and Cruise

  • We're backed by Andreessen Horowitz, NEA, and Addition, with $250M+ raised to date

  • Recent partnerships with Azure, CoreWeave, and Google Cloud are putting AI-native compute directly into enterprise environments

  • Competitive salary and equity, plus health/dental/vision coverage (many plans up to 99% employer-covered)

  • We offer flexible time off, paid parental leave, and mental health support

Anyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.

Anyscale Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer (Ray Data)
Software Engineer (Ray Data)

Anyscale • San Francisco (CA)

On-site
USD 140,000 - 190,000
Software Engineer (Ray Data)
Software Engineer (Ray Data)

Anyscale, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Software Engineer (Ray Core)
Software Engineer (Ray Core)

Anyscale • San Francisco (CA)

On-site
USD 120,000 - 160,000
Head of Platform Infrastructure (Foundations)
Head of Platform Infrastructure (Foundations)

Cerebras • San Francisco (CA)

On-site
USD 260,000 - 380,000
Machine Learning Engineer, Customer Engineering
Machine Learning Engineer, Customer Engineering

Anyscale • San Francisco (CA)

On-site
USD 180,000 - 240,000
Engineering Manager, Platform Infrastructure (Foundations)
Engineering Manager, Platform Infrastructure (Foundations)

Anyscale • San Francisco (CA)

On-site
USD 260,000 - 340,000
Engineering Manager, Platform Infrastructure (Foundations)
Engineering Manager, Platform Infrastructure (Foundations)

Cerebras • San Francisco (CA)

On-site
USD 240,000 - 310,000
Senior / Staff Product Manager (Ray Data)
Senior / Staff Product Manager (Ray Data)

Anyscale • San Francisco (CA)

On-site
USD 120,000 - 160,000
Software Engineer, Ray Data — Scalable AI Pipelines
Software Engineer, Ray Data — Scalable AI Pipelines

Anyscale • San Francisco (CA)

On-site
USD 140,000 - 190,000
Distributed LLM Inference Engineer
Distributed LLM Inference Engineer

Cerebras • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans with 99% premium coverage
401k Retirement Plan
+6