Software Engineer - Data Infrastructure

Luma

Redwood City (CA)

On-site

USD 170,000 - 360,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Luma seeks a Data Infrastructure Engineer in Research to build and scale data infrastructure supporting multimodal AI systems. You will develop high-throughput data pipelines and collaborate with ML researchers to enable efficient experimentation and robust platform services.

You will focus on distributed systems, performance, and reliability, contributing to open-source data tools and adopting best practices across teams in a fast-paced, product-driven environment.

Qualifications

  • Proficiency in Python and willingness to learn Python for large-scale data infrastructure
  • Experience with distributed computing frameworks (e.g., Ray, Spark, Beam) and building high-throughput data systems
  • Ability to design and optimize data pipelines for ML research and internal teams
  • Strong problem-solving skills and understanding of data engineering at scale
  • Collaborative, product-focused mindset; comfortable in fast-paced environments
  • Experience sourcing, integrating, and optimizing data from diverse and large datasets
  • Comfortable working in a fast-paced, product-focused environment with a strong execution mindset
  • Open to candidates across seniority levels, from mid-level contributors to senior engineers and managers

Responsibilities

  • Build and maintain scalable data infrastructure for high-throughput ML workflows
  • Collaborate with ML researchers and product teams to ensure data systems meet evolving needs
  • Develop and optimize large-scale data pipelines and batch processing jobs
  • Contribute to architecture and implementation of reliable, high-performance data platforms
  • Integrate open-source tools and continuously improve data infrastructure through monitoring and tuning
  • Participate in cross-functional projects to improve data reliability, scalability, and operational excellence
  • Support evaluation and adoption of new programming languages and frameworks
  • Engage in continuous improvement of data infrastructure through monitoring and performance tuning
  • Collaborate with research & engineering teams to refine best practices for data infrastructure development

Skills

Python
Distributed frameworks
Data pipelines
Problem solving
Collaboration
Data sourcing
Fast-paced env
Senior-friendly

Job description

About The Role

As a Data Infrastructure Engineer in Research at Luma, you will play a critical role in building and scaling the data infrastructure that supports our cutting-edge multimodal AI systems. Your work will focus on developing high-throughput, large-scale data processing pipelines tailored for machine learning research and internal ML platform needs. You will collaborate closely with ML researchers and product teams to create reliable, efficient, and easy-to-use data infrastructure that empowers innovation and accelerates development. This role requires a strong foundation in distributed systems and data engineering, with an emphasis on supporting complex machine learning workflows rather than traditional product data infrastructure.

About The Role

As a Data Infrastructure Engineer in Research at Luma, you will play a critical role in building and scaling the data infrastructure that supports our cutting-edge multimodal AI systems. Your work will focus on developing high-throughput, large-scale data processing pipelines tailored for machine learning research and internal ML platform needs. You will collaborate closely with ML researchers and product teams to create reliable, efficient, and easy-to-use data infrastructure that empowers innovation and accelerates development. This role requires a strong foundation in distributed systems and data engineering, with an emphasis on supporting complex machine learning workflows rather than traditional product data infrastructure.

Responsibilities
  • Build and maintain scalable data infrastructure for high-throughput machine learning workflows
  • Collaborate with ML researchers and product teams to ensure data systems meet evolving needs
  • Develop and optimize large-scale data pipelines and batch processing jobs
  • Contribute to the architecture and implementation of reliable, high-performance data platforms
  • Integrate open-source tools and continuously improve data infrastructure through monitoring and tuning
  • Participate in cross-functional projects to improve data reliability, scalability, and operational excellence
  • Support the evaluation and adoption of new programming languages and frameworks relevant to data infrastructure
  • Engage in continuous improvement of data infrastructure through monitoring, troubleshooting, and performance tuning
  • Collaborate with research & engineering teams to help define and refine best practices for data infrastructure development
Qualifications
  • Proficiency in Python (or similar languages with willingness to learn Python) and experience with large-scale, high-throughput data infrastructure
  • Familiarity with distributed computing frameworks (e.g., Ray, Spark, Beam)
  • Ability to design and optimize data pipelines for ML research and internal teams
  • Strong problem-solving skills and understanding of data engineering at scale
  • Collaborative, product-focused mindset; comfortable in fast-paced environments
  • Experience sourcing, integrating, and optimizing data from diverse and large datasets
  • Comfortable working in a fast-paced, product-focused environment with a strong execution mindset
  • Open to candidates across seniority levels, from mid-level individual contributors to senior engineers and managers.
Nice to have
  • Prior experience working with complex data infrastructure or AI/ML platforms highly desirable
  • Experience with open source data infrastructure projects is a plus
  • Experience working in the robotics industry preferred

Compensation Range: $170K - $360K

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Research Data Infrastructure Engineer
ML Research Data Infrastructure Engineer

Luma • Redwood City (CA)

On-site
USD 170,000 - 360,000
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000
Software Engineer [ Data Pipelines & Interface ]
Software Engineer [ Data Pipelines & Interface ]

Metamorphic • Palo Alto (CA)

On-site
USD 160,000 - 240,000
Visa sponsorship
Competitive compensation
Mentorship and career development
Software Engineer, Data Infrastructure
Software Engineer, Data Infrastructure

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision insurance
Unlimited paid time off (PTO)
Paid parental leave
+1
ML Infra Engineer (Data Systems)
ML Infra Engineer (Data Systems)

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Data Infrastructure
Member of Technical Staff, Data Infrastructure

Inception • San Francisco (CA)

On-site
USD 140,000 - 190,000
Machine Learning Engineer, Distributed Data Systems - Robotics
Machine Learning Engineer, Distributed Data Systems - Robotics

Jobzhr • San Francisco (CA)

Hybrid
USD 295,000 - 445,000
Relocation assistance
Hybrid work model
In-office 3 days/week
Research Member of Technical Staff- Data Infrastructure
Research Member of Technical Staff- Data Infrastructure

Rhoda AI • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Senior Software Engineer, ML Infrastructure
Senior Software Engineer, ML Infrastructure

LMArena • San Francisco (CA)

On-site
USD 170,000 - 260,000
Competitive compensation
Comprehensive health benefits
Opportunity to work on cutting-edge AI
+1
Software Engineer, Data Infrastructure - Research
Software Engineer, Data Infrastructure - Research

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 160,000