Software Engineer, ML Infrastructure, Content Retrieval Platform, Level 4

Relha LLC

Palo Alto, Northern (CA, KY)

Hybrid

USD 157,000 - 235,000

Full time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Paid parental leave
Comprehensive medical coverage
Emotional and mental health support
Equity / RSU

Job summary

Snap Inc. is seeking a Software Engineer for ML Infrastructure in California to design and operate large-scale retrieval and ML systems. You will work on feature generation, serving pipelines, and high-performance inference, collaborating with ML engineers to deploy models into production.

The role emphasizes scalable data workflows, distributed systems expertise, and leveraging Python and Java for performance-critical components. Office-based with a strong emphasis on privacy and security.

Qualifications

  • Bachelor’s degree in a technical field or equivalent experience.
  • 2+ years of post-Bachelor’s software development experience or 1+ year with a Master’s.
  • Experience building large scale production ML systems and distributed systems.
  • Familiarity with ML frameworks and data processing tools.

Responsibilities

  • Design and optimize infrastructure for ML workloads at scale.
  • Build feature generation and serving pipelines for online inferencing and offline data generation.
  • Develop high-performance inference systems for model serving.
  • Build scalable data collection, labeling, processing, and evaluation pipelines.
  • Collaborate with ML engineers to deploy models into production.
  • Ensure code correctness, security, and production-ready quality.

Skills

Python
Java
Distributed systems
Problem solving
Collaboration
Scaleable ML infrastructure
Model serving

Education

Bachelor's degree in Computer Science or related field
Masters/PhD in technical field (preferred)

Tools

TensorFlow
PyTorch
Caffe2
Spark
Flink
Ray

Job description

Software Engineer, ML Infrastructure, Content Retrieval Platform, Level 4

Snap Inc is a technology company. We believe the camera presents the greatest opportunity to improve the way people live and communicate. Snap contributes to human progress by empowering people to express themselves, live in the moment, learn about the world, and have fun together.

The Company operates Snapchat, a visual messaging app that enhances your relationships with friends, family, and the world, and Specs Inc., a wholly-owned subsidiary dedicated to making computing more human, in addition to Bitmoji, Saturn, and other digital services.

Snap Engineering teams build fun and technically sophisticated products that reach hundreds of millions of Snapchatters around the world, every day. We’re deeply committed to the well-being of everyone in our global community, which is why our values are at the root of everything we do. We move fast, with precision, and always execute with privacy at the forefront.

We're looking for a Software Engineer to join the Content Retrieval Platform team, part of the Content ML organization. We develop and operates on large scale retrieval system for content recommendation, partner with Applied ML teams for architecture changes and model developments. We utilize Java/Go as the main indexing and serving languages with performant pieces in C++. We are building agentic repo starting from retrieval root and incubating L0.5, generative retrieval, and redesign of inference feature sourcing.

What you’ll do:

Design and optimize infrastructure systems for machine learning workloads at scale and drive reliability and efficiency improvements across Snapchat’s ML Infrastructure

Build and enhance feature generation and serving pipelines that power online inferencing and offline training data generation

Develop high-performance inference systems to ensure fast and efficient AI model serving

Build infrastructure to perform scalable ML model training, evaluation, and inference in the cloud

Develop high-performance inference systems to ensure fast and efficient AI model serving

Build comprehensive data management systems for scalable data collection, labeling, processing, and evaluation

Work closely with ML engineers to deploy cutting-edge models into production

Utilize AI tools and high velocity engineering workflows to design and ship scalable services while upholding rigorous standards for code correctness, security, and production ready quality code

Knowledge, Skills & Abilities:
Strong programming skills in Python, Java

Strong problem-solving skills with a focus on system performance, scalability, and efficiency

Good understanding of distributed systems and the infrastructure components of large-scale ML

Experience with big data processing frameworks such as Spark, Flink, or Ray
Ability to collaborate and work well with others
Proven track record of operating highly-available systems at significant scale
Ability to proactively learn new concepts and apply them at work

Adaptability in learning and applying evolving AI systems and tools to remain at the forefront of engineering trends and modern development practices

Minimum Qualifications:

Bachelor’s degree in a technical field such as computer science or equivalent experience

2+ years of post-Bachelor’s software development experience; or Master’s degree in a technical field + 1+ year of post-grad software development experience; or PhD in a relevant technical field

Experience building large scale production machine learning systems, distributed systems or big data processing

Preferred Qualifications:

Masters/PhD in a technical field such as computer science or equivalent industry experience

Experience working with ML Training platforms or optimizing AI model inference

Familiarity with ML frameworks such as TensorFlow, PyTorch, Caffe2, Spark ML, scikit-learn, or related frameworks

If you have a disability or special need that requires accommodation, please don’t be shy and provide us some information.

"Default Together" Policy at Snap: At Snap Inc. we believe that being together in person helps us build our culture faster, reinforce our values, and serve our community, customers and partners better through dynamic collaboration. To reflect this, we practice a “default together” approach and expect our team members to work in an office 4+ days per week.

At Snap, we believe that having a team of diverse backgrounds and voices working together will enable us to create innovative products that improve the way people live and communicate. Snap is proud to be an equal opportunity employer, and committed to providing employment opportunities regardless of race, religious creed, color, national origin, ancestry, physical disability, mental disability, medical condition, genetic information, marital status, sex, gender, gender identity, gender expression, pregnancy, childbirth and breastfeeding, age, sexual orientation, military or veteran status, or any other protected classification, in accordance with applicable federal, state, and local laws. EOE, including disability/vets.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance and the Los Angeles Fair Chance Initiative for Hiring, where applicable).

Our Benefits:

Snap Inc. is its own community, so we’ve got your back! We do our best to make sure you and your loved ones have everything you need to be happy and healthy, on your own terms. Our benefits are built around your needs and include paid parental leave, comprehensive medical coverage, emotional and mental health support programs, and compensation packages that let you share in Snap’s long-term success!

  • paid parental leave
  • comprehensive medical coverage
  • emotional and mental health support programs
  • compensation packages that let you share in Snap’s long-term success!
Compensation

In the United States, work locations are assigned a pay zone which determines the salary range for the position. The successful candidate’s starting pay will be determined based on job‑related skills, experience, qualifications, work location, and market conditions. The starting pay may be negotiable within the salary range for the position. These pay zones may be modified in the future.

Zone A (CA, WA, NYC):

The base salary range for this position is $157,000-$235,000 annually.

Zone B:

The base salary range for this position is $149,000-$223,000 annually.

Zone C:

The base salary range for this position is $133,000-$200,000 annually.

This position is eligible for equity in the form of RSUs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, ML Infrastructure, Content Retrieval Platform, Level 4
Software Engineer, ML Infrastructure, Content Retrieval Platform, Level 4

Socket.dev • Palo Alto (CA)

On-site
USD 133,000 - 235,000
Paid parental leave
Comprehensive medical coverage
Mental health support
Software Engineer, ML Infrastructure, Level 4
Software Engineer, ML Infrastructure, Level 4

Snap Inc. • Los Angeles (CA)

On-site
USD 157,000 - 235,000
Paid parental leave
Comprehensive medical coverage
Mental health support programs
+1
Software Engineer, ML Infrastructure, Level 4
Software Engineer, ML Infrastructure, Level 4

Snap Inc. • Palo Alto (CA)

On-site
USD 133,000 - 235,000
Paid parental leave
Comprehensive medical coverage
Mental health support programs
+1
Machine Learning Engineer, Level 4
Machine Learning Engineer, Level 4

Snap Inc. • New York (NY)

On-site
USD 173,000 - 259,000
Equity (RSUs)
Comprehensive medical coverage
Paid parental leave
Principal Software Engineer, Machine Learning Infrastructure
Principal Software Engineer, Machine Learning Infrastructure

Snap • New York (NY)

On-site
USD 276,000 - 414,000
Equity in RSUs
Machine Learning Engineer, Level 4
Machine Learning Engineer, Level 4

Snap Inc. • Palo Alto (CA)

On-site
USD 173,000 - 259,000
Paid parental leave
Comprehensive medical coverage
Emotional and mental health support
+1
Machine Learning Engineer, Level 3
Machine Learning Engineer, Level 3

Relha LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 118,000 - 176,000
Parental leave
Comprehensive medical coverage
Mental health support
Staff Machine Learning Engineer, Diffusion, Generative Modeling and Inference
Staff Machine Learning Engineer, Diffusion, Generative Modeling and Inference

Snap • New York (NY)

On-site
USD 229,000 - 343,000
RSU equity
Software Engineer, Backend, Level 4
Software Engineer, Backend, Level 4

Snap Inc. • Los Angeles (CA)

On-site
USD 157,000 - 235,000
RSUs
Medical coverage
Paid parental leave
Software Engineer, Backend, Level 5
Software Engineer, Backend, Level 5

Snap Inc. • Los Angeles (CA)

Hybrid
USD 209,000 - 313,000
Equity RSUs
Medical coverage