Principal Scientist - Data Pipeline Engineer

Adobe

San Jose (CA)

On-site

USD 268,000 - 388,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Adobe is seeking a Principal ML Engineer to architect and scale the multimodal data processing pipelines and infrastructure behind Firefly’s multimodal foundation models (image, video, audio). You will build distributed, GPU‑accelerated systems turning billions of raw assets into training‑ready data at scale, influencing throughput, reliability, and data quality for model training.

This senior individual contributor role requires deep expertise in data engineering, ML infrastructure, and

Qualifications

  • 10+ years in data engineering, ML infrastructure, or distributed systems.
  • Experience with large-scale data processing frameworks.
  • Proficient in Python and systems languages (C++, Rust, Go, or Java).
  • Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, or ML.

Responsibilities

  • Architect and optimize large‑scale distributed data pipelines processing billions of assets.
  • Scale up inference throughput across pipelines to improve training data speed and cost.
  • Design and manage data infrastructure for storage, indexing, and retrieval at scale.
  • Collaborate with modeling teams to align data curation with training goals.

Skills

Python
Distributed systems
Data pipelines
Ray
Spark
C++
Go
Java
Data processing
ML inference

Education

Computer Science degree

Tools

Ray
Spark

Job description

About The Role

We’re looking for a Principal ML Engineer to architect and scale the multimodal data processing pipelines and infrastructure behind Adobe Firefly’s multimodal foundation models (image, video, audio). In this role, you’ll sit at the intersection of data engineering and applied ML building distributed, GPU‑accelerated systems that turn billions of raw assets into training‑ready data at scale.

Your work will directly determine how fast and how well Adobe models can learn, directly impacted by the throughput and reliability of our data pipelines, and the quality of data that reaches training. This is a senior individual contributor role with broad technical influence across data, infrastructure, and modeling teams.

What You’ll Do
Optimize Data Processing Pipelines at Scale
  • Architect and optimize large‑scale distributed pipelines that process billions of images, video, and audio assets through ML workflows into training‑ready data
  • Scale up inference throughput across the pipeline (batching, parallelism, hardware utilization) to turn raw collected data into training data faster and more cheaply
  • Identify and eliminate bottlenecks across ingestion, processing, and delivery, from storage and I/O to compute scheduling
Architect Scalable Data Infrastructure
  • Design systems that reliably store, index, and serve billions of data points, each requiring substantial processing spanning large‑scale databases, distributed storage, and high‑throughput compute
  • Apply deep expertise in distributed systems and frameworks such as Ray (or equivalent) to orchestrate large‑scale, GPU/CPU‑heavy data workloads
  • Own architecture decisions including database and storage choices, job scheduling, GPU cluster utilization that let the platform scale alongside data and model growth
Drive Data Curation for Model Training
  • Bring a strong ML background, especially inference optimization for VLMs and LLMs and data curation for training
  • Partner closely with modeling teams to understand what data improves training outcomes, and translate that into pipeline and curation requirements
  • Operate as a hands‑on technical leader who bridges data engineering and applied ML
What You Need To Succeed
  • 10+ years of experience in data engineering, ML infrastructure, or distributed systems, including work at large scale (billions of records or assets)
  • Strong software engineering background, with hands‑on expertise in distributed systems and frameworks such as Ray, Spark, or equivalent large‑scale data processing frameworks
  • Proficiency in Python, plus strong experience in a systems‑level language (C++, Rust, Go, or Java) with strong debugging skills across distributed and ML‑centric runtime environments
  • Deep knowledge of databases and storage systems at scale such as data lakes, indexing, and retrieval across billions of data points
  • Strong ML background, particularly expertise in optimizing GPU inference pipelines for VLMs, LLMs, or other large models (batching, quantization, serving, throughput/latency tradeoffs)
  • Experience with data curation for model training: understanding what makes data valuable for training generative or multimodal models, not just how to move it efficiently
  • Comfort operating across the full stack, from low‑level systems and GPU optimization to higher‑level data strategy and curation decisions
  • Ability to communicate clearly and partner effectively across data, infrastructure, and modeling teams
  • Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, Machine Learning, or a related field
Equal Employment Opportunity Statement

Adobe is proud to be an Equal Employment Opportunity employer. We do not discriminate based on gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, veteran status, or any other protected characteristic. Learn more.

Expected Pay Range

The U.S. pay range for this position is $206,300 -- $388,000 annually. Pay within this range varies by work location and may also depend on job‑related knowledge, skills, and experience. In California, the pay range is $268,000 - $388,000. In Washington, it is $247,200 - $357,900.

Fair Chance Ordinances

Adobe will consider qualified applicants with arrest or conviction records for employment in accordance with state and local laws and fair‑chance ordinances.

Colorado Application Window Notice

If this role is open to hiring in Colorado (as listed on the job posting), the application window will remain open until at least the date and time stated above in Pacific Time, in compliance with Colorado pay transparency regulations. If this role does not have Colorado listed as a hiring location, no specific application window applies, and the posting may close at any time based on hiring needs.

Massachusetts Legal Notice

It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Scientist, ML - Overall Architect
Principal Scientist, ML - Overall Architect

Adobe Inc. • San Jose (CA)

On-site
USD 206,000 - 379,000
Senior Machine Learning Engineer, Services/MLOps
Senior Machine Learning Engineer, Services/MLOps

Adobe • San Jose (CA)

On-site
USD 183,000 - 266,000
Comprehensive benefits programs
Career growth opportunities
Senior Applied Scientist
Senior Applied Scientist

Adobe Inc. • San Jose (CA)

On-site
USD 162,000 - 302,000
Senior Applied Scientist / Engineer, Training & Inference
Senior Applied Scientist / Engineer, Training & Inference

Adobe Inc. • San Jose (CA)

On-site
USD 216,000 - 313,000
Senior Machine Learning Engineer, Applied Science Data Frameworks
Senior Machine Learning Engineer, Applied Science Data Frameworks

Adobe Inc. • San Jose (CA)

On-site
USD 152,000 - 265,000
Research Engineer
Research Engineer

Adobe • Seattle (WA)

On-site
USD 146,000 - 212,000
Machine Learning Engineer, Express AI Foundations
Machine Learning Engineer, Express AI Foundations

Adobe Inc. • San Jose (CA)

On-site
USD 161,000 - 235,000
Senior ML Infra Engineer for Foundation Models at Scale
Senior ML Infra Engineer for Foundation Models at Scale

Adobe • San Jose (CA)

On-site
USD 151,800 - 265,350
Senior Machine Learning Engineer, Applied Science Data Frameworks
Senior Machine Learning Engineer, Applied Science Data Frameworks

Adobe • San Jose (CA)

On-site
USD 183,000 - 265,000
Machine Learning Engineer
Machine Learning Engineer

Adobe Inc. • San Francisco (CA)

On-site
USD 163,000 - 237,000