Principal Scientist - Data Pipeline Engineer

Adobe Inc.

San Jose (CA)

On-site

USD 268,000 - 388,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Adobe Inc. in California seeks a Principal ML Engineer to architect and scale multimodal data pipelines for Firefly’s foundation models. You will build distributed, GPU-accelerated systems turning billions of raw assets into training-ready data at scale.

As a senior IC, you’ll influence data, infrastructure, and modeling teams, drive throughput and data curation, and ensure reliability while enabling rapid model learning progress.

Qualifications

  • 10+ years in data engineering, ML infra, or distributed systems at large scale.
  • Strong software engineering background with distributed systems and Ray/Spark experience.
  • Proficiency in Python and systems languages (C++, Rust, Go, or Java).
  • Deep knowledge of scalable databases and data storage.
  • Experience optimizing GPU inference for VLMs/LLMs and data curation for training.

Responsibilities

  • Architect and optimize large-scale distributed pipelines processing billions of assets into training-ready data.
  • Scale inference throughput across pipelines and optimize hardware utilization.
  • Design and own architecture for databases, storage, and GPU/CPU workloads.
  • Collaborate with modeling teams to translate data requirements into pipelines.
  • Bridge data engineering and applied ML across multiple teams.

Skills

Distributed systems
Ray/Spark
Python
Systems programming
Databases and data lakes
GPU inference optimization
Data curation for model training
Cross-team collaboration
Debugging distributed systems

Education

Bachelor's/Master's/PhD in CS/Engineering/ML

Tools

Ray
Spark

Job description

ABOUT THE ROLE

We're looking for a Principal ML Engineer to architect and scale the multimodal data processing pipelines and infrastructure behind Adobe Firefly's multimodal foundation models (image, video, audio). In this role, you'll sit at the intersection of data engineering and applied ML building distributed, GPU-accelerated systems that turn billions of raw assets into training-ready data at scale.

Your work will directly determine how fast and how well Adobe models can learn, directly impacted by the throughput and reliability of our data pipelines, and the quality of data that reaches training. This is a senior individual contributor role with broad technical influence across data, infrastructure, and modeling teams.

WHAT YOU'LL DO
OPTIMIZE DATA PROCESSING PIPELINES AT SCALE
  • Architect and optimize large-scale distributed pipelines that process billions of images, video, and audio assets through ML workflows into training-ready data.
  • Scale up inference throughput across the pipeline (batching, parallelism, hardware utilization) to turn raw collected data into training data faster and more cheaply.
  • Identify and eliminate bottlenecks across ingestion, processing, and delivery, from storage and I/O to compute scheduling.
ARCHITECT SCALABLE DATA INFRASTRUCTURE
  • Design systems that reliably store, index, and serve billions of data points, each requiring substantial processing spanning large-scale databases, distributed storage, and high-throughput compute.
  • Apply deep expertise in distributed systems and frameworks such as Ray (or equivalent) to orchestrate large-scale, GPU/CPU-heavy data workloads.
  • Own architecture decisions including database and storage choices, job scheduling, GPU cluster utilization that let the platform scale alongside data and model growth.
DRIVE DATA CURATION FOR MODEL TRAINING
  • Bring a strong ML background, especially inference optimization for VLMs and LLMs and data curation for training.
  • Partner closely with modeling teams to understand what data improves training outcomes, and translate that into pipeline and curation requirements.
  • Operate as a hands‑on technical leader who bridges data engineering and applied ML.
WHAT YOU NEED TO SUCCEED
  • 10+ years of experience in data engineering, ML infrastructure, or distributed systems, including work at large scale (billions of records or assets).
  • Strong software engineering background, with hands‑on expertise in distributed systems and frameworks such as Ray, Spark, or equivalent large-scale data processing frameworks.
  • Proficiency in Python, plus strong experience in a systems-level language (C++, Rust, Go, or Java) with strong debugging skills across distributed and ML-centric runtime environments.
  • Deep knowledge of databases and storage systems at scale such as data lakes, indexing, and retrieval across billions of data points.
  • Strong ML background, particularly experience in optimizing GPU inference pipelines for VLMs, LLMs, or other large models (batching, quantization, serving, throughput/latency tradeoffs).
  • Experience with data curation for model training: understanding what makes data valuable for training generative or multimodal models, not just how to move it efficiently.
  • Comfort operating across the full stack, from low‑level systems and GPU optimization to higher‑level data strategy and curation decisions.
  • Ability to communicate clearly and partner effectively across data, infrastructure, and modeling teams.
  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Machine Learning, or a related field.

Expected Pay Range:

Our compensation reflects the cost of labor across several U.S. geographic markets, and we pay differently based on those defined markets. The U.S. pay range for this position is $206,300 – $388,000 annually. Pay within this range varies by work location and may also depend on job-related knowledge, skills, and experience. Your recruiter can share more about the specific salary range for the job location during the hiring process. In California, the pay range for this position is $268,000 – $388,000. In Washington, the pay range for this position is $247,200 – $357,900.

Adobe is proud to be an Equal Employment Opportunity employer. We do not discriminate based on gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, veteran status, or any other protected characteristic.

Adobe aims to make our Careers website and recruiting process accessible to any and all users. If you have a disability or special need that requires accommodation to navigate our website or complete the application process, email accommodations@adobe.com or call +1 408-536-3015.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Scientist - Data Pipeline Engineer
Principal Scientist - Data Pipeline Engineer

Adobe • San Jose (CA)

On-site
USD 268,000 - 388,000
Senior Applied Scientist
Senior Applied Scientist

Adobe Inc. • San Jose (CA)

On-site
USD 162,000 - 302,000
Senior Machine Learning Engineer, Services/MLOps
Senior Machine Learning Engineer, Services/MLOps

Adobe • San Jose (CA)

On-site
USD 183,000 - 266,000
Comprehensive benefits programs
Career growth opportunities
Manager, Applied Science for Data
Manager, Applied Science for Data

Adobe Inc. • San Jose (CA)

On-site
USD 211,000 - 307,000
Director, ML Services Engineering
Director, ML Services Engineering

Adobe Inc. • San Jose (CA)

On-site
USD 265,000 - 385,000
Health insurance
Retirement savings plan
Flexible working hours
Principal Machine Learning Engineer
Principal Machine Learning Engineer

Adobe • San Jose (CA)

On-site
USD 262,000 - 379,000
Sr Manager, Machine Learning Engineering
Sr Manager, Machine Learning Engineering

Adobe • San Jose (CA)

On-site
USD 178,000 - 352,000
Senior Machine Learning Engineer, Services/MLOps
Senior Machine Learning Engineer, Services/MLOps

Adobe • San Francisco (CA)

On-site
USD 183,000 - 265,000
Senior Machine Learning Engineer ServicesMLOps
Senior Machine Learning Engineer ServicesMLOps

Adobe • San Jose (CA)

On-site
USD 151,800 - 265,350
Staff Applied Scientist, Firefly Foundry
Staff Applied Scientist, Firefly Foundry

Adobe Inc. • San Jose (CA)

On-site
USD 211,000 - 307,000