Staff ML Data Engineer (Datagrid)

Procore Technologies, Inc.

San Francisco (CA)

Hybrid

USD 227.000 - 313.000

Vollzeit

Vor 12 Tagen
Bewerbungsgenerator

Mach aus dieser Rolle ein Bewerbungsgespräch — ein Lebenslauf und ein Anschreiben, die darauf ausgerichtet sind, was dieser Arbeitgeber sucht.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Procore Technologies, Inc. is seeking a Staff ML Data Engineer in San Francisco to design and run data systems powering frontier-scale ML research and AI products.

You will collaborate with ML researchers, engineers, and architects to build production-ready pipelines for multimodal data and ensure robust data quality and observability. The role is hybrid with onsite presence in SF at least three days per week and offers technical leadership while remaining hands-on.

Qualifikationen

  • Bachelor’s or Master’s in CS/Engineering or related field.
  • 8+ years designing complex data systems in production/research environments.
  • Strong SQL and Python; experience with data-intensive or distributed systems.
  • Proven experience building scalable data pipelines for ML training, evaluation, or inference.
  • Understanding of data modeling, dataset lifecycle, and data quality best practices.
  • Ability to work in ambiguous spaces and collaborate with researchers/architects.
  • Demonstrated ability to lead via technical contribution, mentorship, and setting standards.
  • Strong communication to explain tradeoffs to researchers and engineers.

Aufgaben

  • Act as the technical lead for data engineering supporting frontier model research and ML systems.
  • Design, build, and maintain scalable batch and streaming pipelines for multimodal data.
  • Translate experimental workflows into reliable, repeatable data systems with researchers and architects.
  • Lead dataset curation, versioning, and lineage workflows to support rapid experimentation.
  • Establish standards for data quality, validation, observability, and cost efficiency across pipelines.
  • Contribute to data architecture decisions spanning research environments and production systems.
  • Identify gaps in data workflows and run proofs-of-concept to evaluate improvements.
  • Mentor engineers through code reviews, design discussions, and hands-on collaboration.

Kenntnisse

SQL
Python
Data pipelines
Data modeling
Dataset lifecycle
Data quality
Observability
Leadership
Communication

Ausbildung

Bachelor's or Master's in Computer Science or related field

Tools

Databricks
Spark
Lakehouse
Kafka
Pub/Sub
Airflow
Dagster
AWS
GCP
CI/CD
Infrastructure-as-code

Jobbeschreibung

We’re looking for a Staff ML Data Engineer to join Procore’s AI & Frontier Models organization. In this role, you’ll be responsible for designing and building the data systems that power frontier‑scale machine learning research and applied AI products, with a particular focus on spatial intelligence and multimodal data. The primary goal of this role is to ensure that researchers and engineers can reliably discover, curate, transform, and operate on large‑scale datasets that move from experimentation to production. As a Staff ML Data Engineer, you’ll work closely with ML researchers, applied ML engineers, and system architects to turn ambiguous research needs into scalable, production‑ready data pipelines. You’ll remain deeply hands‑on while providing technical leadership in data architecture, quality, and operational excellence. This is an opportunity to shape how Procore builds, evaluates, and deploys frontier models by ensuring the underlying data systems are robust, observable, and designed for iteration. This role reports reports into the Manager, Software Engineering, and is based in our San Francisco office, supporting Procore's Datagrid AI Division. Given the collaborative and fast moving nature of this work, we are seeking candidates who are available to work onsite in a hybrid model at a minimum of 3 days per week. This is an immediate opening!

What you’ll do

Act as the technical lead for data engineering efforts supporting frontier model research and applied ML systems. Design, build, and maintain scalable batch and streaming pipelines for multimodal data (e.g., documents, images, spatial metadata). Partner closely with researchers and architects to translate experimental workflows into reliable, repeatable data systems. Lead the development of dataset curation, versioning, and lineage workflows that support rapid experimentation and reproducibility. Establish and uphold standards for data quality, validation, observability, and cost efficiency across AI data pipelines. Contribute to data architecture decisions spanning research environments and production systems. Identify gaps or inefficiencies in existing data workflows and run proofs‑of‑concept to evaluate improvements. Mentor other engineers through code reviews, design discussions, and hands‑on collaboration.

What we’re looking for

Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience. 8+ years of experience designing and operating complex data systems in production or research‑adjacent environments. Strong proficiency in SQL and Python; experience with data‑intensive or distributed systems. Proven experience building scalable data pipelines that support machine learning training, evaluation, or inference workflows. Solid understanding of data modeling, dataset lifecycle management, and data quality best practices. Comfort operating in highly ambiguous problem spaces and collaborating closely with researchers and architects. Demonstrated ability to lead through direct technical contribution, mentorship, and setting engineering standards. Strong communication skills, with the ability to explain technical tradeoffs to both research and engineering audiences.

Nice to have experience with technologies such as:
  • ML & Research Data: Large‑scale dataset curation, annotation workflows, experiment tracking, reproducibility tooling
  • Data Platforms: Databricks, Spark, lakehouse architectures, cloud data warehouses
  • Streaming & Pipelines: Kafka, Pub/Sub, event‑driven data architectures
  • Orchestration & Observability: Airflow, Dagster, data quality and lineage tools
  • Cloud & Infrastructure: AWS or GCP, containerized data workloads, CI/CD, infrastructure‑as‑code
  • Performance & Cost: Optimizing data pipelines for GPU‑backed training and large‑scale inference workloads
Additional Information

Base Pay Range: 227,332.00 - 312,581.50 USD Annual This role may also be eligible for Equity Compensation and/or Bonus Incentive Compensation. Procore is committed to offering competitive, fair, and commensurate compensation. Actual compensation will be based on a candidate’s job‑related skills, experience, education or training, and location. For Los Angeles County (unincorporated) Candidates: Procore will consider for employment all qualified applicants, including those with arrest or conviction records, in accordance with the requirements of applicable federal, state, and local laws, including the City of Los Angeles’ Fair Chance Initiative for Hiring Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act. A criminal history may have a direct, adverse, and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment: 1. appropriately managing, accessing, and handling confidential information including proprietary and trade secret information, as well as accessing Procore's information technology systems and platforms; 2. interacting with and occasionally having unsupervised contact with internal/external customers, stakeholders, and/or colleagues; and 3. exercising sound judgment.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Staff ML Data Engineer (Datagrid)
Staff ML Data Engineer (Datagrid)

Procore Technologies • San Francisco (CA)

Hybrid
USD 227.000 - 313.000
Equity compensation
Bonus incentive
Staff ML Data Engineer (Datagrid)
Staff ML Data Engineer (Datagrid)

DroneDeploy, Inc. • San Francisco (CA)

Vor Ort
USD 227.000 - 313.000
Equity compensation
Bonus incentive
Hybrid on-site in San Francisco
Staff ML Data Engineer (Datagrid)
Staff ML Data Engineer (Datagrid)

Procore • San Francisco (CA)

Vor Ort
USD 227.332 - 312.581
Equity compensation
Bonus incentive compensation
Hybrid work model (3 days onsite)
Staff Applied Research Scientist (Datagrid)
Staff Applied Research Scientist (Datagrid)

Procore Technologies, Inc. • San Francisco (CA)

Hybrid
USD 227.000 - 313.000
Equity compensation
Bonus incentive
Staff Applied Research Scientist (Datagrid)
Staff Applied Research Scientist (Datagrid)

Procore Technologies • San Francisco (CA)

Hybrid
USD 227.000 - 313.000
Equity compensation
Bonus incentive
AI Application Architect
AI Application Architect

DroneDeploy, Inc. • Austin (TX)

Vor Ort
USD 233.000 - 321.000
Equity compensation
Bonus incentives
AI Application Architect
AI Application Architect

Procore Technologies, Inc. • Austin (TX)

Vor Ort
USD 233.000 - 321.000
AI Application Architect
AI Application Architect

Procore • Austin (TX)

Vor Ort
USD 233.000 - 321.000
Senior Principal Software Engineer, Ruby
Senior Principal Software Engineer, Ruby

Procore Technologies, Inc. • Austin (TX)

Hybrid
USD 233.000 - 321.000
Equity compensation
Bonus incentive
Senior Product Manager, Domain Team Enablement
Senior Product Manager, Domain Team Enablement

Procore • San Francisco (CA)

Vor Ort
USD 193.844 - 266.536
Equity compensation
Bonus incentive