Data Engineer: Scalable AI Pipelines & Safety

OpenAI

Mountain View (CA)

On-site

USD 150,000 - 190,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Relocation assistance

Job summary

OpenAI is seeking a Data Engineer to lead the design, build, and management of data pipelines and core tables powering analyses, safety systems, product growth, and revenue insights. This role collaborates closely with researchers behind ChatGPT to train models and deliver to users.

You will design scalable data infrastructure, implement fault-tolerant ingestion, and ensure data security and compliance, enabling data-driven decisions across multiple teams and products.

Qualifications

  • 3+ years of data engineering experience and 8+ years in software engineering.
  • Proficiency in Python, Scala, or Java.
  • Experience with distributed processing (Hadoop, Flink) and storage (S3, HDFS).
  • Experience with ETL schedulers like Airflow, Dagster, or Prefect.
  • Solid Spark knowledge: write, debug, optimize Spark code.

Responsibilities

  • Design, build and manage data pipelines and data warehouse integration.
  • Develop canonical datasets for product metrics: user growth, engagement, revenue.
  • Collaborate with Infra, Data Science, Product, Marketing, Finance, and Research to meet data needs.
  • Implement fault-tolerant data ingestion and processing systems.
  • Contribute to data architecture and engineering decisions.
  • Ensure data security, integrity, and compliance with standards.

Skills

Python
Scala
Java
Spark
Hadoop
Flink
Airflow
Dagster
Prefect
Distributed storage (S3/HDFS)

Job description

OpenAI is seeking a Data Engineer to lead the design, build, and management of data pipelines and core tables powering analyses, safety systems, product growth, and revenue insights. This role collaborates closely with researchers behind ChatGPT to train models and deliver to users.

You will design scalable data infrastructure, implement fault-tolerant ingestion, and ensure data security and compliance, enabling data-driven decisions across multiple teams and products.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

OpenAI • Los Angeles (CA)

On-site
USD 120,000 - 160,000
Data Engineer - Infrastructure Analytics at Scale
Data Engineer - Infrastructure Analytics at Scale

Neura Market • San Francisco (CA)

On-site
USD 140,000 - 190,000
Lead Data Engineer, Pipelines & Analytics
Lead Data Engineer, Pipelines & Analytics

OpenAI • Los Angeles (CA)

On-site
USD 120,000 - 160,000
Data Engineer, Experimentation Platform (Hybrid)
Data Engineer, Experimentation Platform (Hybrid)

OpenAI • Bellevue (WA)

Hybrid
USD 293,000 - 325,000
Lead Data Engineer - AI Experimentation & Metrics Pipelines
Lead Data Engineer - AI Experimentation & Metrics Pipelines

Slope • Seattle (WA)

Hybrid
USD 293,000 - 325,000
Data Engineer
Data Engineer

OpenAI • Mountain View (CA)

On-site
USD 150,000 - 190,000
Relocation assistance
Data Foundations Engineer – Scalable Pipelines & Open AI
Data Foundations Engineer – Scalable Pipelines & Open AI

Reflection • San Francisco (CA)

On-site
USD 150,000 - 210,000
Top-tier compensation
Stock options
Health & wellness
+3
Data Engineer: Scalable Pipelines & Analytics Impact
Data Engineer: Scalable Pipelines & Analytics Impact

Apply • United States

On-site
USD 85,000 - 110,000
Data Infrastructure Engineer for Scalable AI Pipelines
Data Infrastructure Engineer for Scalable AI Pipelines

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Data Engineer — AI-Driven Pipelines & Analytics
Data Engineer — AI-Driven Pipelines & Analytics

Airtable • City of Syracuse (NY)

On-site
USD 179,000 - 222,000