Principal Data Engineer: AI-Ready Pipelines & Platforms

JobCubby

Michigan

Hybrid

USD 160,000 - 212,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

GM Vehicle program

Job summary

General Motors is seeking a Principal Data Engineer (Level 8) to lead cross-team data initiatives, set direction for key data domains, and drive scalable data infrastructure and AI-ready data products. You will transform raw data into trusted datasets, shape architecture, and mentor engineers and data scientists across multiple teams.

The role emphasizes data engineering leadership, DevOps/DataOps/MLOps practices, and collaboration with business partners to deliver high-impact AI initiatives and

Qualifications

  • Bachelor’s degree in Computer Science, Software Engineering, Data Engineering, or related field, or equivalent experience.
  • 8+ years of relevant full-time experience in data engineering or closely related roles; or equivalent depth of knowledge.
  • Strong, hands-on experience in data engineering, including: End-to-end pipeline development (ingestion, transformation, orchestration, monitoring).
  • Data modeling (batch and streaming), data integration, and production support for enterprise data platforms.
  • Building and operating highly reliable, scalable data products in production environments.
  • Extensive experience designing and optimizing batch and streaming data pipelines using Databricks, Apache Spark, Delta Lake, and modern cloud data patterns.
  • Proven experience supporting AI or machine-learning use cases through high-quality data preparation, feature-ready datasets, experimentation workflows, and model-development enablement.
  • Proficiency in: Python or Scala.
  • SQL, including performance tuning and working with large-scale datasets.
  • Relational and non-relational data storage technologies (e.g., data warehouses, NoSQL, key-value stores, document stores).
  • Deep experience with cloud platforms – Azure strongly preferred; AWS or GCP also considered.
  • Extensive experience designing, building, and optimizing scalable batch and streaming data pipelines using Databricks (Apache Spark, Delta Lake) to support Medallion Architecture and other modern data patterns.
  • Demonstrated experience with modern cloud data platforms, distributed processing, and production-grade data pipelines, including: Data quality frameworks and observability.
  • Orchestration tools and job scheduling.
  • Performance optimization and cost management.
  • Proven ability to work independently and lead through influence, managing broad, ambiguous technical challenges and delivering high-impact solutions with minimal guidance.
  • Experience driving cross-functional collaboration across engineering, analytics, product, and business teams to deliver data solutions that enable measurable business outcomes, including AI and advanced analytics.
  • Strong understanding of statistics, machine learning, experimentation, and data mining concepts used to drive informed decisions.
  • Demonstrated ability to: Prepare and explore data at scale.
  • Support model development workflows.
  • Help validate analytical outputs and production model behavior in partnership with data scientists.
  • Ability to translate complex analytical and AI needs into scalable, maintainable data solutions and help move data science work from exploration into repeatable, governed, and automated delivery patterns.
  • Strong communication and storytelling skills to connect technical work with business value and to convert complex findings and trade-offs into clear, actionable recommendations for diverse audiences.

Responsibilities

  • Design, build, and productionize reliable, scalable, and secure data pipelines and data products in Azure Databricks that support AI, analytics, and operational use cases across multiple business domains.
  • Lead the end-to-end transformation of raw data from numerous, heterogeneous source systems into trusted, well-structured, and governed datasets suitable for downstream analytics, model development, and AI enablement.
  • Define and champion architecture, design patterns, and best practices (e.g., Medallion Architecture, Delta Lake standards, data quality and observability) that can be adopted across teams.
  • Drive strategic improvements in internal processes, delivery patterns, and technical solutions that support broader functional and enterprise data strategy, increasing efficiency, reliability, and speed of delivery.
  • Solve complex, ambiguous, and non-standard data engineering problems using advanced analytical and problem-solving techniques, demonstrating strong ownership, risk-aware decision making, and sound technical judgment.
  • Partner closely with data scientists, analysts, software engineers, product stakeholders, and business leaders to ensure data is accessible, trustworthy, and aligned to high-impact business outcomes and AI initiatives.
  • Lead AI and data science enablement by defining and delivering high-quality, feature-ready data, experimentation workflows, and scalable patterns for model development, deployment, monitoring, and continuous improvement.
  • Provide technical leadership for large, multi-sprint or multi-team initiatives, including defining scope, breaking down work, and ensuring cohesive, high-quality delivery across contributors.
  • Influence key engineering decisions, technology choices, and long-term roadmaps for your area of responsibility, balancing innovation with operational excellence and sustainability.
  • Mentor and coach engineers and data scientists through deep technical guidance, code and design reviews, knowledge sharing, and strong engineering practices consistent with and extending beyond Level 7 expectations.
  • Help evolve team culture, practices, and tooling around DevOps, DataOps, and MLOps, including CI/CD for data pipelines, testing strategies, observability, governance, and reliability.
  • Communicate complex technical concepts, trade-offs, and recommendations clearly to both technical and non-technical audiences, enabling informed decisions at multiple levels of the organization.
  • Establish repeatable patterns for exposing governed, curated, contract-backed data products and semantic models to analytics and AI applications.
  • Contribute to multi-agent systems, agent orchestration, supervisor-agent patterns, and reusable AI services that turn governed data and domain context into actionable intelligence.
  • This role will help establish the data and AI foundation for trusted, scalable, and reusable intelligence across GM. By combining strong data engineering with production AI capabilities, you will help teams move from fragmented data and exploratory analysis to governed data products, reliable AI experiences, and faster, better-informed decisions.

Skills

End-to-End pipeline development
Data modeling
Databricks
Apache Spark
Delta Lake
Cloud platforms
Python
SQL
Data quality and observability
MLOps

Education

Bachelor’s degree in Computer Science or related field

Tools

Azure Databricks
Databricks (Apache Spark)
Delta Lake

Job description

General Motors is seeking a Principal Data Engineer (Level 8) to lead cross-team data initiatives, set direction for key data domains, and drive scalable data infrastructure and AI-ready data products. You will transform raw data into trusted datasets, shape architecture, and mentor engineers and data scientists across multiple teams.

The role emphasizes data engineering leadership, DevOps/DataOps/MLOps practices, and collaboration with business partners to deliver high-impact AI initiatives and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Data Engineer: AI-Ready Pipelines & Platform Lead
Staff Data Engineer: AI-Ready Pipelines & Platform Lead

General Motors • Michigan

Hybrid
USD 160,000 - 212,000
Company vehicle program
Relocation benefits
Staff Data Engineer — AI & Data Platform Lead
Staff Data Engineer — AI & Data Platform Lead

General Motors • United States

Hybrid
USD 160,000 - 212,000
Company vehicle program
Relocation support
Principal Data Engineer: AI-Ready Data Platform Lead
Principal Data Engineer: AI-Ready Data Platform Lead

Releady • Denver (CO)

Hybrid
Senior AI Data Engineer for Quality & Productivity
Senior AI Data Engineer for Quality & Productivity

General Motors • Michigan

Hybrid
USD 120,000 - 160,000
Hybrid AI/ML Software Engineer — Data Pipelines
Hybrid AI/ML Software Engineer — Data Pipelines

General Motors • Warren (MI), Northern (KY)

Hybrid
USD 76,000 - 110,000
Health & wellbeing programs
GM vehicle discounts
Relocation benefits possible
Agentic Data Engineer: AI-Powered Vehicle Data Platforms
Agentic Data Engineer: AI-Powered Vehicle Data Platforms

General Motors • Austin (TX)

Hybrid
USD 248,000 - 315,000
Health benefits
GM vehicle discounts
Tuition assistance
Data Engineering Manager - Lakehouse & AI Platform
Data Engineering Manager - Lakehouse & AI Platform

General Motors • Warren (MI)

Hybrid
USD 180,000 - 230,000
Data & AI Systems Engineer for Finance Insights
Data & AI Systems Engineer for Finance Insights

General Motors • Warren (MI)

Hybrid
USD 90,000 - 130,000
Hybrid work model
Relocation benefits
Principal Engineer - AI-Driven Developer Platforms
Principal Engineer - AI-Driven Developer Platforms

Jobtailor • Missouri

On-site
USD 150,000 - 190,000
Senior Data Engineer, Vehicle Telemetry Platform & AI
Senior Data Engineer, Vehicle Telemetry Platform & AI

General Motors • Warren (MI)

Hybrid
USD 129,000 - 169,000
Health & wellbeing benefits
Retirement plan
GM vehicle discounts
+3