Staff Data Engineer

Jobtailor

Missouri

On-site

USD 150,000 - 190,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Jobtailor is seeking a senior Data Engineer to design, build, and productionize scalable data pipelines and products on Azure Databricks. You will transform raw data into governed datasets, champion architecture patterns, and drive data quality and observability across multi-team initiatives.

You will mentor engineers and data scientists, collaborate with product stakeholders, and enable AI/ML workflows with feature-ready data, reproducible experiments, and scalable deployment patterns.

Qualifications

  • Bachelor’s degree in Computer Science, Software Engineering, Data Engineering, or related field, or equivalent experience.
  • 8+ years of relevant full-time experience in data engineering or closely related roles, or equivalent depth of knowledge.
  • Hands-on experience with end-to-end data pipeline development, including ingestion, transformation, orchestration, and monitoring.
  • Experience with data modeling, data integration, and production support for enterprise data platforms.
  • Experience building and operating reliable, scalable data products in production.
  • Extensive experience with Databricks, Apache Spark, Delta Lake, batch and streaming pipelines, and modern cloud data patterns.
  • Experience supporting AI or machine-learning use cases through data preparation, feature-ready datasets, experimentation workflows, and model-development enablement.
  • Proficiency in Python or Scala.
  • Proficiency in SQL, including performance tuning and large-scale datasets.
  • Experience with relational and non-relational data storage technologies.
  • Deep experience with cloud platforms; Azure strongly preferred, AWS or GCP also considered.
  • Experience with data quality frameworks, observability, orchestration, job scheduling, performance optimization, and cost management.
  • Ability to work independently and lead through influence on broad, ambiguous technical challenges.
  • Experience driving cross-functional collaboration across engineering, analytics, product, and business teams.
  • Understanding of statistics, machine learning, experimentation, and data mining concepts.
  • Ability to prepare and explore data at scale, support model development workflows, and validate analytical outputs and production model behavior.
  • Ability to translate complex analytical and AI needs into scalable, maintainable data solutions.
  • Strong communication and storytelling skills.
  • Advanced degree in a related field preferred.
  • Experience as a principal- or staff-level engineer in large-scale cloud-based data environments preferred.
  • Experience leading or technically directing engineers or data scientists preferred.
  • Background in manufacturing, automotive, industrial IoT, or similar domains preferred.
  • Experience with MLOps, data governance, privacy, security controls, event-driven and streaming architectures, and enterprise data platforms preferred.
  • GM does not provide immigration-related sponsorship for this role; applicants must not require GM immigration sponsorship now or in the future

Responsibilities

  • Design, build, and productionize reliable, scalable, and secure data pipelines and data products in Azure Databricks
  • Transform raw data from heterogeneous source systems into trusted, structured, and governed datasets
  • Define and champion architecture, design patterns, Medallion Architecture, Delta Lake standards, data quality, and observability practices
  • Drive strategic improvements to data processes, delivery patterns, and technical solutions
  • Solve complex and non-standard data engineering problems
  • Partner with data scientists, analysts, software engineers, product stakeholders, and business leaders
  • Enable AI and data science through feature-ready data, experimentation workflows, and scalable model development, deployment, monitoring, and improvement patterns
  • Provide technical leadership for large multi-sprint and multi-team initiatives
  • Influence engineering decisions, technology choices, and long-term roadmaps
  • Mentor and coach engineers and data scientists through technical guidance, reviews, and knowledge sharing
  • Evolve DevOps, DataOps, and MLOps practices, including CI/CD, testing, observability, governance, and reliability
  • Communicate technical concepts, trade-offs, and recommendations to technical and non-technical audiences
  • Establish governed, curated, contract-backed data products and semantic models for analytics and AI applications
  • Contribute to multi-agent systems, agent orchestration, supervisor-agent patterns, and reusable AI services
  • Help establish GM’s foundation for trusted, scalable, reusable data and AI intelligence

Skills

Azure Databricks
Apache Spark
Delta Lake
Data Pipeline Development
MLOps
Data Engineering
Data Modeling
Data Integration
SQL Performance Tuning
Python
Scala
Data Quality Frameworks
Observability
Job Scheduling
Cloud Platforms

Education

Bachelor’s degree in Computer Science, Software Engineering, Data Engineering, or related field

Tools

CI/CD
DataOps
Machine Learning
Event-Driven Architectures
Enterprise Data Platforms

Job description


  • Design, build, and productionize reliable, scalable, and secure data pipelines and data products in Azure Databricks

  • Transform raw data from heterogeneous source systems into trusted, structured, and governed datasets

  • Define and champion architecture, design patterns, Medallion Architecture, Delta Lake standards, data quality, and observability practices

  • Drive strategic improvements to data processes, delivery patterns, and technical solutions

  • Solve complex and non-standard data engineering problems

  • Partner with data scientists, analysts, software engineers, product stakeholders, and business leaders

  • Enable AI and data science through feature-ready data, experimentation workflows, and scalable model development, deployment, monitoring, and improvement patterns

  • Provide technical leadership for large multi-sprint and multi-team initiatives

  • Influence engineering decisions, technology choices, and long-term roadmaps

  • Mentor and coach engineers and data scientists through technical guidance, reviews, and knowledge sharing

  • Evolve DevOps, DataOps, and MLOps practices, including CI/CD, testing, observability, governance, and reliability

  • Communicate technical concepts, trade-offs, and recommendations to technical and non-technical audiences

  • Establish governed, curated, contract-backed data products and semantic models for analytics and AI applications

  • Contribute to multi-agent systems, agent orchestration, supervisor-agent patterns, and reusable AI services

  • Help establish GM’s foundation for trusted, scalable, reusable data and AI intelligence


Requirements


  • Bachelor’s degree in Computer Science, Software Engineering, Data Engineering, or related field, or equivalent experience

  • 8+ years of relevant full-time experience in data engineering or closely related roles, or equivalent depth of knowledge

  • Hands-on experience with end-to-end data pipeline development, including ingestion, transformation, orchestration, and monitoring

  • Experience with data modeling, data integration, and production support for enterprise data platforms

  • Experience building and operating reliable, scalable data products in production

  • Extensive experience with Databricks, Apache Spark, Delta Lake, batch and streaming pipelines, and modern cloud data patterns

  • Experience supporting AI or machine-learning use cases through data preparation, feature-ready datasets, experimentation workflows, and model-development enablement

  • Proficiency in Python or Scala

  • Proficiency in SQL, including performance tuning and large-scale datasets

  • Experience with relational and non-relational data storage technologies

  • Deep experience with cloud platforms; Azure strongly preferred, AWS or GCP also considered

  • Experience with data quality frameworks, observability, orchestration, job scheduling, performance optimization, and cost management

  • Ability to work independently and lead through influence on broad, ambiguous technical challenges

  • Experience driving cross-functional collaboration across engineering, analytics, product, and business teams

  • Understanding of statistics, machine learning, experimentation, and data mining concepts

  • Ability to prepare and explore data at scale, support model development workflows, and validate analytical outputs and production model behavior

  • Ability to translate complex analytical and AI needs into scalable, maintainable data solutions

  • Strong communication and storytelling skills

  • Advanced degree in a related field preferred

  • Experience as a principal- or staff-level engineer in large-scale cloud-based data environments preferred

  • Experience leading or technically directing engineers or data scientists preferred

  • Background in manufacturing, automotive, industrial IoT, or similar domains preferred

  • Experience with MLOps, data governance, privacy, security controls, event-driven and streaming architectures, and enterprise data platforms preferred

  • GM does not provide immigration-related sponsorship for this role; applicants must not require GM immigration sponsorship now or in the future


Core Competencies

Demonstrates expertise in designing and building scalable data pipelines and products using Azure Databricks, Apache Spark, and Delta Lake. Proficient in data modeling, integration, and supporting AI use cases through feature-ready datasets and experimentation workflows.


Highest-signal resume keywords


  • Azure Databricks

  • Apache Spark

  • Delta Lake

  • Data Pipeline Development

  • MLOps


Hard Skills


  • Data Engineering

  • Data Modeling

  • Data Integration

  • SQL Performance Tuning

  • Python

  • Scala

  • Data Quality Frameworks

  • Observability

  • Job Scheduling

  • Cloud Platforms


Soft Skills


  • Strong Communication Skills

  • Cross-Functional Collaboration

  • Technical Leadership

  • Mentoring

  • Problem Solving


Industry Keywords


  • Manufacturing

  • Automotive

  • Industrial IoT

  • Data Governance

  • Privacy Controls


Tools & Technologies


  • CI/CD

  • DataOps

  • Machine Learning

  • Event-Driven Architectures

  • Enterprise Data Platforms

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer – AI Quality and Productivity
Senior Software Engineer – AI Quality and Productivity

Jobtailor • Missouri

On-site
USD 125,000 - 180,000
Data Engineer
Data Engineer

Jobtailor • Town of Montana (WI)

On-site
USD 120,000 - 180,000
Senior Data Engineer
Senior Data Engineer

Jobtailor • New Jersey

On-site
USD 120,000 - 170,000
Principal Data Engineer
Principal Data Engineer

Jobtailor • Vienna (VA)

On-site
USD 150,000 - 190,000
AI/ML Data Scientist, GPSSC
AI/ML Data Scientist, GPSSC

Jobtailor • Missouri

On-site
USD 120,000 - 170,000
Azure Databricks Engineer (Dallas, TX)
Azure Databricks Engineer (Dallas, TX)

Cedent • Dallas (TX)

On-site
USD 120,000 - 150,000
Staff Data Engineer
Staff Data Engineer

Jobtailor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Data Engineer
Staff Data Engineer

General Motors • Warren (MI)

Hybrid
USD 140,000 - 210,000
Lead Machine Learning Engineer, Python, AWS, SQL, GenAI
Lead Machine Learning Engineer, Python, AWS, SQL, GenAI

Jobtailor • New York (NY)

On-site
USD 140,000 - 210,000
Data Architect C2C requirements AWS, Databricks & Generative AI
Data Architect C2C requirements AWS, Databricks & Generative AI

Tech Mirrors • Dallas (TX)

On-site
USD 140,000 - 180,000