Sr Data Engineer

Mcgraw Hill

New York (NY)

Remote

USD 135,000 - 160,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

McGraw Hill is seeking a Senior Data Engineer for its Data & Analytics team. The role focuses on building scalable data pipelines and architectures on AWS and Databricks, delivering reliable insights across financial, product, and customer domains.

You will design end-to-end data solutions, implement SCDs with Delta Lake, and collaborate with stakeholders using Agile practices. This remote US role offers competitive compensation and benefits.

Qualifications

  • 5+ years of data engineering experience with cloud-based data platforms.
  • Strong hands-on expertise in Databricks and Delta Lake architectures.
  • Experience with SCDs (types 1–3) and lakehouse patterns.
  • Proficiency in SQL and Python/Scala for data processing.
  • Experience with CI/CD and IaC (Terraform) for data pipelines.

Responsibilities

  • Design and deliver data solutions on Databricks with Delta Lake lakehouses.
  • Implement Delta Lake MERGE for SCD types 1–3 and maintain a unified lakehouse model.
  • Operate Databricks Workflows, job scheduling, and SLA adherence.
  • Integrate with Git-based workflows, Jira, and Agile/Kanban delivery.
  • Apply Unity Catalog governance and cloud-native lakehouse design on AWS.
  • Translate business requirements into production-grade data pipelines.

Skills

Databricks
AWS
Python
SQL
Scala
Spark
Terraform
Git
Jira

Tools

Delta Lake
Unity Catalog
Delta Live Tables
MLflow
Apache Airflow
Databricks SQL
S3
Redshift

Job description

Overview

Build the Future

At McGraw Hill, we are dedicated to delivering digital learning experiences that transform education for learners and educators. Our focus is on creating seamless, impactful products that truly benefit our users while supporting growth and collaboration across teams. We foster a culture that values innovation, teamwork, and a balance between career growth and personal well-being.

How can you make an impact?

The Senior Data Engineer in Data and Analytics is responsible for advancing McGraw-Hill Education's (MHE) business intelligence and data platform capabilities, delivering scalable, reliable, and actionable insights across financial, product, customer, user, and third-party data domains. This role is deeply hands‑on - designing, building, and optimizing end-to-end data pipelines and architectures on AWS (including services such as S3, Glue, Redshift, Lambda, EMR, and Step Functions) and Databricks (leveraging Delta Lake, Unity Catalog, and MLflow where applicable).

The Senior Data Engineer will architect and implement dynamic reporting, analytics, and data modeling solutions that drive measurable outcomes in the education domain, while ensuring the performance, efficiency, and reliability of the broader Data Platform. The ideal candidate brings a strong data engineering foundation with deep, hands‑on expertise in AWS cloud infrastructure and Databricks, including experience with Delta Lake architecture, medallion (Bronze/Silver/Gold) data design patterns, and Databricks Workflows for pipeline orchestration. Advanced proficiency in SQL and experience with Python or Scala for large-scale data transformation are essential. Familiarity with infrastructure-as-code (e.g., Terraform) and CI/CD practices for data pipelines is a strong plus.

This role requires close collaboration with business stakeholders, data analysts, and product teams to translate complex data requirements into robust, production‑grade engineering solutions - ensuring timely, high-quality delivery across all data initiatives.

This is a remote position open to applicants authorized to work for any employer within the United States.
What You'll Do
  • Senior Data Engineer must have prior hands‑on experience designing and delivering data solutions on Databricks, including building and maintaining lakehouses using Delta Lake with a medallion (Bronze/Silver/Gold) architecture.
  • Strong knowledge working with data from financial and operational systems, with proven experience implementing Slowly Changing Dimensions (SCD Types 1, 2, and 3) using Delta Lake MERGE operations and Databricks SQL within a unified lakehouse model.
  • Experience running and optimizing cloud data platforms on Databricks, including cluster configuration, autoscaling policies, job scheduling via Databricks Workflows, and adherence to daily runbook SLAs through proactive monitoring and alerting.
  • Strong experience with Git-based version control integrated into Databricks (Databricks Repos / Git folders) and project management tools such as Jira, operating within Agile/Kanban delivery frameworks.
  • Strong experience with modern data architecture principles, including Unity Catalog for data governance, Delta Sharing, and cloud-native lakehouse design patterns on AWS with Databricks.
  • Ability to translate business requirements into technical designs and deliver production‑grade data solutions within Databricks, from initial scoping through deployment.
  • Design and develop parallel and distributed ETL/ELT pipelines using Apache Spark (PySpark/Scala) on Databricks, applying partitioning, caching, and broadcast join strategies for optimal resource efficiency and throughput.
  • Understand data mapping and transformation requirements and implement them using Databricks-native constructs including Spark transformations (aggregations, joins, unions, window functions, lookups, and pivot/unpivot operations) and Delta Live Tables (DLT) for declarative pipeline development.
  • Develop and maintain Databricks Workflows and job orchestration logic (including dependency management, retry policies, and alerting), replacing traditional shell-based wrapper patterns with cloud‑native, maintainable pipeline automation.
  • Proven experience designing and building integrations that support standard data modeling constructs - fact tables, dimension tables, star and snowflake schemas, and aggregations - implemented as Delta tables within Unity Catalog.
  • Ability to provide end‑to‑end technical guidance across the full software development life cycle, from requirements gathering and architecture design through implementation, testing, and production deployment on Databricks.
  • Ability to produce high‑quality solution design documentation, including data flow diagrams, pipeline architecture specs, and Unity Catalog data asset definitions, ensuring clarity for both technical and business stakeholders.
What You Bring
  • Deep expertise in modern data lakehouse architecture, including Delta Lake, medallion design patterns, Unity Catalog governance, and the transition from traditional data warehousing to cloud‑native lakehouse solutions on Databricks.
  • 5+ years of experience in Data Engineering, with a focus on the following tools and technologies:
    • Databricks - Delta Live Tables (DLT), Databricks Workflows, Unity Catalog, Delta Lake (MERGE, OPTIMIZE, VACUUM, Z-ordering), Databricks SQL, and MLflow
    • AWS services - S3, Redshift, Glue, Lambda, EMR, Athena (with Iceberg), Step Functions, and IAM - integrated with Databricks as the primary compute and transformation layer
    • Scripting and programming languages - Python (PySpark), Scala (Spark), or SQL as primary languages for pipeline development and data transformation within Databricks
  • 3+ years of experience working with cloud platforms - primarily AWS - architecting and operating Databricks environments including workspace configuration, cluster policies, instance profiles, and cost optimization strategies.
  • 1+ years of experience with workflow automation and pipeline orchestration using Databricks Workflows, Apache Airflow (with the Databricks provider), or equivalent cloud-native schedulers, replacing traditional Unix shell scripting with scalable, observable pipeline management.
Preferred Experience & Skills:
  • Experience with Publication and Education domain.
  • Prior experience or familiarity with Tableau/Alteryx.
Why work for us?

The work you do at McGraw Hill will be work that matters. We are collectively building experiences that will help shape the future of education. Play your part and experience a sense of fulfilment that will inspire you to even greater heights.

The pay range for this position is between $135,000 - $160,000 annually, however, base pay offered may vary depending on job-related knowledge, skills, experience, and location. An annual bonus plan may be provided as part of the compensation package, in addition to a full range of medical and/or other benefits, depending on the position offered. Click here to learn more about our benefit offerings.

McGraw Hill recruiters always use a "@mheducation.com" email address and/or from our Applicant Tracking System, iCIMS. Any variation of this email domain should be considered suspicious. Additionally, McGraw Hill recruiters and authorized representatives will never request sensitive information in email.

McGraw Hill uses an automated employment decision tool (AEDT) to assist in the screening process by recommending candidates with "like skills" based on resume and job data. To request an alternative screening process, please select "Opt-Out" when asked to "Consent to use of Automated Employment Decision Tools" during the application.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr Data Engineer
Sr Data Engineer

McGraw Hill • United States

Remote
USD 135,000 - 160,000
Sr Software Engineer
Sr Software Engineer

McGraw Hill • United States

Remote
USD 124,000 - 155,000
Sr Software Engineer, Backend
Sr Software Engineer, Backend

Mcgraw Hill • New York (NY)

On-site
USD 110,000 - 160,000
Director, Engineering, Applied AI
Director, Engineering, Applied AI

Mcgraw Hill • Columbus (OH)

Remote
USD 175,000 - 220,000
Sr Software Engineer, Backend
Sr Software Engineer, Backend

Mcgraw Hill • United States

Remote
USD 110,000 - 160,000
Lead Data Scientist
Lead Data Scientist

Mheducation • United States

On-site
USD 136,000 - 190,000
Annual bonus plan
Full range of medical benefits
Director, Engineering - AI Enablement
Director, Engineering - AI Enablement

McGraw Hill • United States

Remote
USD 175,000 - 220,000
Annual bonus
Medical benefits
Benefits package
Sr Software Engineer, Backend
Sr Software Engineer, Backend

Mheducation • United States

On-site
USD 110,000 - 160,000
Sr Software Engineer, Backend
Sr Software Engineer, Backend

Mcgraw Hill • Columbus (OH)

On-site
USD 110,000 - 160,000
Senior Data Engineer — Lakehouse & Databricks on AWS
Senior Data Engineer — Lakehouse & Databricks on AWS

Mheducation • United States

On-site
USD 135,000 - 160,000
Annual bonus
Medical benefits