Data Engineer

GM Financial

Irving (TX)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

401K matching
Bonding leave for new parents (12+ wks
Tuition assistance
Training
GM employee auto discount
Community service pay
Nine company holidays

Job summary

GM Financial is expanding its data engineering team to design and deploy scalable data processing pipelines in cloud environments. The role focuses on batch and streaming transformations to support data science and analytics, while collaborating with scientists, architects, and IT partners to deliver features and datasets for ML training and model execution.

Candidates should have hands-on experience with Hadoop, Spark, Kafka and cloud technologies, plus strong SQL and Python skills.

Qualifications

  • Experience with processing large data sets using Hadoop, HDFS, Spark, Kafka, Flume or similar distributed systems.
  • Experience with ingesting various source data formats such as JSON, Parquet, SequenceFile, Cloud Databases, MQ, Relational Databases such as Oracle.
  • Experience with Cloud technologies (Azure, AWS, GCP) and native toolsets such as Azure ARM Templates, Hashicorp Terraform, AWS Cloud Formation.
  • Understanding of cloud computing technologies, business drivers and emerging computing trends.
  • Thorough understanding of Hybrid Cloud Computing: virtualization technologies, IaaS, PaaS and SaaS, and current landscape.
  • Knowledge of Object Storage technologies (Data Lake Storage Gen2, S3, Minio, Ceph, ADLS).
  • Experience with containerization (Docker, Kubernetes, Spark on Kubernetes, Spark Operator).
  • Working knowledge of Agile development / SAFe, Scrum, ALM.
  • Strong background with source control (GIT/Subversion); Build Systems (Maven/Gradle/Webpack); Code Quality (Sonar); Artifact Repos (Artifactory); CI/CD (Azure DevOps).
  • Experience with NoSQL stores such as CosmosDB, MongoDB, Cassandra, Redis, Riak or NoSQL search like MarkLogic/Lily.

Responsibilities

  • Code, test, deploy, orchestrate, monitor, document and troubleshoot cloud-based data engineering processing and automation per best practices and security standards.
  • Collaborate with data scientists, data architects, ETL developers, and business partners to extract features from data sources.
  • Evaluate, research, and experiment with batch and streaming data engineering technologies and assess business impact.
  • Showcase capabilities of emerging technologies and enable adoption across teams.
  • Contribute to defining and refining data engineering processes and procedures.
  • Educate ETL developers on cloud-based initiatives to enable transition to data engineering.

Skills

Distributed data processing
Data ingestion
Python
SQL
Spark
Hadoop
Kafka
Cloud computing
REST APIs

Education

Bachelor’s Degree in related field

Tools

Hadoop
Spark
Kafka
Terraform
Azure
AWS
GCP
Docker
Kubernetes
Azure DevOps

Job description

We are expanding our efforts into complementary data technologies for decision support in areas of ingesting and processing large data sets including data commonly referred to as semi-structured or unstructured data. Our interests are in enabling data science and search-based applications on large and low latent data sets in both a batch and streaming context for processing. To that end, this role will engage with team counterparts in exploring and deploying technologies for creating data sets using a combination of batch and streaming transformation processes. These data sets support both off-line and in-line machine learning training and model execution. Other data sets support search engine-based analytics. Exploration and deployment of technologies activities include identifying opportunities that impact business strategy, collaborating on the selection of data solutions software, and contributing to the identification of hardware requirements based on business requirements. Responsibility also includes coding, testing, and documentation of new or modified scalable analytic data systems including automation for deployment and monitoring. This role participates along with team counterparts to develop solutions in an end-to-end framework on a group of core data technologies. Other aspects of the role include developing standards and processes for data engineering projects and cloud initiatives.

JOB DUTIES
  • Code, test, deploy, Orchestrate, monitor, document and troubleshoot cloud-based data engineering processing and associated automation in accordance with best practices and security standards throughout the development lifecycle
  • Work closely with data scientists, data architects, ETL developers, other IT counterparts, and business partners to identify, collect, and format data from external sources, internal systems and the data warehouse and lakehouse to extract features of interest
  • Significantly contribute to evaluation, research, and experimentation efforts with batch and streaming data engineering technologies to keep pace with industry innovation while assessing business impact and viability for use cases associated with efforts in hand
  • Work with data engineering related groups to inform on and showcase capabilities of emerging technologies and to enable the adoption of these new technologies and associated techniques
  • Significantly contribute to the definition and refinement of processes and procedures for the data engineering practice
  • Educate and develop ETL developers on data engineering cloud-based initiatives so as to enable transition to data engineer and practice
Qualifications
What makes you a dream candidate?
  • Experience with processing large data sets using Hadoop, HDFS, Spark, Kafka, Flume or similar distributed systems
  • Experience with ingesting various source data formats such as JSON, Parquet, SequenceFile, Cloud Databases, MQ, Relational Databases such as Oracle
  • Experience with Cloud technologies (such as Azure, AWS, GCP) and native toolsets such as Azure ARM Templates, Hashicorp Terraform, AWS Cloud Formation
  • Understanding of cloud computing technologies, business drivers and emerging computing trends
  • Thorough understanding of Hybrid Cloud Computing: virtualization technologies, Infrastructure as a Service, Platform as a Service and Software as a Service Cloud delivery models and the current competitive landscape
  • Working knowledge of Object Storage technologies to include but not limited to Data Lake Storage Gen2, S3, Minio, Ceph, ADLS etc
  • Experience with containerization to include but not limited to Dockers, Kubernetes, Spark on Kubernetes, Spark Operator
  • Working knowledge of Agile development /SAFe, Scrum and Application Lifecycle Management
  • Strong background with source control management systems (GIT or Subversion); Build Systems (Maven, Gradle, Webpack); Code Quality (Sonar); Artifact Repository Managers (Artifactory), Continuous Integration/ Continuous Deployment (Azure DevOps)
  • Experience with NoSQL data stores such as CosmosDB, MongoDB, Cassandra, Redis, Riak or other technologies that embed NoSQL with search such as MarkLogic or Lily Enterprise
  • Creating and maintaining ETL processes
  • Knowledgeable of best practices in information technology governance and privacy compliance
  • Experience with REST APIs
  • Advanced knowledge of Databricks platform and associated features including workflows, unity catalog, delta live tables, time travel, SQL statement execution API, etc

Understanding of Databricks medallion architecture

Advanced knowledge of programming concepts and languages including SQL and Python/PySpark

Additional Skills
  • Troubleshoot complex problems and works across teams to meet commitments
  • Excellent computer skills and proficiency in digital data collection
  • Ability to work in an Agile/Scrum team environment
  • Strong interpersonal, verbal, and writing skills
  • Digital technology solutions (DMPs, CDPs, Tag Management Platforms, Cross-Device Tracking, SDKs, etc)
  • Knowledge of Real Time-CDP and Journey Analytics solutions
  • Understanding of big data platforms and architectures, data stream processing pipeline/platform, data lake and data lake houses
  • SQL experience: querying data and sharing what insights can be derived
  • Understanding of cloud solutions such as Google Cloud Platform, Microsoft Azure & Amazon AWS cloud architecture & services
  • Understanding of GDPR, privacy & security topics
Experience and Education
  • 2-4 years of hands-on experience with data engineering required
  • 2-4 years of hands-on experience with processing large data sets required
  • 2-4 years of hands-on experience with SQL, data modeling, relational databases and/or no SQL databases required
  • Bachelor’s Degree in related field or equivalent work experience required
What We Offer
  • 401K matching
  • bonding leave for new parents (12 weeks, 100% paid)
  • tuition assistance
  • training
  • GM employee auto discount
  • community service pay
  • nine company holidays
Our Culture

Our team members define and shape our culture — an environment that welcomes innovative ideas, fosters integrity, and creates a sense of community and belonging. Here we do more than work — we thrive.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. Data Engineer
Sr. Data Engineer

Save A Lot • St. Ann (MO)

On-site
USD 110,000 - 160,000
401K match
Paid Time Off
Medical insurance
+6
Sr. Data Engineer
Sr. Data Engineer

Save-A-Lot, Ltd. • Missouri

On-site
USD 100,000 - 140,000
401K match up to 4%
Paid Time Off
Medical Insurance options including FS
+5
Data Engineer
Data Engineer

Compunnel, Inc. • Norfolk (VA)

On-site
USD 110,000 - 150,000
Data Engineer
Data Engineer

DataJobs • New York (NY)

On-site
USD 70,000 - 95,000
Paid time off
Medical, dental, and vision coverage
Retirement plans (401(k))
+2
Data Engineer
Data Engineer

Prodigy Resources • Denver (CO)

On-site
USD 110,000 - 170,000
Software Engineer
Software Engineer

The Coca-Cola Company • Atlanta (GA)

On-site
USD 110,000 - 160,000
Data Engineer 1
Data Engineer 1

DAIKIN COMFORT TECHNOLOGIES MFG INC • Waller (TX)

On-site
USD 90,000 - 130,000
Data Engineer
Data Engineer

CEI • Virginia (MN)

Hybrid
USD 100,000 - 120,000
Data Engineer
Data Engineer

Northbound Executive Search • Austin (TX)

On-site
USD 80,000 - 120,000
Data Engineer
Data Engineer

ReturnPro • Town of Florida (NY)

On-site
USD 110,000 - 150,000