Big Data Engineer

Ex

Pittsburgh (Allegheny County)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Ex is seeking a Big Data Engineer with 7-10 years of hands-on experience to design, build, and maintain scalable data pipelines in an on-premises Big Data environment. You will own architecture decisions, set technical direction, and mentor other engineers on the team.

The role collaborates with data analysts and data scientists to translate data requirements into robust designs, while staying current with AI/ML capabilities and best practices.

Qualifications

  • 7–10 years in data engineering with leadership experience.
  • Strong Python and Spark production experience.
  • Hands-on in Hadoop ecosystem and on-premises environments.
  • Ability to design and guide data architecture decisions.
  • Exposure to AI/ML concepts and opportunities.

Responsibilities

  • Lead design, development, and maintenance of scalable data pipelines.
  • Own architectural decisions and set technical standards for the team.
  • Mentor data engineers through code reviews and coaching.
  • Build data workflows using Python, Spark, and Hadoop ecosystem.
  • Work with Hive, HDFS, and Impala to query large datasets.
  • Manage batch scheduling with CA7 or Control-M.
  • Ensure data quality, integrity, and performance across platforms.
  • Collaborate with analysts and data scientists to translate requirements.

Skills

Python
Apache Spark
Hadoop ecosystem
CA7/Control-M
SQL
Data architecture
Mentoring
AI/ML exposure
On-premises

Tools

CA7
Control-M

Job description

Job Description: Big Data Engineer Summary

We are looking for an experienced Big Data Engineer with 7-10 years of hands-on experience to design, build, and maintain scalable data pipelines and processing systems in an on-premises Big Data environment. Beyond strong individual contribution, the ideal candidate will own architecture and design decisions, set technical direction, and mentor and support other developers on the team. The role works closely with cross-functional teams to deliver reliable, high-quality data solutions that support business and analytics needs.

Roles & Responsibilities
  • Lead the design, development, and maintenance of robust, scalable data pipelines for ingestion, transformation, and processing of large datasets in an on-premises environment.
  • Own architectural and design decisions for data solutions, evaluating trade-offs and defining technical standards for the team.
  • Mentor, guide, and support other data engineers through code reviews, design reviews, technical coaching, and hands-on problem-solving.
  • Build and optimize data workflows using Python, Spark, and the Hadoop ecosystem.
  • Work extensively with Hadoop ecosystem components (Hive, HDFS, Impala) to manage and query large-scale data.
  • Manage and optimize batch scheduling and job orchestration using enterprise schedulers such as CA7 or Control-M.
  • Ensure data quality, integrity, and performance across data platforms.
  • Collaborate with data analysts, data scientists, and business stakeholders to translate data requirements into sound technical designs.
  • Troubleshoot and resolve complex issues in data pipelines and production environments, acting as an escalation point for the team.
  • Champion best practices for coding standards, version control, testing, and documentation.
  • Stay current with emerging technologies, particularly AI/ML capabilities, and identify opportunities to apply them to data engineering workflows.
Technical Skills Must Have
  • 7-10 years of overall experience in data engineering, with a proven track record in technical leadership (design ownership, mentoring, guiding development teams).
  • Python – strong hands-on development experience building production-grade data solutions.
  • Big Data / Hadoop ecosystem (Hadoop, Hive, Impala, HDFS) – deep, hands-on experience in on-premises environments.
  • Apache Spark – solid experience developing and tuning large-scale distributed data processing jobs.
  • Job scheduling / orchestration – hands-on experience with CA7 or Control-M (or comparable enterprise schedulers).
  • Strong understanding of data structures, ETL processes, and SQL.
  • Extensive experience with large-scale data processing and distributed systems.
  • Demonstrated ability to make sound architecture/design decisions and to mentor and support other developers.
  • Exposure to AI/ML concepts or tools, with a strong willingness to learn and grow in this space.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Big Data Lead
Big Data Lead

Veriipro • United States

On-site
USD 180,000 - 240,000
Big Data Consultant
Big Data Consultant

Unisys • Rockville (MD)

On-site
USD 120,000 - 170,000
Big Data Engineer
Big Data Engineer

TechDigital Group • Jersey City (NJ)

On-site
USD 90,000 - 150,000
Big Data Engineer
Big Data Engineer

Select Minds LLC • Lansing (MI)

On-site
USD 100,000 - 160,000
Onsite
Competitive salary
Opportunity for advancement
Big Data Consultant
Big Data Consultant

EXL • Dallas (TX)

Hybrid
USD 120,000 - 180,000
Senior Data Engineer – 8+ Years Experience
Senior Data Engineer – 8+ Years Experience

Hudson Manpower • New York (NY)

On-site
USD 140,000 - 190,000
Big Data Developer
Big Data Developer

Unisys • Rockville (MD)

Hybrid
USD 120,000 - 180,000
Big Data Engineer
Big Data Engineer

Princeton IT Services, Inc • Mount Laurel Township (NJ)

On-site
USD 120,000 - 150,000
Data Engineer
Data Engineer

Infinite Computer Solutions • Town of Texas (WI)

On-site
USD 120,000 - 170,000
Senior Data Engineer
Senior Data Engineer

Novatalent • United States

On-site
USD 110,000 - 160,000