Big Data Engineer

Compunnel, Inc.

Tysons (VA)

On-site

USD 150,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Compunnel, Inc. is seeking a Big Data Engineer to design, develop, and optimize large-scale data processing systems at petabyte scale.

You will build scalable data pipelines using Spark, PySpark, Hadoop, Hive, Trino, and AWS services, collaborating with data scientists and analysts to translate business needs into robust solutions. The role emphasizes performance tuning, reliability, automation, and data quality, with opportunities to apply AI tooling and modern development practices in a

Qualifications

  • Bachelor's degree with 5+ years of related experience.
  • Strong hands-on experience with Big Data technologies including Hadoop and Apache Spark.
  • Experience with Spark development, internals, and performance tuning.
  • Experience with AWS services such as S3, EMR, Glue, Lambda, and Athena.
  • Experience developing scalable data processing solutions using Python, Scala, or Java.
  • Experience with test automation, build automation, and configuration management practices.

Responsibilities

  • Design, develop, and maintain large-scale data processing pipelines using Hadoop, Spark, Python, and Scala.
  • Implement scalable and reliable data ingestion, storage, transformation, and analysis solutions.
  • Develop and optimize data pipelines designed to process massive data volumes at petabyte scale.
  • Optimize Spark applications through partitioning, caching, broadcast joins, and other performance tuning techniques.
  • Analyze and troubleshoot Spark architecture components including executors, tasks, stages, and DAGs.
  • Work with AWS services including S3, EMR, Glue, Lambda, and Athena to build and support cloud-based data solutions.
  • Develop scalable Python, PySpark, or Scala-based data processing applications.
  • Collaborate with cross-functional teams to understand business requirements and translate them into technical solutions.
  • Develop automated testing frameworks and implement continuous testing for data quality assurance.

Skills

Big Data tech
Spark
Python
Scala
Java
SQL
AWS
Agile
CI/CD
Testing
AI tools

Education

Bachelor's degree in Computer Science / related field

Tools

Hadoop
Apache Spark
Hive
Trino
Python
Scala

Job description

The Big Data Engineer will design, develop, and optimize large-scale data processing systems capable of operating at petabyte scale. The role focuses on building scalable and reliable data pipelines using Spark, Python/PySpark, Hadoop, Hive, Trino, and AWS services. The engineer will work closely with cross-functional teams, data scientists, and analysts to translate business requirements into technical solutions while improving data processing performance, reliability, automation, and quality.

Key Responsibilities
  • Design, develop, and maintain large-scale data processing pipelines using technologies such as Hadoop, Spark, Python, and Scala.
  • Implement scalable and reliable data ingestion, storage, transformation, and analysis solutions.
  • Develop and optimize data pipelines designed to process massive data volumes at petabyte scale.
  • Optimize Spark applications through partitioning, caching, broadcast joins, and other performance tuning techniques.
  • Analyze and troubleshoot Spark architecture components including executors, tasks, stages, and DAGs.
  • Troubleshoot data processing failures related to resource limitations, data quality, scalability, and data skew.
  • Develop solutions to identify and mitigate data skew and other distributed processing challenges.
  • Work with AWS services including S3, EMR, Glue, Lambda, and Athena to build and support cloud-based data solutions.
  • Develop scalable Python, PySpark, or Scala-based data processing applications.
  • Collaborate with cross-functional teams to understand business requirements and translate them into technical solutions.
  • Optimize existing data pipelines for performance, scalability, and reliability.
  • Develop automated testing frameworks and implement continuous testing for data quality assurance.
  • Write and maintain automated unit, integration, and end-to-end tests.
  • Conduct unit, integration, and system testing to ensure the robustness and accuracy of data pipelines.
  • Monitor and troubleshoot production data pipelines and resolve operational issues.
  • Work with data scientists and analysts to support data-driven decision-making.
  • Apply AI development tools and prompt engineering techniques to improve software development and data engineering workflows.
  • Evaluate emerging Big Data and AI technologies and contribute to continuous improvement of data architecture and engineering practices.
Required Qualifications
  • Bachelor's degree in Computer Science, Information Systems, or a related discipline with 5+ years of related experience, or equivalent training and/or work experience.
  • Strong software development background with experience delivering enterprise-quality solutions.
  • Strong hands‑on experience with Big Data technologies including Hadoop and Apache Spark.
  • Experience working with Spark development, internals, and performance tuning.
  • Strong SQL skills, including window functions, multi-table joins, aggregations, and complex queries.
  • Experience working with massive data volumes at petabyte scale.
  • Strong understanding of distributed data processing and troubleshooting.
  • Experience with Hadoop, Hive, and Trino.
  • Hands‑on experience with AWS services such as S3, EMR, Glue, Lambda, and Athena.
  • Experience developing scalable and maintainable data processing solutions using Python, Scala, or Java.
  • Experience with object-oriented programming and database technologies and concepts.
  • Experience developing enterprise-quality solutions in iterative or Agile environments.
  • Experience with test automation, build automation, and configuration management practices.
  • Experience managing and troubleshooting production data pipelines or ETL systems.
  • Strong written and verbal technical communication skills.
  • Strong analytical, problem‑solving, organizational, and prioritization skills.
  • Ability to learn new technologies and develop proficiency rapidly in a fast‑paced environment.
  • Hands‑on experience with AI development tools such as GitHub Copilot, Amazon Q Developer, ChatGPT, Claude, or comparable tools.
  • Experience with Agile methodologies such as Scrum and Kanban.
Preferred Qualifications
  • Master's degree in Computer Science, Information Systems, or a related discipline.
  • Experience in the Financial Services industry.
  • Experience with CI/CD practices.
  • Experience writing and maintaining comprehensive test cases.
  • Experience with AI workflow design and redesigning development processes to leverage AI capabilities.
  • Experience with prompt engineering for AI coding assistants and analysis tools.
  • Experience interpreting AI-generated insights and applying them to development or data engineering improvements.
  • Experience supporting organizational adoption of AI-enabled development workflows.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Big Data Developer
Big Data Developer

Unisys • Rockville (MD)

On-site
USD 150,000 - 190,000
Bigdata Engineer
Bigdata Engineer

Disys - Oak Brook • Tampa (FL)

On-site
USD 90,000 - 120,000
Big Data Platform Engineer
Big Data Platform Engineer

Compunnel, Inc. • Rockville (MD)

On-site
USD 140,000 - 190,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
Senior Data Engineer
Senior Data Engineer

Selby Jennings • New York (NY)

On-site
USD 150,000 - 185,000
Data Engineer
Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 125,000 - 140,000
Data Engineer
Data Engineer

Compunnel, Inc. • San Francisco (CA)

On-site
USD 100,000 - 130,000
Data Engineer
Data Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 90,000 - 120,000
Data Engineer
Data Engineer

VTG Defense • McLean (VA)

On-site
USD 100,000 - 130,000
Data Engineer
Data Engineer

The Judge Group • New York (NY)

On-site
USD 150,000 - 190,000