Big Data Platform Engineer

Compunnel, Inc.

Rockville (MD)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Compunnel, Inc. seeks a seasoned Big Data Platform Engineer to design, develop, and optimize large-scale data processing platforms in a cloud-native environment. You will build scalable data pipelines and architect Kubernetes infra to run Spark workloads on AWS, including EMR on EKS.

The role requires deep expertise in Spark, Kubernetes, Hadoop, Hive, Trino, and modern DevOps practices, with a focus on performance, cost efficiency, and reliability.

Qualifications

  • Design and develop enterprise Big Data solutions on Kubernetes with Spark.
  • Experience running Spark on Kubernetes and EMR on EKS.
  • Strong AWS services and security in cloud environments.
  • Proficient in Python or Scala for data processing.
  • Experience with Hadoop, Hive, and Trino ecosystems.
  • Proven ability to optimize data pipelines for scale and cost.

Responsibilities

  • Design, develop, and maintain large-scale data pipelines using Spark, Hadoop, Python, and Scala.
  • Architect, deploy, and optimize containerized big data workloads on EMR on EKS.
  • Build Kubernetes infrastructure to support Spark apps and data platforms.
  • Develop scalable data ingestion, storage, transformation, and analytics.
  • Monitor, troubleshoot, and resolve production data pipelines and platforms.
  • Create automated tests to ensure data quality and reliability.
  • Collaborate with architects, data scientists, and engineers on enterprise solutions.
  • Research emerging Big Data and cloud AI technologies to improve capabilities.

Skills

Kubernetes
Apache Spark
Python
Scala
AWS
Data pipelines
SQL
Performance tuning

Education

Bachelor's degree in CS/IS
Master's degree (preferred)

Tools

Docker
Terraform
CloudFormation
GitHub Actions
ArgoCD
Istio
Prometheus
Grafana

Job description

We are seeking an experienced Big Data Platform Engineer to design, develop, and optimize large-scale data processing platforms and cloud-native data solutions. The ideal candidate will have deep expertise in Apache Spark, Kubernetes, AWS, and distributed data processing technologies. This role focuses on building scalable data pipelines, architecting Kubernetes infrastructure, optimizing big data workloads, and supporting enterprise data platforms through modern software engineering and DevOps practices.

Key Responsibilities
  • Design, develop, and maintain large-scale data processing pipelines using Apache Spark, Hadoop, Python, and Scala.
  • Architect, deploy, and optimize containerized big data workloads using Amazon EMR on EKS.
  • Design and build Kubernetes infrastructure to support scalable Spark applications and enterprise data platforms.
  • Develop scalable solutions for data ingestion, storage, transformation, and analytics.
  • Optimize existing data pipelines for performance, scalability, reliability, and cost efficiency.
  • Monitor, troubleshoot, and resolve production issues affecting data pipelines and big data platforms.
  • Develop automated testing frameworks and implement continuous testing to ensure data quality and platform reliability.
  • Create and maintain unit, integration, and end-to-end tests for data processing applications.
  • Manage Kubernetes clusters, pods, deployments, services, namespaces, ConfigMaps, Secrets, and storage resources.
  • Collaborate with architects, data scientists, analysts, and engineering teams to deliver enterprise data solutions.
  • Research and evaluate emerging Big Data, cloud, and AI technologies to improve platform capabilities.
  • Participate in architecture discussions and contribute to technical design decisions.
Required Qualifications
  • Bachelor's degree in Computer Science, Information Systems, or a related field, or equivalent professional experience.
  • Minimum of 5 years of experience designing and developing enterprise Big Data solutions.
  • Strong experience building Kubernetes infrastructure for enterprise applications.
  • Hands‑on experience with Kubernetes architecture, including pods, services, deployments, namespaces, ConfigMaps, Secrets, networking, storage, and security.
  • Experience running Apache Spark workloads on Kubernetes, including Amazon EMR on EKS.
  • Strong experience with Apache Spark, including Spark architecture, performance tuning, partitioning, caching, broadcast joins, DAG execution, executors, and resource optimization.
  • Experience designing and supporting large‑scale data pipelines processing massive datasets.
  • Experience with Hadoop, Spark, Hive, and Trino.
  • Experience troubleshooting data skew, resource constraints, scalability issues, and Spark job failures.
  • Strong experience with AWS services, including Amazon S3, EMR, EMR on EKS, Glue, Lambda, Athena, Amazon EKS, CloudWatch, and CloudTrail.
  • Experience with AWS IAM Roles for Service Accounts (IRSA), VPC networking, subnets, and security groups.
  • Experience with Kubernetes resource management, scheduling, auto‑scaling, Helm charts, kubectl, and YAML manifests.
  • Experience integrating Spark with Kubernetes operators and dynamic resource allocation.
  • Strong programming experience using Python or Scala.
  • Experience writing clean, modular, scalable, and high‑performance code.
  • Strong understanding of functional programming concepts, concurrency, memory management, and collections.
  • Strong SQL skills, including window functions, complex joins, aggregations, and query optimization.
  • Experience using AI development tools such as GitHub Copilot, Amazon Q Developer, ChatGPT, or Claude.
  • Experience applying AI tools to improve software development workflows and engineering productivity.
  • Strong analytical, troubleshooting, communication, and problem‑solving skills.
  • Experience working in Agile software development environments.
Preferred Qualifications
  • Experience managing enterprise ETL and production data pipeline platforms.
  • Experience with CI/CD tools such as Jenkins, GitLab CI, GitHub Actions, or ArgoCD.
  • Experience with Infrastructure as Code (Terraform or AWS CloudFormation).
  • Experience with Docker and container image optimization.
  • Experience with service mesh technologies such as Istio or Linkerd.
  • Experience with monitoring and observability tools including Prometheus, Grafana, or the ELK Stack.
  • AWS certifications such as AWS Certified AI Practitioner, AWS Certified Solutions Architect, or AWS Certified Data Analytics/Specialty.
  • Kubernetes certifications such as Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD).
  • Experience with GitOps deployment practices.
  • Master's degree in Computer Science, Information Systems, or a related field.
  • Experience working within the Financial Services industry.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Data Engineer, Data Platform
Sr. Data Engineer, Data Platform

Mirion Technologies • United States

On-site
USD 120,000 - 160,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
Senior Data Engineer
Senior Data Engineer

Peyton Resource Group • Houston (TX)

On-site
USD 120,000 - 150,000
Bigdata Engineer
Bigdata Engineer

Disys - Oak Brook • Tampa (FL)

On-site
USD 90,000 - 120,000
Big Data Engineer - RQ308
Big Data Engineer - RQ308

Experis • McLean (VA)

On-site
USD 110,000 - 150,000
Big Data Lead
Big Data Lead

Veriipro • United States

On-site
USD 180,000 - 240,000
Data Engineer
Data Engineer

The Judge Group • New York (NY)

On-site
USD 150,000 - 190,000
Data Engineer
Data Engineer

Infinite Computer Solutions • Town of Texas (WI)

On-site
USD 120,000 - 170,000
Data Platform Engineer
Data Platform Engineer

Compunnel, Inc. • Pennsylvania

On-site
USD 100,000 - 130,000
Big Data Consultant
Big Data Consultant

Unisys • Rockville (MD)

On-site
USD 120,000 - 170,000