Big Data Engineer

TechDigital Group

Jersey City (NJ)

On-site

USD 90,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

An innovative firm is looking for a talented Big Data Engineer to join their dynamic team. In this role, you will design and implement scalable data pipelines, develop machine learning models, and integrate complex big data systems. Your expertise in tools like Apache Spark, AWS Glue, and various databases will be crucial in optimizing data workflows and ensuring seamless data processing. This position offers the opportunity to work with cutting-edge technologies in cloud environments, contributing to impactful projects that drive data-driven decision-making. If you're passionate about big data and eager to make a difference, this is the perfect opportunity for you.

Qualifications

  • Expertise in building scalable ETL pipelines and big data systems.
  • Hands-on experience with machine learning frameworks and cloud platforms.

Responsibilities

  • Build scalable ETL pipelines using Apache Spark and AWS Glue.
  • Process large datasets with Hadoop and Kafka for real-time data streaming.

Skills

Apache Spark
AWS Glue
Google Dataflow
Hadoop
Kafka
MySQL
PostgreSQL
MongoDB
Cassandra
Machine Learning

Tools

Talend
Apache Flink
TensorFlow
PyTorch
Scikit-learn
Git

Job description

Mandatory Skills:
Apache Spark, Hive, Kafka, Amazon Glue, Google Dataflow, Talend MDM, Hadoop, Presto, Strong experience with MySQL, PostgreSQL, MongoDB, Cassandra.

Role: Big Data Engineer

Job Overview:
We're seeking a highly skilled Data Engineer, Big Data Engineer to build scalable data pipelines, develop ML models, and integrate big data systems. You'll work with structured, semi-structured, and unstructured data, focusing on optimizing data systems, building ETL pipelines, and deploying AI models in cloud environments.

Key Responsibilities:

  1. Data Ingestion: Build scalable ETL pipelines using Apache Spark, Talend, AWS Glue, Google Dataflow, Apache NiFi. Ingest data from APIs, file systems, and databases.
  2. Data Transformation/Validation: Use Pandas, Apache Beam, and Dask for data cleaning, transformation, and validation. Automate data quality checks with Pytest, Unittest.
  3. Big Data Systems: Process large datasets with Hadoop, Kafka, Apache Flink, Apache Hive. Stream real-time data using Kafka, Google Cloud PubSub.
  4. Task Queues: Manage asynchronous processing with Celery, RQ, RabbitMQ, or Kafka. Implement retry mechanisms and track task status.
  5. Scalability: Optimize for performance with distributed processing (Spark, Flink), parallelization (joblib), and data partitioning.
  6. Cloud Storage: Work with AWS, Azure, GCP, Databricks. Store and manage data with S3, BigQuery, Redshift, Synapse Analytics, and HDFS.

Required Skills:

  1. ETL Data Processing: Expertise in Apache Spark, AWS Glue, Google Dataflow, Talend.
  2. Big Data Tools: Proficient with Hadoop, Kafka, Apache Flink, Hive, Presto.
  3. Databases: Strong experience with MySQL, PostgreSQL, MongoDB, Cassandra.
  4. Machine Learning: Hands-on with TensorFlow, PyTorch, Scikit-learn, XGBoost.
  5. Cloud Platforms: Experience with AWS, Azure, GCP, Databricks.
  6. Task Management: Familiar with Celery, RQ, RabbitMQ, Kafka.
  7. Version Control: Git for source code management.

Desirable Skills:

  1. Real-time Data Processing: Experience with Apache Pulsar, Google Cloud PubSub.
  2. Data Warehousing: Familiarity with Redshift, BigQuery, Synapse Analytics.
  3. Scalability Optimization: Knowledge of load balancing (NGINX, HAProxy) and parallel processing.
  4. Data Governance: Use of MLflow, DVC, or other tools for model and data versioning.

Tools Technologies:

  1. ETL: Apache Spark, Talend, AWS Glue, Google Dataflow.
  2. Big Data: Hadoop, Kafka, Apache Flink, Presto.
  3. Databases: MySQL, PostgreSQL, MongoDB, Cassandra.
  4. Cloud: AWS, GCP, Azure, Databricks.
  5. Storage: S3, BigQuery, Redshift, Synapse Analytics, HDFS.
  6. Version Control: Git.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Big Data Engineer
Big Data Engineer

Princeton IT Services, Inc • Mount Laurel Township (NJ)

On-site
USD 120,000 - 150,000
Big Data Lead
Big Data Lead

Veriipro • United States

On-site
USD 180,000 - 240,000
Big Data Engineer
Big Data Engineer

Ex • Pittsburgh

On-site
USD 120,000 - 180,000
Big Data Consultant
Big Data Consultant

Unisys • Rockville (MD)

On-site
USD 120,000 - 170,000
Data Engineer ETL
Data Engineer ETL

Compunnel, Inc. • Durham (NC)

On-site
USD 100,000 - 130,000
Big Data Engineer - RQ308
Big Data Engineer - RQ308

Experis • McLean (VA)

On-site
USD 110,000 - 150,000
Senior Data Engineer – 8+ Years Experience
Senior Data Engineer – 8+ Years Experience

Hudson Manpower • New York (NY)

On-site
USD 140,000 - 190,000
Bigdata Engineer
Bigdata Engineer

Disys - Oak Brook • Tampa (FL)

On-site
USD 90,000 - 120,000
Senior Data Engineer – 8+ Years Experience
Senior Data Engineer – 8+ Years Experience

Hudson Manpower • New Jersey

On-site
USD 140,000 - 190,000
Data Engineer
Data Engineer

ATC • New York (NY)

On-site
USD 140,000 - 180,000