Real-Time Big Data Engineer — Spark, Hive & ML

Remote Jobs

United States

Remote

USD 120,000 - 160,000

Full time

47 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Remote Jobs is seeking a skilled Big Data Engineer to design, develop, and deploy scalable data platforms in the Hadoop ecosystem (HBase, Hive, Kudu, Spark). You will build and optimize data pipelines for storage, analytics, and real-time processing.

You will architect HBase schemas, develop Hive queries, implement Spark Streaming/SQL pipelines, and collaborate with data scientists to productionize ML models.

Responsibilities

  • Design, develop, and implement scalable and distributed big data solutions using Hadoop ecosystem technologies such as HBase, Hive, Kudu, and Spark.
  • Architect HBase schemas and data models to accommodate evolving business requirements and ensure optimal performance for data storage and retrieval.
  • Develop complex Hive queries and data processing pipelines to transform raw data into structured formats suitable for analysis and reporting.
  • Implement data ingestion pipelines using Spark Streaming and Spark SQL for real-time processing of streaming data sources, ensuring high throughput and low latency.
  • Optimize Spark applications for performance and resource utilization, including tuning RDD transformations, optimizing data partitioning strategies, and leveraging in-memory caching.
  • Utilize advanced features of Spark MLlib for machine learning tasks such as classification, regression, clustering, and collaborative filtering.
  • Design and deploy Kudu tables for fast analytical queries and real-time analytics, leveraging Kudu's unique combination of fast analytics and fast data ingestion.
  • Collaborate with data scientists to integrate machine learning models into Spark workflows and productionize them for real-time predictions and analytics.
  • Troubleshoot performance bottlenecks, data quality issues, and system failures in big data applications and infrastructure, and implement solutions to address them.
  • Stay abreast of emerging technologies and best practices in big data processing and analytics, and evaluate their potential impact on our architecture and solutions.

Tools

Hadoop
HBase
Hive
Kudu
Spark
Spark Streaming
Spark SQL
Spark MLlib

Job description

Remote Jobs is seeking a skilled Big Data Engineer to design, develop, and deploy scalable data platforms in the Hadoop ecosystem (HBase, Hive, Kudu, Spark). You will build and optimize data pipelines for storage, analytics, and real-time processing.

You will architect HBase schemas, develop Hive queries, implement Spark Streaming/SQL pipelines, and collaborate with data scientists to productionize ML models.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Data Platform Engineer: Big Data & Spark
Remote Data Platform Engineer: Big Data & Spark

Bright Vision Technologies • Carmel (IN)

Remote
USD 100,000 - 150,000
100% Remote
Remote Big Data Engineer - Spark, Hadoop & Analytics
Remote Big Data Engineer - Spark, Hadoop & Analytics

Bright Vision Technologies • Elk Grove (CA)

Remote
USD 90,000 - 110,000
Remote Big Data Engineer – Spark & Hadoop Expert
Remote Big Data Engineer – Spark & Hadoop Expert

Bright Vision Technologies • Reston (VA)

On-site
USD 100,000 - 150,000
Big Data Developer
Big Data Developer

Remote Jobs • United States

Remote
USD 120,000 - 160,000
Senior Spark Developer — Remote, Big Data & Cloud
Senior Spark Developer — Remote, Big Data & Cloud

Bright Vision Technologies • Austin (TX)

Remote
USD 125,000 - 185,000
Remote Data Platform Engineer - Big Data & Spark Expert
Remote Data Platform Engineer - Big Data & Spark Expert

Bright Vision Technologies • Columbus (OH), New Albany (OH), Hilliard (OH)

Remote
USD 100,000 - 150,000
Senior Data Engineer - Big Data & ML Infra (Remote)
Senior Data Engineer - Big Data & ML Infra (Remote)

ZipRecruiter • Los Angeles (CA)

Hybrid
USD 130,000 - 175,000
Competitive compensation
Exceptional benefits package
Flexible Vacation & Paid Time Off
+1
Lead Big Data Engineer — Spark, Hadoop, Hybrid
Lead Big Data Engineer — Spark, Hadoop, Hybrid

EXL • Dallas (TX)

Hybrid
USD 120,000 - 180,000
Senior Big Data Engineer - PySpark, Spark, Hive
Senior Big Data Engineer - PySpark, Spark, Hive

Realign Llc • Charlotte (NC)

On-site
USD 120,000 - 180,000
Remote Hadoop Big Data Engineer
Remote Hadoop Big Data Engineer

United States Digital Space LLC • United States

Remote
USD 62,000 - 103,000