Big Data Developer

Remote Jobs

United States

Remote

USD 120,000 - 160,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Remote Jobs is seeking a skilled Big Data Engineer to design, develop, and deploy scalable data platforms in the Hadoop ecosystem (HBase, Hive, Kudu, Spark). You will build and optimize data pipelines for storage, analytics, and real-time processing.

You will architect HBase schemas, develop Hive queries, implement Spark Streaming/SQL pipelines, and collaborate with data scientists to productionize ML models.

Responsibilities

  • Design, develop, and implement scalable and distributed big data solutions using Hadoop ecosystem technologies such as HBase, Hive, Kudu, and Spark.
  • Architect HBase schemas and data models to accommodate evolving business requirements and ensure optimal performance for data storage and retrieval.
  • Develop complex Hive queries and data processing pipelines to transform raw data into structured formats suitable for analysis and reporting.
  • Implement data ingestion pipelines using Spark Streaming and Spark SQL for real-time processing of streaming data sources, ensuring high throughput and low latency.
  • Optimize Spark applications for performance and resource utilization, including tuning RDD transformations, optimizing data partitioning strategies, and leveraging in-memory caching.
  • Utilize advanced features of Spark MLlib for machine learning tasks such as classification, regression, clustering, and collaborative filtering.
  • Design and deploy Kudu tables for fast analytical queries and real-time analytics, leveraging Kudu's unique combination of fast analytics and fast data ingestion.
  • Collaborate with data scientists to integrate machine learning models into Spark workflows and productionize them for real-time predictions and analytics.
  • Troubleshoot performance bottlenecks, data quality issues, and system failures in big data applications and infrastructure, and implement solutions to address them.
  • Stay abreast of emerging technologies and best practices in big data processing and analytics, and evaluate their potential impact on our architecture and solutions.

Tools

Hadoop
HBase
Hive
Kudu
Spark
Spark Streaming
Spark SQL
Spark MLlib

Job description

  1. Design, develop, and implement highly scalable and distributed big datasolutions using Hadoop ecosystem technologies such as HBase, Hive, Kudu, andSpark.
  2. Architect HBase schemas and data models to accommodate evolving businessrequirements and ensure optimal performance for data storage and retrievaloperations.
  3. Develop complex Hive queries and data processing pipelines to transform rawdata into structured formats suitable for analysis and reporting.
  4. Implement data ingestion pipelines using Spark Streaming and Spark SQL forreal-time processing of streaming data sources, ensuring high throughput andlow latency.
  5. Optimize Spark applications for performance and resource utilization,including tuning RDD transformations, optimizing data partitioning strategies,and leveraging in-memory caching.
  6. Utilize advanced features of Spark MLlib for machine learning tasks such asclassification, regression, clustering, and collaborative filtering.
  7. Design and deploy Kudu tables for fast analytical queries and real-timeanalytics, leveraging Kudus unique combination of fast analytics and fast dataingestion.
  8. Collaborate with data scientists to integrate machine learning models intoSpark workflows and productionize them for real-time predictions and analytics.
  9. Troubleshoot performance bottlenecks, data quality issues, and systemfailures in big data applications and infrastructure, and implement solutionsto address them.
  10. Stay abreast of emerging technologies and best practices in big data processingand analytics, and evaluate their potential impact on our architecture andsolutions.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Big Data Developer
Big Data Developer

eNcloud • Boston (MA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Big Data Developer
Big Data Developer

Unisys • Rockville (MD)

On-site
USD 150,000 - 190,000
Real-Time Big Data Engineer — Spark, Hive & ML
Real-Time Big Data Engineer — Spark, Hive & ML

Remote Jobs • United States

Remote
USD 120,000 - 160,000
Big Data Engineer
Big Data Engineer

Compunnel, Inc. • Tysons (VA)

On-site
USD 150,000 - 190,000
Hadoop Developer
Hadoop Developer

Veriipro • Addison (TX)

On-site
USD 120,000 - 150,000
Big Data Developer Sr
Big Data Developer Sr

eNcloud Services LLC • Katy (TX)

On-site
USD 100,000 - 170,000
Big Data Solution Engineer
Big Data Solution Engineer

Summitworks • Durham (NC)

On-site
USD 90,000 - 120,000
Bigdata Engineer
Bigdata Engineer

Disys - Oak Brook • Tampa (FL)

On-site
USD 90,000 - 120,000
Big Data Sr. Developer cum Architect
Big Data Sr. Developer cum Architect

US Tech Solutions • Fort Worth (TX)

On-site
USD 110,000 - 130,000
BigData Consultant
BigData Consultant

Jobsbridge • Santa Ana (CA)

On-site
USD 100,000 - 130,000