Big Data Developer

TechWish

McLean (VA)

On-site

USD 90,000 - 130,000

Full time

8 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TechWish is seeking a seasoned data engineer to design, build, and maintain large-scale data processing pipelines using Hadoop, Spark, Python, and Scala. You will implement ingestion, storage, transformation, and analysis of scalable solutions while ensuring reliability and data quality across production systems.

The role requires collaboration with data scientists and analysts to enable data-driven decision making, ongoing learning, and staying current with industry trends in big data

Qualifications

  • Bachelor's degree in Computer Science, Information Systems or related discipline with at least five (5) years of related experience
  • Demonstrated technical expertise in object-oriented and database technologies
  • Experience with developing enterprise-grade solutions in an Agile environment
  • Strong knowledge of test automation, build automation and configuration management
  • Excellent written and verbal technical communication skills
  • Ability to work in a fast-paced environment
  • Experience with Java, Scala or Python
  • Experience with Big Data tools like Hadoop, Spark, Hive & Trino

Responsibilities

  • Design, develop, and maintain large-scale data processing pipelines using Big Data technologies
  • Implement data ingestion, storage, transformation, and analysis of scalable, efficient solutions
  • Stay current with industry trends to improve data architecture
  • Collaborate with cross-functional teams to translate business requirements into technical solutions
  • Optimize and enhance pipelines for performance, scalability, reliability
  • Develop automated testing frameworks and continuous testing for data quality
  • Conduct unit, integration, and system testing for robustness
  • Work with data scientists and analysts to support data-driven decisions
  • Monitor and troubleshoot data pipelines in production environments

Skills

Big Data
Spark
Python
Scala
SQL
Cloud AWS
Data Pipelines
Testing & QA

Education

Bachelor's degree in CS/IS or related
5+ years of related experience

Tools

Hadoop
Hive
Trino
CI/CD

Job description

  • Design, develop, and maintain large-scale data processing pipelines using Big Data technologies (e.g., Hadoop, Spark, Python, Scala).
  • Implement data ingestion, storage, transformation, and analysis of solutions that are scalable, efficient, and reliable.
  • Stay current with industry trends and emerging Big Data technologies to continuously improve the data architecture
  • Collaborate with cross-functional teams to understand business requirements and translate them into technical solutions.
  • Optimize and enhance existing data pipelines for performance, scalability, and reliability.
  • Develop automated testing frameworks and implement continuous testing for data quality assurance.
  • Conduct unit, integration, and system testing to ensure the robustness and accuracy of data pipelines.
  • Work with data scientists and analysts to support data-driven decision-making across the organization.
  • Ability to write and maintain automated unit, integration, and end-to-end tests
  • Monitor and troubleshoot data pipelines in production environments to identify and resolve issues.
  • Design, develop, and maintain large-scale data processing pipelines using Big Data technologies (e.g., Hadoop, Spark, Python, Scala).
  • Implement data ingestion, storage, transformation, and analysis of solutions that are scalable, efficient, and reliable.
  • Stay current with industry trends and emerging Big Data technologies to continuously improve the data architecture
  • Collaborate with cross-functional teams to understand business requirements and translate them into technical solutions.
  • Optimize and enhance existing data pipelines for performance, scalability, and reliability.
  • Develop automated testing frameworks and implement continuous testing for data quality assurance.
  • Conduct unit, integration, and system testing to ensure the robustness and accuracy of data pipelines.
  • Work with data scientists and analysts to support data-driven decision-making across the organization.
  • Ability to write and maintain automated unit, integration, and end-to-end tests
  • Monitor and troubleshoot data pipelines in production environments to identify and resolve issues.
Responsibilities
  • Design, develop, and maintain large-scale data processing pipelines using Big Data technologies (e.g., Hadoop, Spark, Python, Scala).
  • Implement data ingestion, storage, transformation, and analysis of solutions that are scalable, efficient, and reliable.
  • Stay current with industry trends and emerging Big Data technologies to continuously improve the data architecture
  • Collaborate with cross-functional teams to understand business requirements and translate them into technical solutions.
  • Optimize and enhance existing data pipelines for performance, scalability, and reliability.
  • Develop automated testing frameworks and implement continuous testing for data quality assurance.
  • Conduct unit, integration, and system testing to ensure the robustness and accuracy of data pipelines.
  • Work with data scientists and analysts to support data-driven decision-making across the organization.
  • Ability to write and maintain automated unit, integration, and end-to-end tests
  • Monitor and troubleshoot data pipelines in production environments to identify and resolve issues.
Education/Experience Requirements
  • Bachelor's degree in Computer Science, Information Systems or related discipline with at least five (5) years of related experience, or equivalent training and/or work experience; Master's degree and past Financial Services industry experience preferred.
  • Demonstrated technical expertise in Object Oriented and database technologies/concepts which resulted in deployment of enterprise quality solutions.
  • Past experience with developing enterprise quality solutions in an iterative or Agile environment.
  • Extensive knowledge of industry leading software engineering approaches including Test Automation, Build Automation and Configuration Management frameworks.
  • Strong written and verbal technical communication skills.
  • Demonstrated ability to develop effective working relationships that improved the quality of work products.
  • Should be well organized, thorough, and able to handle competing priorities.
  • Ability to maintain focus and develop proficiency in new skills rapidly.
  • Ability to work in a fast paced environment.
  • Experience with object oriented programming languages such as Java, Scala or Python.
Essential Technical Skills
  • AI Tool Proficiency: Hands-on experience with AI development tools (GitHub Copilot, Q Developer, ChatGPT, Claude, etc.)
  • Technical Background: Strong software development background with ability to contribute to technical discussions
  • Agile Methodology: Extensive experience with Scrum, Kanban, and continuous improvement practices
Big Data technologies
  • Experience with Big data technologies such as Hadoop, Spark, Hive & Trino
  • Evaluate understanding of common issues like:
  • Data skew and strategies to mitigate it.
  • Working with massive data volumes in PetaBytes.
  • Troublehsooting job failures due to resource limitations, bad data, scalability challenged.
  • Look for real-world debugging and mitigation stories.
AI Skills
  • Prompt Engineering: Proficiency in crafting effective prompts for AI coding assistants and analysis tools
  • AI Workflow Design: Experience redesigning development processes to leverage AI capabilities
  • Data Analysis: Ability to interpret AI-generated insights and translate them into actionable team improvements
  • Change Management: Experience leading teams through AI adoption and workflow transformation
SQL Skills (Window Functions, Joins, Complex Queries)
  • Assess comfort with SQL window functions, multi-table joins, aggregations.
  • Provide examples or ask them to write/optimize SQL queries on the spot.
  • Probe how they handle edge cases like NULLs, duplicates, ordering, etc.
Apache Spark (Development, Internals & Tuning)
  • Test their understanding of Spark s core architecture executors, tasks, stages, DAG.
  • Focus on Spark performance tuning techniques: partitioning, caching, broadcast joins, etc.
  • Ask scenario-based questions on troubleshooting slow running/stuck jobs or resource issues in Spark.
  • Explore their experience optimizing Spark jobs for large-scale datasets.
Cloud Technologies
  • Check exposure to AWS services like S3, EMR, Glue, Lambda, Athena, etc.
  • Ask how they ve used S3 with Spark (e.g., dealing with file formats, consistency issues).
  • EKS, Serverless knowledge, etc
Programming - Python or Scala
  • Assess ability to write clean, modular, and performant code.
  • Look for experience in functional programming concepts (e.g., immutability, higher-order functions).
  • Ask about real-world use cases where they wrote scalable data processing code.
  • Evaluate understanding of collections, concurrency, and memory management.
Good To Have
  • Experience with managing production data pipelines/ETL systems
  • Experience with CI/CD
  • Experience writing test cases
  • AWS certifications
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Big Data Developer
Big Data Developer

Unisys • Rockville (MD)

Hybrid
USD 120,000 - 180,000
Principal Data Engineer
Principal Data Engineer

Jobtailor • Vienna (VA)

On-site
USD 150,000 - 190,000
Bigdata Engineer
Bigdata Engineer

Disys - Oak Brook • Tampa (FL)

On-site
USD 90,000 - 120,000
Senior Staff Data Engineer
Senior Staff Data Engineer

Jobtailor • Town of Texas (WI)

On-site
USD 150,000 - 200,000
Senior Data Engineer
Senior Data Engineer

TALENT Software Services • United States

On-site
USD 140,000 - 170,000
Lead Software Engineer - Big Data
Lead Software Engineer - Big Data

JPMorgan Chase & Co. • Plano (TX)

On-site
USD 140,000 - 190,000
Software Developer
Software Developer

UCRYA • Town of Florida (NY)

On-site
USD 90,000 - 130,000
Data Engineer III
Data Engineer III

Jobtailor • Fremont (CA)

On-site
USD 120,000 - 180,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
Junior Software Engineer
Junior Software Engineer

Jobtailor • San Diego (CA)

On-site
USD 90,000 - 140,000