PySpark Developer

Inizio Partners Corp

Hartford (CT)

On-site

USD 90,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading data solutions company is seeking a Python and PySpark Developer in Hartford, CT. You'll design and optimize big data pipelines, collaborate with cross-functional teams, and ensure efficient data processing. Ideal candidates have strong skills in Python and PySpark, with experience in data frameworks and cloud services. An academic background in Computer Science is required. This position offers an engaging environment with opportunities for professional growth.

Responsibilities

  • Design and implement scalable data pipelines using PySpark.
  • Develop reusable and efficient code for data extraction, transformation, and loading.
  • Optimize data workflows for performance and cost efficiency.
  • Process and analyze structured and unstructured datasets.
  • Build and maintain data lakes and data warehouses.
  • Collaborate with cross-functional teams on business requirements.
  • Troubleshoot performance bottlenecks in big data pipelines.
  • Write clean and well-documented code.
  • Ensure compliance with data governance policies.

Skills

Proficient in Python
Experience in data processing libraries like Pandas and NumPy
Strong experience with PySpark and Apache Spark
Hands-on experience with big data platforms
Familiarity with cloud services like AWS, Azure, or Google Cloud
Strong knowledge of SQL and NoSQL databases
Experience with workflow orchestration tools
Problem-solving ability
Strong communication skills

Education

Bachelors or Masters degree in Computer Science

Tools

Apache Airflow
Hadoop
Databricks

Job description

We are seeking a highly skilled and experienced Python and PySpark Developer to join our team. The ideal candidate will be responsible for designing, developing, and optimizing big data pipelines and solutions using Python, PySpark, and distributed computing frameworks. This role involves working closely with data engineers, data scientists, and business stakeholders to process, analyze, and derive insights from large-scale datasets.

Key Responsibilities
  • Design and implement scalable data pipelines using PySpark and other big data frameworks.
  • Develop reusable and efficient code for data extraction, transformation, and loading (ETL).
  • Optimize data workflows for performance and cost efficiency.
Data Analysis & Processing
  • Process and analyze structured and unstructured datasets.
  • Build and maintain data lakes, data warehouses, and other storage solutions.
Collaboration & Problem Solving
  • Collaborate with cross-functional teams to understand business requirements and translate them into technical solutions.
  • Troubleshoot and resolve performance bottlenecks in big data pipelines.
Code Quality & Documentation
  • Write clean, maintainable, and well-documented code.
  • Ensure compliance with data governance and security policies.
Required Skills & Qualifications
Programming Skills
  • Proficient in Python with experience in data processing libraries like Pandas and NumPy.
  • Strong experience with PySpark and Apache Spark.
  • Hands‑on experience with big data platforms such as Hadoop, Databricks, or similar.
  • Familiarity with cloud services like AWS (EMR, S3), Azure (Data Lake, Synapse), or Google Cloud (BigQuery, Dataflow).
Database Expertise
  • Strong knowledge of SQL and NoSQL databases.
  • Experience working with relational databases like PostgreSQL, MySQL, or Oracle.
Data Workflow Tools
  • Experience with workflow orchestration tools like Apache Airflow or similar.
Problem Solving & Communication
  • Ability to solve complex data engineering problems efficiently.
  • Strong communication skills to work effectively in a collaborative environment.
Preferred Qualifications
  • Knowledge of data Lakehouse architectures and frameworks.
  • Familiarity with machine learning pipelines and integration.
  • Experience in CI/CD tools and DevOps practices for data workflows.
  • Certification in Spark, Python, or cloud platforms is a plus.
Education
  • Bachelors or Masters degree in Computer Science, Data Engineering, or a related field.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

The Value Maximizer • United States

On-site
USD 90,000 - 120,000
Pyspark Developer
Pyspark Developer

Tata Consultancy Services • Irving (TX)

On-site
USD 100,000 - 130,000
Data Engineer/Python Developer
Data Engineer/Python Developer

TechDigital Group • Minnesota

On-site
USD 80,000 - 120,000
AWS Python Developer with Pyspark Newark, NJ, New Jersey
AWS Python Developer with Pyspark Newark, NJ, New Jersey

Polarits • Newark (NJ)

On-site
USD 140,000 - 190,000
Senior PySpark Data Engineer
Senior PySpark Data Engineer

Covetus • Irving (TX)

On-site
USD 100,000 - 130,000
Senior Python Developer - AI & Full Stack Specialist
Senior Python Developer - AI & Full Stack Specialist

Aktra • McLean (VA)

On-site
USD 100,000 - 130,000
Data Engineer
Data Engineer

SDLC Technologies • Charlotte (NC)

On-site
USD 90,000 - 150,000
PySpark / Java Developer
PySpark / Java Developer

Veriipro • Whitpain Township (PA)

On-site
USD 90,000 - 120,000
Senior Data Engineer – PySpark & Python
Senior Data Engineer – PySpark & Python

Highbrow LLC • Charlotte (NC)

On-site
USD 110,000 - 160,000
Pyspark developer (Data engineer)
Pyspark developer (Data engineer)

MTK Technologies • Irving (TX)

On-site
USD 90,000 - 120,000