Data Engineer - PySpark and Apache

Infosys Limited

Atlanta (GA)

Hybrid

USD 120,000 - 180,000

Full time

13 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Long-term disability
Health and dependent care accounts
Insurance (various)

Job summary

Infosys Limited is seeking a Technology Consultant 2 in Atlanta, GA to design, develop, and optimize data pipelines and analytics solutions. You will work with PySpark, Python, SQL, and Spark, contributing to end-to-end data processing and integration projects.

The role emphasizes collaboration with data scientists and engineering teams, implementing ETL/ELT workflows, and ensuring robust production systems while using AWS, Airflow, Snowflake, and Databricks.

Qualifications

  • Design, develop, and maintain scalable data pipelines using PySpark, Python, SQL, and Apache Spark.
  • Experience in data engineering, big data processing, or a closely related role.
  • Hands-on proficiency in Python, SQL, PySpark, and Apache Spark.

Responsibilities

  • Contribute to requirements elicitation by documenting assigned parts of business requirements.
  • Facilitate software design discussions and document decisions to guide development teams.
  • Participate in coding and integrate new features into existing applications while maintaining stability.
  • Conduct code reviews and maintain code repositories.
  • Implement test strategies, analyze results, and coordinate bug fixes to uphold quality.
  • Develop user training programs, documentation, and support frameworks for new software.

Skills

PySpark
Python
SQL
Apache Spark
Data pipelines
ETL/ELT pipelines
Airflow
AWS
Snowflake
Databricks
MLflow
Problem solving

Education

Bachelor’s degree in a related field

Tools

Apache Airflow
AWS (S3, Glue, Redshift, Lambda)
Snowflake
Databricks
MLflow

Job description

Technical Skills 2

Technical Skills 3

Technology|Cloud Platform|AWS App Development

Overview

The Infosys Data and Analytics (DNA) unit is at the forefront of transforming data into actionable insights, driving business growth and operational efficiency. We specialize in leveraging advanced AI and analytics to create innovative solutions that address complex business challenges. Our team is dedicated to pioneering the future of data-driven decision-making, enabling organizations to unlock new opportunities and achieve sustainable success. Join us to be part of a dynamic team that is revolutionizing the way businesses harness the power of data and AI. At Infosys DNA, you'll have the opportunity to work with cutting-edge technologies, collaborate with industry experts, and contribute to transformative projects that shape the future of business. We are committed to fostering a culture of continuous learning and growth, ensuring that our team members thrive in a dynamic and supportive environment. If you're passionate about AI and eager to make a significant impact, the Infosys DNA unit is the perfect place for you to grow and excel.

In the assigned Job Role of Technology Consultant 2, your Area Of Responsibility will be as below:
  • Contribute to the requirements elicitation process by documenting assigned parts of business requirements, in line with guidance provided
  • Facilitate software application design discussions, and document design decisions to guide the technical team towards building software solutions
  • Participate in coding and integrate new features or updates into existing applications, with a focus on maintaining system stability
  • Conduct code reviews, do changes to the codebase and maintain code repositories
  • Implement test strategies, analyse results, and coordinate bug fixes to uphold the software quality standards
  • Develop user training programs, documentation, and support frameworks to ensure a smooth transition to new software applications
  • Actively participate in resolving production issues and recommend preventive strategies to enhance system reliability
  • Maintain detailed records of code, testing techniques, and support activities to enrich the knowledge base and assist other similar projects
Your contribution to the team:
  • A collaborative spirit and excellent communication skills.
  • The ability to handle end to end SDLC phases from requirement gathering to implementation.
  • A knack for translating complex requirements into actionable development tasks.
  • A passion for design and hands-on coding experience
  • A proactive approach to testing, troubleshooting, and refining our applications.
  • The ability to work with cross-functional teams and do software integration.
Required Skill and Experience
  • Design, develop, and maintain scalable data pipelines using PySpark, Python, SQL, and Apache Spark.
  • experience in data engineering, big data processing, or a closely related role.
  • Strong hands-on proficiency in Python, SQL, PySpark, and Apache Spark.
  • Experience designing, developing, and maintaining scalable ETL or ELT data pipelines.
  • Experience with Apache Airflow for workflow orchestration and scheduling.
  • Hands-on experience with AWS services such as S3, Glue, Redshift, or Lambda.
  • Experience working with Snowflake and/or Databricks.
  • Proficiency in data transformation, data validation, performance tuning, and troubleshooting.
  • Working experience with relational databases such as PostgreSQL or MySQL.
  • Ability to process and integrate large datasets across multiple file formats and sources.
  • Strong analytical, problem-solving, collaboration, and communication skills
  • Build and optimize batch-processing workflows for high-volume structured and semi-structured datasets.
  • Develop ETL and data-integration solutions using Databricks, AWS Glue, Apache Airflow, and cloud storage services.
  • Ingest, standardize, transform, and validate data from sources such as CSV, JSON, Parquet, transactional systems, and cloud platforms.
  • Create automated data-quality checks to improve completeness, consistency, and accuracy of downstream data.
  • Optimize Spark workloads and complex SQL queries, including CTEs and window functions, to improve processing and reporting performance.
  • Develop and manage Airflow DAGs for workflow scheduling, orchestration, monitoring, and operational reliability.
  • Prepare data for machine learning through cleansing, feature selection, feature engineering, and scalable transformation pipelines.
  • Support supervised machine learning model development, hyperparameter tuning, evaluation, deployment, and retraining.
  • Use MLflow for experiment tracking and model lifecycle management where applicable.
  • Collaborate with data scientists, analysts, business teams, and engineering stakeholders to deliver reliable data products.
  • Troubleshoot pipeline failures, resolve processing bottlenecks, and continuously improve performance and automation.
Preferred Skill and Experience
  • Experience with supervised machine learning, feature engineering, model training, and model evaluation.
  • Hands-on experience with Pandas, NumPy, Scikit-learn, and Matplotlib.
  • Experience with MLflow for model tracking, deployment, and lifecycle management.
  • Experience implementing data-quality frameworks such as Great Expectations.
  • Knowledge of fraud detection, predictive analytics, financial data, billing, or invoice-processing use cases.
  • Experience integrating Spark workflows with cloud object storage and optimizing distributed processing.
  • AWS Certified Data Engineer - Associate or a comparable cloud/data engineering certification
Additional Required Qualifications
  • Bachelor’s degree or foreign equivalent required from an accredited institution. Will also consider three years of progressive experience in the specialty in lieu of every year of education.
  • This position may require relocation and/or travel to work/project location.
  • Candidates authorized to work for any employer in the United States without employer-based visa sponsorship are welcome to apply. Infosys is unable to provide immigration sponsorship for this role now or in the future.

Along with competitive pay, as a full-time Infosys employee you are also eligible for the following benefits:

  • Long-term/Short-term Disability
  • Health and Dependent Care Reimbursement Accounts
  • Insurance (Accident, Critical Illness , Hospital Indemnity, Legal)
  • 401(k) plan and contributions dependent on salary level
About Us

Infosys is a global leader in next-generation digital services and consulting. We enable clients in more than 50 countries to navigate their digital transformation. With over four decades of experience in managing the systems and workings of global enterprises, we expertly steer our clients through their digital journey. We do it by enabling the enterprise with an AI-powered core that helps prioritize the execution of change. We also empower the business with agile digital at scale to deliver unprecedented levels of performance and customer delight. Our always-on learning agenda drives their continuous improvement through building and transferring digital skills, expertise, and ideas from our innovation ecosystem.

EEO

Infosys provides equal employment opportunities to applicants and employees without regard to race; color; sex; gender identity; sexual orientation; religious practices and observances; national origin; pregnancy, childbirth, or related medical conditions; status as a protected veteran or spouse/family member of a protected veteran; or disability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Data Engineer (AWS & Pyspark)
Sr. Data Engineer (AWS & Pyspark)

Infosys Limited • Charlotte (NC), Northern (KY)

On-site
USD 95,000 - 135,000
Long-term/Short-term Disability
Health & Dependent Care Reimbursement
Insurance (Accident, Critical Illness,
+2
Sr. Data Engineer (AWS & Pyspark)
Sr. Data Engineer (AWS & Pyspark)

Infosys • Charlotte (NC)

On-site
USD 90,000 - 140,000
Medical/Dental/Vision/Life Insurance
Long-term/Short-term Disability
Health and Dependent Care Reinvestment
+3
Pyspark Engineer
Pyspark Engineer

Infosys Limited • Richardson (TX)

On-site
USD 110,000 - 150,000
Disability benefits
Health and Dependent Care Reimburse
Insurance (Accident, Critical Illness,
+1
Pyspark Developer
Pyspark Developer

Infosys • Hartford (CT)

On-site
USD 120,000 - 160,000
Pyspark Engineer
Pyspark Engineer

Infosys Limited • Irving (TX)

On-site
USD 90,000 - 130,000
Long-term Disability
Health and Dependent Care Reimbursement Accounts
401(k) plan
Data Engineer (ETL/Spark Technologies)
Data Engineer (ETL/Spark Technologies)

Infosys Limited • Charlotte (NC), Northern (KY)

On-site
USD 90,000 - 130,000
Disability insurance
Health and dependent care accounts
Insurance (various)
Senior Data Engineer -Python and AWS
Senior Data Engineer -Python and AWS

Infosys • West Palm Beach (FL)

On-site
USD 90,000 - 130,000
Medical/Dental/Vision Insurance
401(k) plan
Paid holidays
+2
Senior Data Engineer -Python and AWS
Senior Data Engineer -Python and AWS

Infosys Limited • West Palm Beach (FL), Northern (KY)

Hybrid
USD 90,000 - 120,000
Health insurance
401(k) retirement plan
Paid time off
+1
Spark Engineer
Spark Engineer

Infosys • Jersey City (NJ)

On-site
USD 90,000 - 120,000
Medical/Dental/Vision/Life Insurance
Long-term/Short-term Disability
Health and Dependent Care Reimburse
+3
Cloud Data Engineer-Snowflake/AWS
Cloud Data Engineer-Snowflake/AWS

Infosys Limited • Detroit (MI), Northern (KY)

On-site
USD 110,000 - 140,000
Long-term/Short-term Disability
Health and Dependent Care Reimbursemen
Insurance (Accident, Critical Illness,
+1