Turn this role into an interview — a resume and cover letter built around what this employer wants.
Infosys in Bengaluru seeks a PySpark / Spark Developer with 5-8 years of experience to design, develop, and maintain scalable data processing solutions using Apache Spark and PySpark.
You will build and optimize ETL/ELT pipelines, analyze and load data from multiple sources, tune jobs for performance, and collaborate with data engineers and analysts. Familiarity with Hive/HDFS and CI/CD practices is a plus.
Required Skills Strong experience in PySpark, Apache Spark, and Python. Good understanding of Spark Core, Spark SQL, DataFrames, and RDDs. Experience with Hive, HDFS, SQL, and data warehousing concepts. Knowledge of Azure Databricks, AWS EMR, or Hadoop Ecosystem. Experience with performance tuning and query optimization. Hands‑on experience with Git and Agile methodologies. Strong analytical and problem‑solving skills.
Experience: 5-8Years Design, develop, and maintain scalable data processing solutions using Apache Spark and PySpark. Build and optimize ETL/ELT pipelines for large‑scale data processing. Develop Spark applications for batch and real‑time data processing. Analyze, transform, and load structured and unstructured data from multiple sources. Tune Spark jobs for performance, scalability, and reliability. Work with distributed computing frameworks and big data technologies. Collaborate with Data Engineers, Data Architects, and Business Analysts to understand requirements. Troubleshoot production issues and provide efficient solutions. Implement data quality checks and monitoring mechanisms. Follow coding standards, version control, and CI/CD best practices.