Skills : Spark, Python, Databricks, AWS, Snowflake OR Kafka
Job Summary
EPAM is looking for an experienced Data Engineer with strong expertise in Spark, Python, Databricks, AWS, Snowflake, and/or Kafka. The candidate will be responsible for designing, developing, and optimizing scalable data pipelines and data processing solutions for enterprise applications.
Key Responsibilities
- Design and develop scalable data pipelines using Python and Apache Spark/PySpark.
- Develop and maintain data engineering solutions on Databricks.
- Build and manage cloud-based data solutions using AWS services.
- Design and optimize data storage and processing solutions using Snowflake.
- Implement real-time/streaming data pipelines using Apache Kafka.
- Perform data ingestion, transformation, cleansing, and validation from multiple sources.
- Optimize Spark jobs, Databricks workflows, SQL queries, and data pipelines for performance.
- Develop reusable ETL/ELT frameworks and ensure data quality and reliability.
- Work with cross-functional teams to understand data requirements and deliver scalable solutions.
- Troubleshoot production data pipeline issues and perform root-cause analysis.
- Follow best practices for security, monitoring, documentation, and data governance.
Must-Have Skills
- Strong hands‑on experience with Python.
- Strong experience with Apache Spark / PySpark.
- Hands‑on experience with Databricks.
- Experience with AWS cloud services.
- Good knowledge of SQL and data warehousing concepts.
- Experience with Snowflake and/or Kafka.
- Strong understanding of ETL/ELT, data pipelines, data transformation, and data integration.
Good to Have
- AWS services such as S3, Glue, EMR, Lambda, Redshift, and CloudWatch.
- Databricks Delta Lake, Workflows, and Unity Catalog.
- Kafka topics, producers, consumers, and streaming.
- CI/CD and DevOps practices.
- Agile/Scrum experience.