Role: AWS Databricks Developer
Experience: 616 Years
Location: Bangalore
CTC: Up to 40 LPA
Job Summary
We are looking for an experienced AWS Databricks Developer with strong hands-on expertise in Databricks, AWS, PySpark, Apache Spark and SQL. The candidate will be responsible for designing, developing, optimizing and maintaining scalable data engineering pipelines and data processing solutions on AWS Databricks.
Key Responsibilities
- Design and develop scalable data pipelines using AWS Databricks, PySpark, Spark and SQL.
- Develop ETL/ELT workflows for ingestion, transformation and processing of large datasets.
- Build and maintain Databricks notebooks, jobs, workflows and Delta Lake tables.
- Integrate Databricks with AWS services such as S3, Glue, Redshift, Lambda and IAM.
- Perform Spark and Databricks performance tuning including optimization of jobs, queries, clusters and data processing.
- Work extensively with Delta Lake for data storage, transformation and management.
- Implement data quality, validation and error-handling mechanisms across pipelines.
- Troubleshoot and resolve production issues related to data pipelines and Databricks workloads.
- Follow coding standards, version control and CI/CD practices.
- Collaborate with data engineers, architects and business stakeholders to deliver data solutions.
- Contribute to technical design, solution architecture and continuous improvement of data platforms.
Mandatory Skills
- AWS Databricks
- Databricks
- PySpark
- Apache Spark
- Python
- SQL
- AWS S3
- AWS Glue
- Delta Lake
- Data Engineering / ETL / ELT
- Databricks Jobs & Workflows
- Data Pipeline Development
- Spark Performance Tuning
Candidate Profile
- 6–16 years of overall IT experience with strong relevant experience in data engineering.
- Strong hands-on experience with AWS and Databricks.
- Excellent expertise in PySpark, Python and SQL.
- Experience developing and supporting large-scale data pipelines.
- Strong understanding of Data Lake, ETL/ELT and Data Warehouse concepts.
- Good understanding of distributed data processing using Apache Spark.
- Experience in production environments with strong debugging and problem-solving skills.
- Ability to work independently as well as collaborate with cross-functional teams.
Key Technology Stack
Cloud: AWS
Data Platform: Databricks
Programming: Python, PySpark, SQL
Processing: Apache Spark
Storage: AWS S3, Delta Lake
AWS Services: Glue, Redshift, Lambda, IAM, CloudWatch
DevOps: Git, CI/CD, Terraform
Data Engineering: ETL/ELT, Data Pipelines, Data Transformation, Data Quality