A leading data engineering firm in Gandhamguda is seeking a Data Engineer to lead the design and optimization of cloud-based data pipelines. The role requires strong expertise in Python, SQL, and big data technologies, alongside a Bachelor's or Master's degree in a related field. The ideal candidate will have at least 4 years of experience and the ability to mentor junior engineers while ensuring data quality and security. Familiarity with Azure Data Factory or AWS Glue is a plus.
Qualifications
4+ years' experience as a data engineer or relevant role.
Understanding of Big Data technologies and Lakehouse architectures.
Deep understanding of data warehousing concepts.
Responsibilities
Lead the design and optimization of cloud-based data pipelines.
Develop infrastructure to publish customer data.
Mentor junior engineers in data engineering practices.
Skills
Python
Apache Spark
SQL
Data modeling
ETL/ELT processes
Cloud data platforms
Data governance
CI/CD pipelines
Education
Bachelor's or Master's degree in Computer Science or related field
Tools
Azure Data Factory
AWS Glue
GCP Dataflow
Snowflake
Databricks
Apache Airflow
Terraform
Job description
Core Responsibilities
Lead the design and optimization of large-scale, cloud-based data pipelines and systems.
Develop infrastructure to collect, transform, combine, and publish/distribute customer data
Mentor junior engineers and contribute to best practices for data engineering standards.
Collaborate closely with data architects, analysts, and stakeholders to deliver reliable data solutions.
Implement and optimize ETL/ELT processes for performance and scalability.
Ensure data quality, integrity, and security across all environments.
Research and implement new tools, frameworks, and automation strategies to enhance productivity.
Support DevOps initiatives through CI/CD pipeline management and infrastructure-as-code practices.
Qualifications
Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
4+ years’ experience as a data engineer or relevant role
Understanding of Big Data technologies, DataLake, DeltaLake, Lakehouse architectures
Advanced proficiency in Python, SQL, and Apache Spark.
Deep understanding of data modeling, ETL/ELT processes, and data warehousing concepts.
Experience with cloud data platforms such as Azure Data Factory, AWS Glue, or GCP Dataflow.
Strong background in performance tuning and handling large-scale datasets.
Familiarity with version control, CI/CD pipelines, and agile development practices.
Proficient with handling different file formats: JSON, Avro, and Parquet
Knowledge of one or more database technologies (e.g. PostgreSQL, Redshift, Snowflake, NoSQL databases).
Preferred Qualifications
Cloud certification from a major provider (AWS, Azure, or GCP).
Hands-on experience with Snowflake, Databricks, Apache Airflow, or Terraform.
Exposure to data governance, observability, and cost optimization frameworks.