About The Role
The role owns the design, implementation, and scaling of the core data infrastructure, building robust ELT pipelines that power enterprise analytics and machine learning applications.
The data engineering team works at the intersection of software engineering and analytics, ensuring high-throughput data ingestion, transformation, and availability across distributed cloud environments.
Key Responsibilities
- Design and build scalable data pipelines in Python and SQL using modern orchestration tools like Apache Airflow or Prefect to process large-scale streaming and batch data sets
- Develop and maintain high-performance data models within cloud data warehouses such as Snowflake or BigQuery following dimensional modeling best practices
- Implement automated data quality checks, anomaly detection, and lineage tracking to maintain strict reliability and governance standards
- Optimize data storage, query performance, and pipeline execution times to minimize cloud infrastructure costs and maximize throughput
- Collaborate with data analysts and data scientists to operationalize analytical datasets and machine learning features for production use
What We Are Looking For
- 3–6 years of professional experience in data engineering or backend development focused on large-scale data systems
- Advanced proficiency in SQL, Python, and shell scripting, with a strong emphasis on writing clean, modular, and testable code
- Hands-on experience with cloud data warehouses (Snowflake, BigQuery, or Redshift) and distributed data processing frameworks (Spark, dbt)
- Experience with cloud infrastructure (AWS, GCP, or Azure) and infrastructure-as-code tools like Terraform or CloudFormation
- Bonus: Familiarity with real-time streaming technologies such as Kafka or Flink, and experience contributing to open-source data tools