We are looking for a Senior Data Engineer with 3-5 years of experience to join our Engineering team. The ideal candidate will have strong expertise in designing and developing scalable data pipelines, building cloud-native data platforms, and enabling analytics and AI-driven solutions.
You will work closely with Product Managers, Data Scientists, and Software Engineers to build reliable, high-performance data solutions that support business intelligence, reporting, and machine learning initiatives.
Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines for structured and semi-structured data
- Build and optimize batch and near real-time data processing workflows
- Process large-scale datasets efficiently while ensuring reliability and scalability
- Monitor and optimize pipeline performance and data availability
Data Modeling & Storage
- Design and maintain efficient data models for analytical and operational use cases
- Develop and optimize data warehouse solutions
- Work with relational and NoSQL databases to manage large-scale datasets
- Optimize queries and storage for performance and scalability
- Build cloud-native data solutions using AWS services
- Develop distributed data processing applications using Apache Spark (PySpark)
- Integrate data from APIs, databases, files, and streaming sources
- Ensure secure, scalable, and reliable data architectures
Data Quality & Governance
- Implement automated data validation and quality checks
- Monitor production pipelines and troubleshoot data issues
- Ensure compliance with data governance, security, and privacy standards
- Maintain documentation and data lineage for engineering solutions
- Collaborate with cross-functional teams to translate business requirements into scalable technical solutions
- Participate in architecture discussions, code reviews, and technical planning
- Mentor junior engineers and promote engineering best practices
- Contribute to continuous improvement initiatives across the data platform
Requirements
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field
- 3-5 years of professional experience in Data Engineering
- Strong programming skills in Python
- Advanced proficiency in SQL with experience optimizing complex queries
- Hands-on experience with Apache Spark (PySpark)
- Experience building scalable ETL/ELT pipelines for production environments
- Strong understanding of data modeling and data warehousing concepts
- Experience with AWS services such as S3, Glue, Athena, Redshift, Lambda, or EMR
- Experience working with relational databases such as PostgreSQL or MySQL
- Familiarity with Git and CI/CD workflows
- Strong analytical and problem-solving skills
- Excellent communication and interpersonal skills
- High level of ownership and accountability
- Ability to work independently in a fast-paced environment
- Strong collaboration and stakeholder management skills
- Continuous learning mindset and passion for modern data technologies
Nice to Have
- Hands-on experience with Databricks for developing and managing data engineering workloads
- Experience with workflow orchestration tools such as Apache Airflow
- Knowledge of Apache Kafka or other streaming technologies
- Experience with Docker and containerized deployments
- Familiarity with Infrastructure as Code (Terraform or CloudFormation)
- Exposure to Delta Lake, Snowflake, or Redshift
- Experience supporting AI/ML data pipelines
We help companies move from AI-curious to AI-powered by building systems that do outcome-based work inside the enterprise.