We're hiring a Data Engineer to build the data backbone that everything else in the company runs on from executive dashboards to machine learning models to customer-facing analytics. Every AI initiative, every business decision, and every product feature increasingly depends on clean, reliable, well-governed data arriving where it needs to be, on time. That's the job: designing and running the pipelines and infrastructure that make that possible, day in and day out.
This role sits at the foundation of the data organization. You'll work closely with data scientists, ML engineers, analysts, and backend engineers, but your core responsibility is distinct from theirs — you build and maintain the systems that move, transform, store, and secure data at scale, so that everyone downstream can trust what they're working with.
Roles & Responsibilities
Pipeline Development
- Design, build, and maintain robust ETL/ELT pipelines that ingest data from multiple sources — databases, APIs, event streams, third-party platforms — into centralized data stores.
- Write efficient, well-tested, production-grade code (primarily Python and SQL) for data transformation and validation logic.
- Build and orchestrate workflows using tools like Apache Airflow, Dagster, or equivalent schedulers, ensuring pipelines are idempotent, monitored, and recoverable.
- Handle both batch and real-time/streaming data processing needs (Kafka, Kinesis, or Spark Streaming) depending on project requirements.
Data Architecture & Warehousing
- Design and evolve data models (star schema, snowflake schema, or data vault) that balance query performance with long-term maintainability.
- Build and optimize data warehouses/lakehouses on platforms such as Snowflake, BigQuery, Redshift, or Databricks.
- Partition, index, and tune large-scale datasets for cost-efficient, fast querying by downstream teams.
- Own schema evolution and versioning as business requirements change, without breaking downstream consumers.
Data Quality & Governance
- Implement automated data quality checks, validation rules, and alerting so bad data is caught before it reaches dashboards or models.
- Maintain clear documentation of data lineage — where data comes from, how it's transformed, and who consumes it.
- Apply data governance and access-control best practices, especially for sensitive or regulated data (PII, financial data).
- Work with compliance/security teams to ensure data handling meets relevant regulatory standards.
Collaboration & Support
- Partner with data scientists and ML engineers to ensure feature pipelines are reliable and reproducible for model training.
- Support analysts and business stakeholders by making data self-serve wherever possible, reducing repetitive ad-hoc requests.
- Continuously monitor pipeline health, troubleshoot failures, and perform root-cause analysis on data incidents.
- Evaluate and recommend new data tools/platforms as the data stack scales.
Desired Candidate Profile
- 2–6 years of experience building production data pipelines, with strong command of SQL and Python.
- Hands-on experience with at least one distributed processing framework (Spark) and one orchestration tool (Airflow or equivalent).
- Working knowledge of at least one major cloud platform (AWS, Azure, or GCP) and its data services.
- Understanding of data modeling principles and experience with a modern cloud data warehouse.
- Strong debugging instincts — comfortable diagnosing why a pipeline silently produced wrong numbers, not just why it crashed.
- Bonus: experience with streaming data (Kafka), infrastructure-as-code (Terraform), or dbt for transformation management.
Growth Path:
Data Engineer Senior Data Engineer Data Platform Lead/Architect Head of Data Engineering
Perks & Benefits:
Health insurance, performance bonus, cloud certification reimbursement, hybrid work options, learning budget for data tooling courses.
Salary variants by level:
- Fresher (0–2 yrs): 5,00,000 – 9,00,000
- Senior (7+ yrs): 40,00,000 – 80,00,000+