- Design and maintain batch and real-time data pipelines using Python, SQL, and suitable orchestration tools.
- Ingest and transform data from APIs, databases, satellite imagery, public datasets, and external sources.
- Build cloud-based data lakes, warehouses, and curated datasets on GCP or AWS.
- Develop data models and optimise storage, partitioning, indexing, and query performance.
- Build geospatial ETL workflows for formats such as NetCDF, Zarr, GeoTIFF/COG, and vector data.
- Prepare reliable datasets for analytics, ML training, inference, RAG, and application APIs.
- Implement data quality checks, lineage, documentation, monitoring, retries, and backfills.
Cloud and DevOps
- Provision and manage cloud infrastructure required for data, geospatial, and AI workloads.
- Build CI/CD pipelines for data pipelines, backend services, and ML workloads.
- Use Docker and services such as Kubernetes, Cloud Run, or equivalent deployment platforms.
- Implement monitoring, logging, IAM, secrets management, and cost optimisation.
- Collaborate with AI, geospatial, platform, and product teams on deployment and scaling issues.
Selection Criteria
Educational Qualification
- B.E./B.Tech in Computer Science, IT, or a related discipline, or equivalent relevant experience.
- Relevant cloud or data engineering certifications are preferred.
Must Have
- 5+ years of experience in data engineering, including production-grade data pipelines.
- Strong proficiency in Python and SQL.
- Experience with ETL/ELT, data modelling, APIs, relational databases, and data lake or warehouse architectures.
- Hands-on experience with GCP or AWS data and infrastructure services.
- Experience with PostgreSQL/PostGIS or similar databases.
- Experience with orchestration tools such as Airflow or Prefect and processing frameworks such as Spark or Dask.
- Working knowledge of Linux, Docker, CI/CD, cloud networking, monitoring, and infrastructure operations.
- Strong understanding of data quality, security, performance, and cost optimisation.
Preferred
- Experience with geospatial data stacks and tools such as PostGIS, GeoPandas, Rasterio, Xarray, GeoServer, or TiTiler.
- Experience supporting AI/ML workloads, GPU environments, model serving, or MLOps platforms.
- Familiarity with vector databases, embedding pipelines, and RAG systems.
- Exposure to Kubernetes and infrastructure-as-code tools such as Terraform.
- Experience with government cloud, Digital Public Infrastructure, climate tech, or public-data systems.
Compensation
Competitive compensation - commensurate with the experience and matching the best of the standards adopted by the industry or other similar organisations for similar roles.
Application Process
CEEW is an equal-opportunity employer, and the selection process does not discriminate on the basis of age, gender, caste, ethnicity, religion, or sexuality. Female candidates are encouraged to apply.