AuxoAI is seeking a Senior Data Engineer to lead the design, development, and optimisation of modern data pipelines and cloud-native platforms. This role is ideal for someone with deep experience building scalable batch and streaming data workflows across cloud and lakehouse environments, strong hands‑on engineering skills, and a drive to mentor junior engineers.
You will work closely with AI engineers, solution architects, and cross‑functional teams to build production‑grade pipelines spanning ingestion, transformation, and curated data delivery — enabling high‑quality data for AI and analytics use cases at scale.
Location: Bangalore / Mumbai / Hyderabad / Gurgaon (Hybrid — 3 days in office)
Responsibilities
- - Design and build scalable batch and streaming data pipelines across bronze, silver, and gold medallion layers.
- - Build historical and incremental ingestion using Auto Loader/Spark Structured Streaming/Kafka feeds, with GCS/Azure/AWS storage and Databricks Jobs orchestration.
- - Develop and maintain Databricks-based pipelines using Spark and Delta Lake for lakehouse architecture, including migration of legacy or on‑premises data sources.
- - Design and maintain analytical data layers in BigQuery or Databricks SQL, applying best practices in partitioning, clustering, and performance tuning.
- - Implement SQL/PySpark transformations for wide and semi‑structured data, including wide‑to‑long processing and typed or hybrid models suited to consumer requirements.
- - Collaborate with AI engineers and data scientists to build pipelines that feed ML models, AI agents, and analytical systems.
- - Implement data governance, quality controls, and security best practices including schema enforcement, lineage tracking, and access controls.
- - Drive engineering best practices across CI/CD, testing, monitoring, and pipeline observability.
- - Partner with solution architects to translate data requirements into technical designs.
- - Mentor junior data engineers and contribute to documentation, code reviews, and agile ceremonies.
Requirements
- - 5+ years of hands‑on experience in data engineering, building and operating production‑grade pipelines.
- - Hands‑on experience with Databricks on GCP, including BigQuery, GCS, Databricks, Spark, Delta Lake, and structured streaming.
- - Hands‑on experience with Databricks and Apache Spark, including Delta Lake and end‑to‑end lakehouse implementations.
- - Strong programming skills in Python and/or Scala, with solid SQL for modelling and transformation.
- - Experience with data modelling, ETL/ELT, pipeline orchestration, and data warehousing concepts, including experience working with large, evolving JSON/map/array payloads, wide‑to‑long transformations, event‑time context joins and schema‑change handling.
- - Familiarity with Git, CI/CD pipelines, and data quality monitoring frameworks.
- - Solid understanding of data architecture, schema design, and performance tuning.
- - Experience with Unity Catalog, source reconciliation, schema evolution, correction handling and replay/recovery testing.
- - Strong problem‑solving and collaboration skills.
Bonus Skills
- - GCP Professional Data Engineer certification.
- - Experience with Vertex AI, Cloud Functions, Dataproc, or real‑time streaming architectures.
- - Experience with factory or industrial data sources — MES systems, IoT sensor streams, or operational telemetry.
- - Familiarity with data governance and cataloguing tools such as Dataplex, Unity Catalog, Atlan, or Collibra.
- - Exposure to Docker, Kubernetes, API integration, and infrastructure‑as‑code (Terraform).