Data Engineer

Straive

Pune District

On-site

INR 350,000 - 750,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Straive seeks an experienced Data Engineering lead to architect and optimize large-scale data pipelines using Python, PySpark, and Databricks in a risk domain. The role involves building GenAI agents, integrating microservices on OpenShift/Kubernetes, and enforcing governance across platforms.

You will work with modern tooling (Kafka, SQL, Data Mesh, Starburst) and cloud-native CI/CD to ensure scalable, compliant data platforms for AI/ML and NLP use cases.

Qualifications

  • 8+ years in large-scale application development.
  • 5+ years in a Python and PySpark Data Engineering lead role.
  • Bachelor's degree in Computer Science, Engineering, or a related field; Master's preferred.
  • Core tech: Python, PySpark, Databricks, Google ADK, LLMs, FastAPI, Spring Boot, Microservices, Kafka, SQL, Data Mesh, Starburst.
  • Infrastructure: Kubernetes, OpenShift, Docker, Cloud-Native infra, CI/CD pipelines.

Responsibilities

  • Architect and maintain enterprise-grade ELT/ETL data pipelines in Python/PySpark.
  • Build and deploy GenAI agents with Google ADK, LLMs, MCP in human-in-the-loop workflows.
  • Design and deploy microservice integrations on OpenShift/Kubernetes with CI/CD pipelines.
  • Implement data federation layers via Lambda/Data Mesh with Starburst for AI/ML use cases.
  • Leverage enterprise AI platforms and tooling to accelerate engineering velocity and governance.

Skills

Python
PySpark
Databricks
Google ADK
LLMs
FastAPI
Spring Boot
Microservices
Kafka
SQL
Data Mesh
Starburst

Education

Bachelor's degree in Computer Science, Engineering, or a related field
Master's degree preferred

Tools

Kubernetes
OpenShift
Docker
CI/CD pipelines

Job description

  • Architect and maintain enterprise-grade ELT and ETL data pipelines using Python, PySpark, Kafka, and Databricks to manage large-scale risk data.
  • Build and deploy GenAI agents utilizing Google ADK, Google Flash 2.5+ LLMs, and Model Context Protocol (MCP) integrated with Human-in-the-Loop workflows.
  • Design, automate, and deploy microservice integrations for data-intensive applications on OpenShift and Kubernetes using robust CI/CD pipelines.
  • Implement data federation layers supporting Lambda and Data Mesh architectures via Starburst to enable AI/ML and NLP use cases.
  • Leverage agentic AI platforms and development assistants such as Devin.AI and GitHub Copilot with prompt engineering to increase engineering velocity.
  • Enforce data governance, risk management policies, and regulatory compliance standards across all data platforms.
Preferred Candidate Profile
  • Work Experience: 8+ years in large-scale application development with 5+ years in a Python and PySpark Data Engineering lead role.
  • Educational Background: Bachelor's degree in Computer Science, Engineering, or a related field (Master's degree preferred).
  • Core Technical Skills: Python, PySpark, Databricks, Google ADK, LLMs, FastAPI, Spring Boot, Microservices, Kafka, SQL, Data Mesh, Starburst.
  • Infrastructure and Cloud: Kubernetes, OpenShift, Docker, Cloud-Native Infrastructure, CI/CD pipelines.
  • Industry Context: Data engineering experience in Banking Risk, Retail Products, Cards, Mortgage, Deposits, or Wealth Management.
  • Assumed Requirements / Certifications: Databricks Certified Data Engineer, AWS Certified Data Analytics, or Azure Data Engineer Associate.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

SG Analytics • Chennai District

Hybrid
INR 3,000,000 - 5,400,000
Senior Data Engineer (Immediate Joiners Only)
Senior Data Engineer (Immediate Joiners Only)

SG Analytics • Pune District, Chennai District

Hybrid
INR 5,000,000 - 7,500,000
Data Engineer - Lead
Data Engineer - Lead

Iris Software • Dadri

On-site
INR 1,500,000 - 2,500,000
Data Engineer (Python + Agentic AI + AWS) - Chennai Location
Data Engineer (Python + Agentic AI + AWS) - Chennai Location

Shree Trishakti Group • Chennai District

Hybrid
INR 1,800,000 - 3,200,000
Senior Fullstack Developer
Senior Fullstack Developer

agilisium • Chennai District

On-site
INR 800,000 - 1,200,000
Lead Data Engineer (Databricks, PySpark & GCP)
Lead Data Engineer (Databricks, PySpark & GCP)

Egen • Hyderabad

On-site
INR 5,500,000 - 7,500,000
Healthcare benefits
Performance bonus
Senior Data Engineer
Senior Data Engineer

DATAECONOMY Inc • Hyderabad

On-site
INR 1,500,000 - 2,100,000
Data Analytics Engineer
Data Analytics Engineer

EXL • Hyderabad, Pune District, Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Data Engineer (AWS, Databricks, PySpark)
Data Engineer (AWS, Databricks, PySpark)

Tata Consultancy Services • Hyderabad, Bengaluru

On-site
INR 4,000,000 - 6,000,000
Data Engineering Manager
Data Engineering Manager

Good co India • India

On-site
INR 2,400,000 - 5,400,000