GenAI Data Engineer with GenAI projects expertise, Python, Azure/AWS

PwC

Hyderabad, Pune District, Bengaluru

Hybrid

INR 1,200,000 - 1,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

PwC Hyderabad is seeking a data engineer to design, develop, and maintain data pipelines and ETL processes for GenAI projects. You will collaborate with data scientists and software engineers to implement machine learning models and scalable data architectures.

The role emphasizes cloud-native development, containerized deployments on Azure/AWS, and building robust data lakes. You will ensure real-time data processing, apply security best practices, and stay current with GenAI advances to drive

Qualifications

  • 3-5 years of GenAI-focused experience in software/ data roles.
  • 3+ years hands-on Python development.
  • Experience with Flask or FastAPI web frameworks.

Responsibilities

  • Design, develop, and maintain data pipelines and ETL processes for GenAI projects.
  • Collaborate with data scientists and software engineers to implement ML models and algorithms.
  • Optimize data infrastructure for scalable data processing.

Skills

Python
GenAI
Spark
SQL
Git
CI/CD
REST APIs
Microservices

Tools

LangChain
Semantic Kernel
LlamaIndex
Azure Kubernetes Service
Azure Container Instances

Job description

Role & responsibilities
Responsibilities:
  • Design, develop, and maintain data pipelines and ETL processes for GenAI projects.
  • Collaborate with data scientists and software engineers to implement machine learning models and algorithms.
  • Optimize data infrastructure and storage solutions to ensure efficient and scalable data processing.
  • Implement event-driven architectures to enable real-time data processing and analysis.
  • Utilize containerization technologies like Kubernetes and Docker for efficient deployment and scalability.
  • Develop and maintain data lakes for storing and managing large volumes of structured and unstructured data.
  • Implement and integrate LLM frameworks (Langchain, Semantic Kernel) for advanced language processing and analysis.
  • Collaborate with cross-functional teams to design and implement solution architectures for GenAI projects.
  • Utilize cloud computing platforms such as Azure or AWS for data processing, storage, and deployment.
  • Monitor and troubleshoot data pipelines and systems to ensure smooth and uninterrupted data flow.
  • Stay up-to-date with the latest advancements in GenAI technologies and recommend innovative solutions to enhance data engineering processes.
  • Collaborate with cross-functional teams to understand business requirements and translate them into technical solutions.
  • Document data engineering processes, methodologies, and best practices.
  • Maintain solution architecture certificates and stay current with industry best practices.
Requirements:
  • Python Proficiency: Minimum 3 years of hands-on experience building applications with Python.
  • Scalable System Design: Solid understanding of designing and architecting scalable Python applications, particularly for Gen AI use cases, with a strong understanding of various components and systems architecture patterns to make cohesive and decoupled, scalable applications.
  • Web Frameworks: Familiarity with Python web frameworks (Flask, FastAPI) for building web applications around AI models.
  • Modular Design & Security: Demonstrated ability to design applications with modularity, reusability, and security best practices in mind (session management, vulnerability prevention, etc.,).
  • Cloud-Native Development: Familiarity with cloud-native development patterns and tools (e.g., REST APIs, microservices, serverless functions).
  • Cloud Deployments: Experience deploying and managing containerized applications on Azure/AWS (Azure Kubernetes Service, Azure Container Instances, or similar).
  • Version Control (Git): Strong proficiency in Git for effective code collaboration and management.
  • CI/CD: Knowledge of continuous integration and deployment (CI/CD) practices on cloud platforms.
  • 3-5 years of relevant technical/technology experience, with a focus on GenAI projects.
  • Strong programming skills in Python.
  • Experience with data processing frameworks like Apache Spark or similar.
  • Proficiency in SQL and database management systems.
Preferred Skills:
  • Gen AI Frameworks: Experience with LLM frameworks or tools for interacting with LLMs such as LangChain, Semantic Kernel, LlamaIndex
  • Data Pipelines: Experience in setting up data pipelines for model training and real-time inference.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Engineer - Data & AI
Lead Engineer - Data & AI

Quest Global • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Data Scientist
Data Scientist

Dun & Bradstreet • Hyderabad

Hybrid
INR 2,500,000 - 5,500,000
GenAI- Python Lead
GenAI- Python Lead

Reyika • Dadri, Delhi

On-site
INR 1,800,000 - 3,000,000
Data Scientist
Data Scientist

Dun & Bradstreet • Chennai District

On-site
INR 2,500,000 - 4,500,000
AI Data Engineer
AI Data Engineer

EXL • Maharashtra

On-site
INR 3,500,000 - 6,000,000
Data Scientist
Data Scientist

CoreOps.AI • Dadri

On-site
INR 1,800,000 - 3,000,000
AI Data Engineer
AI Data Engineer

ExlService Holdings, Inc. • Gurgaon

On-site
INR 1,500,000 - 2,100,000
Data Engineer - Python /GenAI
Data Engineer - Python /GenAI

Virtusa • Maharashtra

On-site
INR 1,500,000 - 2,500,000
Data Engineer - AI Labs
Data Engineer - AI Labs

IDFC FIRST Bank • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Data Scientist
Data Scientist

Dun & Bradstreet India • Chennai District

On-site
INR 2,500,000 - 4,500,000