Senior AI/ML Engineer

Jobtailor

Bengaluru

On-site

INR 5,500,000 - 9,000,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Jobtailor in Bengaluru, India, seeks a seasoned data engineer to build and operate ML-ready data systems, including data prep, feature generation, and training pipelines. You will own data versioning, validation, and reproducibility for ML workflows, and build pipelines for model training, testing, validation, deployment, and production inference.

You will partner with ML engineers to operationalize models reliably, design ETL/ELT pipelines ingesting real-time streams into lakehouse

Qualifications

  • D5+ years in data engineering or closely related field.
  • Expert-level in Python, Scala, or Java.
  • Built production ML/AI pipelines covering training, testing, validation, deployment.
  • Experience with Databricks and lakehouse architecture at scale.
  • Strong SQL and data modeling skills.

Responsibilities

  • Build and operate ML-ready data systems with data prep and feature generation.
  • Own data versioning, validation, and reproducibility for ML workflows.
  • Deploy and manage pipelines for training, testing, validation, deployment, and inference.
  • Collaborate with ML engineers to operationalize models reliably.
  • Design and operate ETL/ELT pipelines feeding lakehouse architecture.
  • Monitor throughput, latency, reliability, and cost of pipelines.

Skills

Data Engineering
ML Pipeline Development
Kubernetes Deployment
Databricks Experience
Advanced SQL Skills

Education

Master's Degree
Bachelor's Degree

Tools

Databricks
Kubernetes
AWS
GCP
Azure
MLflow
Airflow
Kubeflow
Feast
Lakehouse Architecture

Job description

  • • Build and operate ML-ready data systems, including data preparation, feature generation, and training pipelines
  • • Own data versioning, validation, and reproducibility for ML workflows
  • • Build and maintain pipelines for model training, testing, validation, deployment, and production inference
  • • Partner with ML engineers to operationalize models reliably
  • • Design, build, and operate ETL/ELT pipelines ingesting real-time event streams, third-party APIs, and media-rich sources into the lakehouse architecture
  • • Own pipeline throughput, latency, reliability, and cost
  • • Deploy, monitor, and troubleshoot containerized data services and distributed processing applications in cloud-native Kubernetes environments
  • • Build SDKs, APIs, and reusable frameworks for engineering and research teams
  • • Implement data validation, reconciliation, and monitoring processes
  • • Maintain data catalogs and metadata for discoverable, trusted, and reusable data assets
  • • Design monitoring, alerting, and logging for pipelines and infrastructure
  • • Identify and resolve operational problems proactively
Requirements
  • 5+ years in data engineering, data platform engineering, AI engineering, or a closely related field
  • Expert-level proficiency in at least one major programming language such as Python, Scala, or Java
  • Personally built or co-built production ML/AI pipelines covering training, testing, validation, and deployment
  • Hands-on production experience with Databricks, Lakehouse architecture, Delta Lake, Spark optimization, and Workflows
  • Experience managing large-scale, heterogeneous datasets on Databricks
  • Deep, production-scale experience with Ray, Apache Spark, or an equivalent distributed processing framework
  • Hands-on experience deploying, operating, and troubleshooting production workloads on Kubernetes
  • Strong experience with AWS, GCP, or Azure and their data services
  • Advanced SQL skills for data manipulation, analysis, and optimization
  • Solid understanding of relational and NoSQL databases
  • Solid understanding of data modeling, schema design, and lake/lakehouse/warehouse patterns
  • Experience with messaging, pub/sub, queues, or streaming platforms
  • Master's or Bachelor's degree in Computer Science, Engineering, Data Science, or a related field, or equivalent professional experience
  • Preferred: Experience with MLOps tooling such as MLflow, Airflow, Kubeflow, or Feast
  • Preferred: Familiarity with feature stores, vector databases, embedding pipelines, or retrieval systems
  • Preferred: Experience with batch and online inference workloads and latency-sensitive environments
  • Preferred: Experience supporting generative AI, LLM, or multimodal AI workloads
  • Preferred: Optimized GPU utilization during model training or fine-tuning
  • Preferred: Understanding of distributed computing fundamentals, concurrency, consistency, and fault tolerance
  • Preferred: CI/CD practices and Infrastructure-as-Code
  • Preferred: Contributions to open-source data or AI projects
Core Competencies

Demonstrates expertise in building and operating ML-ready data systems, including data preparation, feature generation, and production ML pipelines. Proficient in deploying and managing data services in cloud-native environments, with a strong focus on data validation, monitoring, and operational reliability.

Highest-signal resume keywords
  • Data Engineering
  • ML Pipeline Development
  • Kubernetes Deployment
  • Databricks Experience
  • Advanced SQL Skills
Hard Skills
  • Python
  • Scala
  • Java
  • Apache Spark
  • Ray
  • SQL
  • Data Modeling
  • ETL/ELT Pipelines
  • Data Validation
  • Distributed Processing
Certifications & Qualifications
  • Master's Degree in Computer Science
  • Bachelor's Degree in Engineering
  • Bachelor's Degree in Data Science
Industry Keywords
  • ML/AI Pipelines
  • Data Catalogs
  • Streaming Platforms
  • Batch Inference
  • Online Inference
  • Generative AI
  • Open-Source Contributions
  • CI/CD Practices
  • Infrastructure-as-Code
  • Fault Tolerance
Tools & Technologies
  • Databricks
  • Kubernetes
  • AWS
  • GCP
  • Azure
  • MLflow
  • Airflow
  • Kubeflow
  • Feast
  • Lakehouse Architecture
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Software Engineer
Lead Software Engineer

Impetus • Bengaluru

On-site
INR 1,000,000 - 2,000,000
Senior Developer, Data Engineer
Senior Developer, Data Engineer

ICE Clear Europe Limited • Hyderabad

On-site
INR 1,500,000 - 2,000,000
Senior Developer, Data Engineer
Senior Developer, Data Engineer

Intercontinental Exchange Holdings, Inc. • Hyderabad

On-site
INR 3,000,000 - 5,500,000
Senior Developer, Data Engineer
Senior Developer, Data Engineer

ICE • Hyderabad

On-site
INR 1,800,000 - 3,000,000
DevSecOps and AI Engineer
DevSecOps and AI Engineer

Jobtailor • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Senior Engineer - Data Engineering
Senior Engineer - Data Engineering

KSB Company • Maharashtra

On-site
INR 600,000 - 1,000,000
Senior Data Engineer (AI/ML)
Senior Data Engineer (AI/ML)

Neolatika • Maharashtra

On-site
INR 2,000,000 - 3,600,000
Technical Lead - Data Engineer (Data&AI)
Technical Lead - Data Engineer (Data&AI)

Srijan Technologies PVT LTD • Gurugram District

On-site
INR 4,000,000 - 7,500,000
Technical Lead - Data Engineer (Data&AI)
Technical Lead - Data Engineer (Data&AI)

Srijan: Now Material • Gurugram District

On-site
INR 3,000,000 - 6,000,000
Lead Data Engineer
Lead Data Engineer

Talentrabbit • Hyderabad

On-site
INR 1,800,000 - 2,500,000