Application Reliability Engineer – Data & AI

Michelin

Maharashtra

On-site

INR 1,000,000 - 2,000,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Michelin in India is seeking a Data and AI Application Reliability Engineer to own post-deployment health, latency and performance of production AI applications and Medallion architectures.

You will collaborate with feature teams to optimize PySpark jobs, build observability dashboards and implement automated self-healing and production readiness gates. Balance proactive engineering with operational reliability, lead incident post-mortems and ensure high user satisfaction.

Qualifications

  • Relevant work experience - 5 yrs

Responsibilities

  • Performance Optimization: Analyze and refactor resource-intensive PySpark jobs, queries, and API endpoints to optimize cost, execution speed, and compute efficiency.
  • Self-Healing Automation: Develop automated recovery routines, DAG rerun triggers, and data quality checks to minimize manual intervention.
  • Engineering Alignment: Partner with Feature Teams and Architects for DoD standards and production readiness gates.
  • Production Observability & Monitoring: Design, implement real-time monitoring and alerting frameworks for Medallion pipelines, feature stores, and AI applications.
  • Incident Management & Root-Cause Analysis: Lead technical resolution for high-priority production incidents and post-mortems.

Skills

Databricks
Python
PySpark
Azure
Medallion architecture
PowerBI
CI/CD
AI/ML fundamentals
Data science fundamentals
Generative AI & LLM

Tools

Azure Data Factory
Key Vault
Azure DevOps
PowerBI

Job description

Technical Expertise

The Data and AI Application Reliability Engineer you are responsible for owning the post-deployment health, latency, and performance of productionized AI application and Medallion architectures followed for applications.

Acting as the problem-solver for technical issues affecting data pipelines, databases, and deployed AI/ML models, ensuring continuous operation and high user satisfaction.

Partner with core Feature teams to perform root-cause analysis, optimize PySpark jobs, and reduce system debt.

Build automated observability dashboards and self-healing mechanisms for automated failure recovery.

As a part of Job you are required to balance your responsibilities between Proactive Engineering & Automation and Operational Reliability/ Production Health.

Key Responsibilities
  • Performance Optimization: Analyze and refactor resource-intensive PySpark jobs, queries, and API endpoints to optimize cost, execution speed, and compute efficiency.
  • Self-Healing Automation: Develop automated recovery routines, DAG rerun triggers, and data quality checks to minimize manual intervention.
  • Engineering Alignment: Partner closely with Feature Teams and Architects to establish strict Definition of Done (DoD) standards and production readiness gates for new deployments.
  • Production Observability & Monitoring: Design, implement, and maintain real-time monitoring and alerting frameworks for Medallion architecture pipelines, feature stores, and AI Applications.
  • Incident Management & Root-Cause Analysis: Lead technical resolution for high-priority production incidents, conducting thorough post-mortems to eliminate recurring failure patterns.
Azure Data Engineering Stack
  • Proficiency in Databricks, Python, and PySpark.
  • Azure (ADLS Gen2, Azure Data Factory, Key Vault, Azure DevOps).
  • Hands-on experience with Medallion architecture.
Cloud and DevOps Fundamentals
  • Understanding of cloud computing concepts and Services, specifically Microsoft Azure.
  • Good handson experience on Python.
Good To Have Technical Abilities
  • Understanding of PowerBI reports will be a plus.
  • DevOps & CI/CD Fundamentals.
  • AI & Machine Learning Fundamentals.
  • Data Science Fundamentals.
  • Generative AI & LLM Fundamentals.
Behaviour
  • Problem Solver: Ability to reverse-engineering complex system behavior and tracking down bugs.
  • Automation-First: An instinct to automate repetitive tasks.
  • Clear Communicator: Ability to explain technical root causes to non-technical stakeholders clearly.

Relevant work experience - 5 yrs

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Solution Architect- Data and AI
Solution Architect- Data and AI

Michelin • Pune District

On-site
INR 4,000,000 - 7,000,000
Solution Architect
Solution Architect

Michelin • Pune District

On-site
INR 2,500,000 - 4,500,000
Solution Architect- Data and AI
Solution Architect- Data and AI

Michelin • Maharashtra

On-site
INR 4,200,000 - 6,400,000
Azure Data & AI Engineer
Azure Data & AI Engineer

Dentsu Global Services • India

On-site
INR 900,000 - 1,500,000
Solution Architect- Data and AI
Solution Architect- Data and AI

Michelin España Portugal SA • Pune District

On-site
INR 5,000,000 - 7,000,000
Solution Architect- Data and AI
Solution Architect- Data and AI

Michelin Reifenwerke AG & Co. KGaA • Pune District

On-site
INR 3,000,000 - 5,000,000
Technical Lead - Data Engineer ( Databricks Azure)
Technical Lead - Data Engineer ( Databricks Azure)

Srijan: Now Material • Gurugram District

On-site
INR 2,400,000 - 4,800,000
Solution Architect- Data and AI
Solution Architect- Data and AI

MICHELIN France • Pune District

On-site
INR 4,200,000 - 7,000,000
IT engineer Data & Analytics DevOps
IT engineer Data & Analytics DevOps

Continental • India

Hybrid
INR 1,200,000 - 1,700,000
Training opportunities
Mobile and flexible working models
Sabbaticals
Azure Platform Site Reliability Engineer
Azure Platform Site Reliability Engineer

Foss United • India

On-site
INR 2,800,000 - 4,600,000