Associate - AI Tooling Ops - Platform Reliability Engineer

Jefferies Financial Group Inc.

Pune District

On-site

INR 1,600,000 - 2,800,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Jefferies Financial Group Inc. is seeking an Associate Platform Reliability Engineer to help ensure the stability and scalability of AI Tooling Ops across global platforms.

The role emphasises automation, observability, and proactive reliability engineering in a fast-paced post-trade environment. You will partner with engineering, infrastructure, and business teams to reduce manual toil, implement scalable solutions, and deliver highly available services with strong monitoring using Grafana,

Qualifications

  • We require a bachelor's degree in CS/Eng/IT or related field.
  • 3+ years in SRE, PRE, DevOps, or production/support roles.
  • Strong coding/scripting in Python, Go, C#, Java, or C++.
  • Solid software engineering fundamentals and system design knowledge.
  • Experience supporting distributed production apps and Linux/Windows.

Responsibilities

  • Collaborate with global reliability teams to design and maintain AI tooling infra on AWS Kubernetes.
  • Monitor health, identify risks, and improve reliability, performance, and availability.
  • Lead incident triage, troubleshooting, and post‑incident reviews to prevent recurrence.
  • Work with stakeholders to implement scalable, resilient solutions across platforms.
  • Develop automation to reduce toil and improve service efficiency.
  • Build deployment, monitoring, and observability capabilities across stacks.
  • Create dashboards, alerts, and service health monitoring with Grafana, Datadog, Prometheus, OpenTelemetry.
  • Analyze logs and metrics to identify bottlenecks and reliability issues.
  • Support Kafka‑based messaging and event‑driven architectures.
  • Troubleshoot complex issues across apps, middleware, databases, and cloud infra.
  • Follow best practices for monitoring, capacity planning, availability management, and Ops excellence.
  • Participate in production support, problem management, release and change management.
  • Partner with APAC/EMEA/AMER teams on strategic tech initiatives.
  • Drive continuous improvement in reliability and automation.

Skills

Python
Go
C#
Java
C++
SRE
DevOps
Automation
Observability
Kubernetes
AWS
Monitoring
Grafana
Datadog
Prometheus
OpenTelemetry
Kafka
CI/CD
IaC

Education

Bachelor's degree in Computer Science/Engineering/ IT

Tools

Grafana
Datadog
Prometheus
OpenTelemetry
Loki
Jaeger
Docker
Kubernetes
OpenShift
Kafka
Redis
MongoDB
Elasticsearch
AWS
Azure
GCP

Job description

Associate Platform Reliability Engineer (AI Tooling Ops)

Location: Pune


Role Overview

We are seeking a highly motivated Associate Platform Reliability Engineer (AI Tooling Ops) to join our global Platform Reliability Engineering team. This is a hands-on engineering role focused on ensuring the stability, reliability, scalability, and operational excellence of critical front-to-back business platforms supporting post-trade processing and operations.


The ideal candidate will have a strong software engineering foundation, production support experience, and a passion for automation, observability, and reliability engineering. You will work closely with development, infrastructure, and business teams to improve system resilience, enhance operational visibility, reduce manual intervention, and deliver highly available services.


Key Responsibilities


  • Partner with a high-performing global reliability engineering team to design, build, and maintain our AI Tooling Infrastructure running on AWS Kubernetes.

  • Monitor platform health, proactively identify risks, and drive improvements in system reliability, performance, and availability.

  • Perform incident triage, troubleshooting, communication, and post-incident reviews to minimize business impact and prevent recurrence.

  • Collaborate with engineering, infrastructure, and business stakeholders to design and implement scalable and resilient solutions.

  • Develop automation to reduce operational toil, eliminate manual support activities, and improve service efficiency.

  • Build and enhance deployment, monitoring, alerting, and observability capabilities across the platform stack.

  • Develop and maintain dashboards, alerts, and service health monitoring using Grafana, Datadog, Prometheus, and OpenTelemetry.

  • Analyze logs, metrics, and distributed traces to proactively identify system bottlenecks and reliability issues.

  • Support enterprise messaging and event-driven architectures, including Kafka-based platforms and integrations.

  • Troubleshoot complex issues across applications, middleware, databases, messaging platforms, infrastructure, and cloud environments.

  • Implement best practices for monitoring, observability, capacity planning, availability management, and operational excellence.

  • Participate in production support, problem management, release management, and change management activities.

  • Collaborate with regional teams across APAC, EMEA, and the Americas on strategic technology initiatives.

  • Drive continuous improvement initiatives focused on reliability, automation, monitoring, and operational efficiency.


Required Qualifications


  • Bachelors degree in Computer Science, Engineering, Information Technology, or a related discipline.

  • 3+ years of experience in Site Reliability Engineering (SRE), Platform Reliability Engineering (PRE), DevOps, Production Support, or Application Support.

  • Strong programming and scripting experience in one or more languages such as Python, Go, C#, Java, or C++.

  • Solid understanding of software engineering principles, data structures, algorithms, and system design.

  • Experience supporting and troubleshooting distributed applications in production environments.

  • Strong working knowledge of Linux/Unix and Windows Server environments.

  • Experience with relational and NoSQL databases, including performance analysis and troubleshooting.

  • Hands-on experience with observability and monitoring platforms including Grafana, Datadog, Prometheus, OpenTelemetry, Loki, and Jaeger.

  • Strong understanding of observability concepts, including Metrics, Logging, Tracing, Alerting, SLOs, SLIs, and Service Health Monitoring.

  • Experience creating operational dashboards, alerts, runbooks, and monitoring solutions to improve platform visibility.

  • Good understanding of event-driven architectures and enterprise messaging platforms such as Kafka.

  • Ability to troubleshoot message flows, APIs, middleware components, and distributed systems.

  • Understanding of incident management, problem management, root‑cause analysis, and operational support processes.

  • Familiarity with source control, CI/CD pipelines, Infrastructure as Code (IaC), and DevOps practices.

  • Excellent analytical and problem-solving skills with the ability to diagnose issues across the technology stack.

  • Strong verbal and written communication skills with the ability to engage both technical and business stakeholders.

  • Self-motivated, detail-oriented, and capable of working independently in a fast-paced environment.


Preferred Qualifications

Observability Monitoring


  • Grafana

  • Prometheus

  • Datadog

  • OpenTelemetry

  • Loki

  • Jaeger


DevOps Automation


  • Git

  • Jenkins

  • Ansible

  • Terraform

  • CI/CD Frameworks


Container Platform Technologies


  • Docker

  • Kubernetes

  • OpenShift


Data Messaging Platforms


  • Kafka

  • Redis

  • MongoDB

  • Elasticsearch


Cloud Technologies


  • AWS

  • Azure

  • GCP

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Associate - SRE - Platform Engineering
Associate - SRE - Platform Engineering

Jefferies Financial Group Inc. • Pune District

On-site
INR 1,600,000 - 2,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Software Engineer II - SRE / DevOps / Observability
Software Engineer II - SRE / DevOps / Observability

Align Knowledge Centre • Pune District

Hybrid
INR 3,000,000 - 6,000,000
58603 Platform Engineer
58603 Platform Engineer

Cephas Consultancy Services Private Limited • Pune District

On-site
INR 600,000 - 800,000
Sr Engg Manager, SRE & Platform Automation
Sr Engg Manager, SRE & Platform Automation

Aziro • Bengaluru

On-site
INR 4,500,000 - 7,000,000
Associate - AI Tooling Ops - Platform Reliability Engineer
Associate - AI Tooling Ops - Platform Reliability Engineer

Jefferies • Maharashtra

On-site
INR 1,500,000 - 2,500,000
Engagement Manager - Support Lead
Engagement Manager - Support Lead

Quantiphi Analytics Solutions • Thiruvananthapuram

On-site
INR 2,500,000 - 3,500,000
Analyst II, Production Support
Analyst II, Production Support

fis • Pune District

On-site
INR 1,500,000 - 2,300,000
Senior Software Engineer
Senior Software Engineer

NVIDIA • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hdfc Bank • Bengaluru

On-site
INR 2,500,000 - 4,000,000