L2 Production Support Genesis

Citi

India

On-site

INR 1,200,000 - 2,400,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Citi is seeking an experienced Applications Support professional in India to deliver L2/L3 support for Big Data and Enterprise Analytics Platforms. You will troubleshoot data pipelines, monitor clusters, manage Autosys batch schedules, and automate daily toil while coordinating with cross-functional teams to resolve incidents effectively.

The role emphasizes incident management, problem management, and hands-on Unix/Linux troubleshooting, with strong SQL skills and experience in IT monitoring

Qualifications

  • Bachelor’s degree in Computer Science, Information Systems, or related field.
  • 5+ years in Application/Production Support or IT operations.
  • Experience in financial services or enterprise IT environments.
  • Proficient in Linux/Unix, SQL, and log/incident analysis.

Responsibilities

  • Provide L2/L3 support for Big Data/EAP applications.
  • Troubleshoot data pipelines, clusters, and memory bottlenecks.
  • Manage batch jobs and Autosys scheduling, including JIL configs.
  • Perform Linux/Unix troubleshooting, scripting, and automations.
  • Lead incident management, root cause analysis, and PIRs.
  • Monitor databases and run complex SQL to diagnose issues.
  • Configure and use ITRS Geneos for real-time monitoring.
  • Collaborate with cross-functional teams to restore services quickly.

Skills

Analytical thinking
Communication
Decision making
Team collaboration

Education

Bachelor’s degree in CS/IS

Tools

Autosys
Unix/Linux
SQL
ITRS Geneos
Hadoop/Hive/Spark
Kafka

Job description

Key Responsibilities

1. Big Data & EAP (Enterprise Application/Analytics Platform) Support

  • Support Distributed Environments: Provide Level 2 (L2) and Level 3 (L3) support for applications hosted on Big Data platforms and Citi’s Enterprise Application/Analytics Platform (EAP).
  • Troubleshoot Data Pipelines: Diagnose and resolve failures in complex data ingestion and processing pipelines, including distributed processing frameworks (e.g., Apache Spark, Hadoop MapReduce).
  • Cluster & Resource Monitoring: Monitor cluster resource utilization (using YARN, Cloudera Manager, or similar tools) to identify and resolve memory bottlenecks, queue congestion, and job failures (e.g., Spark Out-Of-Memory errors).
  • Data Querying & Validation: Query and validate large-scale datasets stored in distributed data warehouses and file systems (e.g., HDFS, Hive, Impala, or HBase).
  • Message Queue Management: Monitor and troubleshoot real-time streaming and messaging platforms (e.g., Apache Kafka), managing consumer groups, offsets, and partition lags.

2. Batch Management & Job Scheduling (Autosys)

  • Monitor and manage batch execution: Oversee the execution of critical daily, weekly, and monthly batch processing cycles scheduled via Autosys.
  • Troubleshoot batch failures: Rapidly diagnose and resolve Autosys job failures, analyzing log files, identifying dependency issues, and performing necessary job overrides, force-starts, or hold/release actions to minimize business impact.
  • Optimize job flows: Collaborate with development and engineering teams to define, configure, and optimize Autosys job definitions using JIL (Job Information Language).

3. Automation & Process Enhancement (Toil Reduction)

  • Identify and eliminate manual bottlenecks: Actively analyze daily support activities to identify repetitive, manual tasks (“toil”) and design automated solutions to eliminate them.
  • Develop automation scripts: Write, test, and deploy robust scripts (using Python, Bash, or PowerShell) to automate routine operations, such as daily health checks, application restarts, log archiving, and data reconciliation.
  • Drive process improvements: Evaluate existing support workflows, runbooks, and escalation paths, implementing enhancements to streamline operations and reduce Mean Time to Repair (MTTR).

4. Incident Management & Production Recovery

  • Own and drive the end-to-end resolution of L2/L3 production incidents, ensuring strict adherence to corporate Service Level Agreements (SLAs) and Service Level Objectives (SLOs).
  • Lead technical triage during Major Incidents (MIM) and high‑severity outages. Coordinate effectively with cross‑functional global teams (Infrastructure, Database, Networks, Development, and Business Operations) to restore services rapidly.
  • Act as the primary technical escalation point during incidents, translating complex technical issues into clear, concise, and business‑friendly updates for senior leadership and stakeholders.
  • Ensure accurate and timely logging, categorization, and tracking of incidents within ServiceNow.

5. Problem Management & Root Cause Analysis (RCA)

  • Lead proactive Problem Management initiatives by analyzing incident trends, identifying systemic patterns, and pinpointing recurring failure points.
  • Conduct deep‑dive technical investigations—including log analysis, database queries, and infrastructure health checks to perform comprehensive Root Cause Analysis (RCA).
  • Author high‑quality Post‑Incident Reviews (PIRs) and RCA documents, detailing the timeline, root cause, impact, and preventative actions.
  • Collaborate closely with Development and Engineering teams to prioritize, track, and implement permanent bug fixes, structural workarounds, and long‑term remediations.

6. Hands‑on Unix/Linux & Application Troubleshooting

  • Perform deep‑dive technical troubleshooting directly within Unix/Linux production environments (analyzing system resources, CPU/memory bottlenecks, process states, and network connectivity).
  • Conduct advanced log analysis using Unix command‑line utilities (e.g., grep, awk, sed, find, tail) to rapidly isolate application errors and system anomalies.
  • Maintain and debug shell scripts (Bash/Korn) used for application startup, shutdown, health checks, and automated maintenance.

7. Database Support & SQL Querying

  • Troubleshoot database‑related application issues by writing and executing complex SQL queries (including multi‑table joins, subqueries, and aggregations) on databases such as Oracle, MS SQL Server, or Sybase.
  • Analyze database performance, identify slow‑running queries, and collaborate with DBAs to resolve locks, blocks, and indexing issues affecting production.

8. Proactive Monitoring & Alerting with ITRS Geneos

  • Actively monitor application health and infrastructure performance using ITRS Geneos (Active Console).
  • Configure, customize, and maintain ITRS Geneos samplers, rules, alerts, and Netprobes to ensure comprehensive coverage of critical system components.

Technical Skills & Experience

Core Requirements (Mandatory & Hands‑on)

Competency / Technology

Required Hands‑on Experience

Big Data & EAP Environments

Working experience supporting Big Data ecosystems and Enterprise Application/Analytics Platforms (EAP). Hands‑on familiarity with Hadoop, HDFS, Hive, Spark, YARN, and Kafka. Ability to troubleshoot distributed job failures and monitor cluster health.

Autosys

Strong hands‑on experience with Autosys (or similar enterprise job schedulers). Proficient in monitoring batch cycles, troubleshooting job failures, managing dependencies, performing run‑time overrides, and writing/modifying JIL (Job Information Language) configurations.

Unix / Linux

Advanced hands‑on experience navigating Unix/Linux file systems, managing processes, analyzing system performance (CPU, memory, disk I/O), and writing/debugging Shell scripts (Bash/Ksh).

SQL & Databases

Proficient in writing complex SQL queries to extract, analyze, and troubleshoot data issues across relational databases (Oracle, Sybase, or MS SQL Server). Understanding of database locks, indexing, and basic performance tuning.

ITRS Geneos

Hands‑on experience using ITRS Geneos for real‑time monitoring. Ability to navigate the Active Console, interpret alerts, and configure basic samplers, rules, and alerts.

Automation & Process Enhancement

Proven experience in automating manual support tasks and optimizing operational processes. Strong ability to identify automation opportunities and implement scripting solutions to reduce operational toil.

Incident & Problem Management

Strong expertise in ITIL processes, specifically Incident and Problem Management. Proven track record of leading major incident triage, managing SLAs, conducting Root Cause Analysis (RCA), and managing the ticket lifecycle in ServiceNow.

Secondary Technical Skills (Highly Desirable)

  • Scripting Languages: Strong proficiency in Python or Perl for building automation tools and utility scripts.
  • Middleware: Familiarity with middleware technologies (TIBCO EMS, IBM MQ, WebSphere, or Tomcat).
  • Cloud & Containers: Basic exposure to containerized environments (OpenShift, Kubernetes) or Cloud platforms (AWS).
  • Change & Release Management: Familiarity with ITIL Change Management processes and post‑release validation.

Qualifications & Soft Skills

  • Education: Bachelor’s degree in Computer Science, Information Systems, or a related field, or equivalent practical experience.
  • Experience: 5+ years of dedicated experience in Application Support or Production Support within a fast‑paced financial services or enterprise environment.
  • Analytical Thinking: Exceptional problem‑solving skills with a methodical approach to diagnosing complex technical issues under pressure.
  • Communication: Strong verbal and written communication skills, with the ability to translate complex technical issues into clear business updates for stakeholders.
Job Family Group:

Technology

Job Family:

Applications Support

Time Type:

Full time

Most Relevant Skills

Please see the requirements listed above.

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.

View Citi’s EEO Policy Statement and the Know Your Rights poster.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

L2 Production Support Genesis
L2 Production Support Genesis

Citigroup Inc. • Chennai District

Hybrid
INR 1,200,000 - 2,200,000
L2 Production Support Genesis
L2 Production Support Genesis

Citi • Chennai District

On-site
INR 900,000 - 1,300,000
L2 Application Support Engineer
L2 Application Support Engineer

Citi • Maharashtra

On-site
INR 800,000 - 1,200,000
Application Support Technology Lead Analyst - Vice President
Application Support Technology Lead Analyst - Vice President

Citi • Chennai District

On-site
INR 3,800,000 - 6,200,000
L2 Application Support Engineer
L2 Application Support Engineer

Citi • Pune District

On-site
INR 1,200,000 - 1,800,000
Senior Manager, Application Support & Data Platforms
Senior Manager, Application Support & Data Platforms

Citi • Chennai District

On-site
INR 4,000,000 - 7,000,000
Applications Support - Assistant Vice President
Applications Support - Assistant Vice President

Citigroup Inc. • Pune District

On-site
INR 1,500,000 - 2,300,000
Application Support Senior Analyst-Assistant Vice President
Application Support Senior Analyst-Assistant Vice President

Citi • Maharashtra

On-site
INR 2,500,000 - 4,200,000
Digital Production Support Analyst
Digital Production Support Analyst

Citi • Chennai District

On-site
INR 1,000,000 - 1,500,000
Applications Support - Assistant Vice President
Applications Support - Assistant Vice President

Citi • Pune District

On-site
INR 1,500,000 - 2,100,000