Senior Manager, Systems Engineering

Oracle

Bengaluru

On-site

INR 3,000,000 - 6,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Oracle in Bengaluru, India seeks a seasoned leader to head AI/GPU host operations at cloud scale. You will translate strategy into disciplined execution across day-to-day operations, incident response, patches, upgrades, and runbook maturity.

You will lead India teams, drive reliability improvements, partner with global OCI colleagues, and foster a culture of ownership and continuous learning in a fast‑growing infrastructure fleet.

Qualifications

  • BS or MS in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
  • 8+ years of experience in software engineering, infrastructure operations, cloud operations, site reliability, production operations, or related technical areas.
  • 5+ years of people management and/or technical leadership experience, including leading operators, engineers, or multiple workstreams.
  • Experience building and scaling teams, including recruiting, coaching, performance management, goal setting, and leadership development.
  • Strong operational background with incident management, service ownership, queue management, operational readiness, process improvement, and post-incident actions.
  • Experience operating or supporting Linux-based infrastructure at scale, hardware/software troubleshooting and lifecycle execution.
  • Familiarity with scripting/automation ecosystems such as Python, Bash.

Responsibilities

  • Lead, grow, and develop India-based teams of operators, developers, and tech leads in a 24x7 environment with clear ownership and accountability.
  • Recruit, hire, coach, and retain high-performing talent and raise the technical and operational bar.
  • Drive organizational excellence, ownership, continuous learning, and disciplined execution during AI/GPU infrastructure growth.
  • Own operational outcomes for AI/GPU host management including health, fleet readiness, triage, service queues, and reliability.
  • Guide teams to diagnose compute host issues across hardware, Linux, services, and networking; ensure quick, durable resolutions.
  • Drive patching, upgrades, staged rollouts, change controls, and readiness reviews with metrics-driven planning.
  • Partner with global OCI teams to align India operations with broader roadmaps and standards.

Skills

Leadership
Incident management
Operational excellence
Automation
Python
Bash
Linux
Cross-functional collaboration
Communication

Education

BS / MS in Computer Science or related field

Tools

Python
Bash

Job description

Job Description

This role reports into senior leadership responsible for global AI/GPU host management strategy, fleet readiness, AIOps adoption, intelligent automation, service reliability, and operational tooling. You will translate that strategy into disciplined execution across day‑to‑day operations, service queues, triage, incident response, patching, upgrades, runbook maturity, telemetry improvements, and operational readiness for OCI's AI/GPU host fleet.

You will be a hands‑on leader during operational escalations, driving crisp communications, rapid diagnosis, and durable corrective actions. You will partner cross‑functionally with geographically distributed OCI engineering, platform, networking, data center, compliance, and operations teams to deliver measurable reliability, efficiency, and execution outcomes.

Key Responsibilities
Organizational Leadership & Talent Development
  • Lead, grow, and develop India‑based teams of operators, developers, and technical leads in a 24x7 operational environment; establish clear ownership boundaries, on‑call expectations, and accountability mechanisms.
  • Recruit, hire, coach, and retain high‑performing talent; set goals, manage performance, develop successors, and raise the technical and operational bar across the team.
  • Create a culture of operational excellence, ownership, continuous learning, pragmatic simplification, and disciplined execution while supporting rapid AI/GPU infrastructure growth.
Service Ownership: AI/GPU Host Operations at Cloud Scale
  • Own operational outcomes for assigned AI/GPU host management services, including host health, fleet readiness, hardware/software triage, service queues, reliability, and operational readiness.
  • Guide teams that diagnose AI compute host issues across hardware, Linux, services, networking, automation, and monitoring layers; ensure issues are resolved quickly and with durable prevention mechanisms.
  • Drive operational execution for service patching, upgrades, staged rollouts, change controls, and readiness reviews using metrics‑driven planning and governance.
  • Partner with global OCI teams to ensure India operations align with broader host management roadmaps, standards, escalation practices, and fleet reliability goals.
Operational Excellence, Metrics, and Governance
  • Define and mature operational KPIs and reporting for service health, incident performance, ticket aging, ticket resolution quality, queue backlog, change execution, and operational readiness.
  • Lead high‑severity incidents and escalations: coordinate rapid triage, communicate status clearly, drive high‑quality post‑incident reviews, and follow through on corrective actions.
  • Improve operating mechanisms for risk management, audit/compliance alignment, change management, handoffs, escalation hygiene, and cross‑team execution tracking.
  • Use data and trend analysis to identify repeat issues, operational bottlenecks, staffing gaps, process defects, and opportunities for reliability improvement.
Engineering Enablement, AIOps, and Automation
  • Drive adoption of AIOps and intelligent automation to reduce manual toil, improve alert quality, accelerate event correlation, standardize remediation, and improve triage accuracy.
  • Partner with engineering and platform teams to prioritize operational tooling, telemetry improvements, workflow enablement, runbook automation, and self‑healing opportunities.
  • Establish measurable automation outcomes such as reduced manual handling, improved MTTR, increased auto‑triage coverage, fewer repeat issues, and better operator/engineer effectiveness.
  • Manage focused software engineering work aligned to operations outcomes, including scripts, dashboards, workflow tooling, triage aids, monitoring enhancements, and reliability automation.
Cross‑Functional & Stakeholder Engagement
  • Work closely with senior leaders, peer managers, technical leads, and globally distributed teams to deliver predictable execution across AI/GPU host operations.
  • Translate complex technical and operational situations into accurate narratives, decisions, risks, and action plans for senior stakeholders.
  • Represent the team in operational reviews, readiness discussions, incident forums, and cross‑functional planning sessions with clarity and ownership.
Qualifications / Experience
  • BS or MS in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
  • 8+ years of experience in software engineering, infrastructure operations, cloud operations, site reliability, production operations, or related technical areas.
  • 5+ years of people management and/or technical leadership experience, including experience leading operators, engineers, senior technical contributors, or multiple operational workstreams.
  • Experience building and scaling teams, including recruiting, hiring, coaching, performance management, goal setting, and leadership development.
  • Strong operational background with incident management, service ownership, queue management, operational readiness, process improvement, and post‑incident corrective actions.
  • Experience operating or supporting Linux‑based infrastructure at scale, including hardware/software troubleshooting and service lifecycle execution.
  • Working familiarity with scripting and automation ecosystems such as Python, Bash, or similar tools sufficient to sponsor, review, and guide operational tooling direction.
  • Understanding of distributed systems fundamentals and the ability to reason across hardware, operating systems, networking, services, monitoring, automation, and customer impact.
  • Familiarity with networking protocols such as TCP/IP and HTTP and with standard cloud infrastructure architectures.
  • Strong organizational and planning skills, including prioritization, scheduling, execution tracking, and operational governance.
  • Strong written and verbal communication skills, including the ability to communicate technical risks, operational status, tradeoffs, and execution plans to senior stakeholders.
Preferred / Nice To Have
  • Experience operating large‑scale cloud infrastructure, AI/ML infrastructure, HPC environments, or GPU fleets.
  • Experience leading teams in a 24x7 production operations environment with globally distributed stakeholders.
  • Experience driving automation, AIOps, telemetry improvements, alert noise reduction, event correlation, runbook standardization, or toil reduction programs.
  • Experience managing operational tooling or focused engineering efforts for workflow enablement, triage automation, monitoring, reliability, or service readiness.
  • Experience with service patching, upgrades, staged rollouts, change management, operational risk management, audit readiness, and compliance alignment.
  • Demonstrated success using metrics such as MTTR, ticket aging, incident volume, repeat issue rate, automation coverage, SLO/SLA performance, and readiness indicators to improve operations at scale.
Leadership Profile

You are a distributed‑systems‑mindful operations and engineering leader who values simplicity, scalability, ownership, and measurable execution. You combine people leadership with technical depth, enabling teams to operate reliable OCI services while continuously improving the operational and engineering foundations.

You are comfortable leading through ambiguity, driving accountability during escalations, and building mechanisms that help teams deliver consistently. You use data to set priorities, improve reliability, reduce toil, and strengthen execution across a rapidly growing AI/GPU infrastructure fleet.

Career Level

M3

Disability Accommodation

We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.

Equal Employment Opportunity

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Manager, Application Software Engineering
Senior Manager, Application Software Engineering

Oracle • Chennai District

On-site
INR 6,000,000 - 11,000,000
Software Developer 3 - Golang
Software Developer 3 - Golang

Oracle • Dadri

On-site
INR 1,500,000 - 2,000,000
Flexible medical options
Life insurance
Retirement options
Senior Manager, Application Software Engineering
Senior Manager, Application Software Engineering

Oracle • Ahmedabad District

On-site
INR 3,500,000 - 7,000,000
Senior Manager, Application Software Engineering
Senior Manager, Application Software Engineering

Oracle • Dadri

On-site
INR 2,500,000 - 4,200,000
Senior Manager, Application Software Engineering
Senior Manager, Application Software Engineering

Oracle • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Senior Manager, Application Software Engineering
Senior Manager, Application Software Engineering

Oracle • Hyderabad

On-site
INR 3,500,000 - 7,500,000
Flexible medical options
Life insurance
Retirement options
Software Development Manager
Software Development Manager

Oracle • Ahmedabad District

On-site
INR 4,000,000 - 7,000,000
Software Development Manager
Software Development Manager

Oracle • Dadri

On-site
INR 4,000,000 - 7,000,000
Software Development Manager
Software Development Manager

Oracle • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Competitive benefits
Flexible medical
Life insurance
+2
Software Development Manager
Software Development Manager

Ll Oefentherapie • India

On-site
INR 4,000,000 - 7,500,000