Software Engineering Technical Leader | AI Cluster Orchestrator & Automation Engineer | 15+ years

Cisco

Maharashtra

On-site

INR 2,600,000 - 3,800,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Cisco is seeking a Software Engineering Technical Leader for AI Cluster Orchestrator & Automation. You will design and implement repeatable end-to-end automation for AI cluster bring-up, validation, lifecycle management, and teardown across compute, network, and storage domains.

We partner with leading vendors to deliver curated AI infrastructure solutions built on proven reference architectures, reducing risk and accelerating value for customers.

Qualifications

  • Bachelor's degree plus 12 years related experience or Master's plus 8 years related experience
  • Experience with Linux systems and AI/GPU cluster architecture knowledge
  • Coding experience using Python and automation/API development
  • Prior experience with PXE, DHCP, Kubernetes, Slurm, BIOS/firmware, networking, and storage integration
  • Experience troubleshooting distributed provisioning failures and system dependencies

Responsibilities

  • Build idempotent orchestration workflows for GPU nodes, service nodes, network fabrics, and storage
  • Automate PXE, NVIDIA BCM, DHCP, Redfish, BIOS, firmware, OS, Kubernetes/operators, and Slurm integration
  • Coordinate dependencies across compute, Cisco networking, storage/Vast, GPU platforms, and service nodes
  • Implement health checks, configuration drift detection, validation gates, rollback, failure recovery, operational observability
  • Document runbooks, APIs, interfaces, and support handoffs

Skills

Python
Linux systems
AI/GPU cluster architecture

Education

Bachelor's degree + 12 years related experience
Master's degree + 8 years related experience

Tools

PXE
NVIDIA BCM
DHCP
Redfish
BIOS/firmware
OS
Kubernetes
Slurm

Job description

Meet the Team

We are a small, agile, and highly collaborative team at the forefront of

Meet the Team

We are a small, agile, and highly collaborative team at the forefront of AI Infrastructure Automation and Benchmarking & Certification. We partner closely with leading hardware and software vendors to design, validate, and deliver curated AI infrastructure solutions to our customers — all built on proven reference architectures that reduce risk and accelerate time-to-value.

Because we're a lean team, every member has real ownership and visibility into outcomes — from automating complex infrastructure workflows to running rigorous benchmarking and certification processes that ensure our solutions perform reliably at scale. We move fast, communicate openly, and lean on each other's expertise daily, making this a great environment for engineers who want to work across the full stack of AI infrastructure rather than being siloed into one narrow function.

Your Impact

As a Software Engineering Technical leader for AI Cluster Orchestrator & Automation, you will design and implementation of repeatable, end-to-end automation for AI cluster bring-up, configuration, validation, lifecycle management, and teardown across compute, network, and storage domains.

You Will
  • Build idempotent orchestration workflows for GPU nodes, service nodes, network fabrics, and storage.
  • Automate PXE, NVIDIA BCM, DHCP, Redfish, BIOS, firmware, OS, Kubernetes/operators, and Slurm integration.
  • Coordinate dependencies across compute, Cisco networking, storage/Vast, GPU platforms, and service nodes.
  • Implement health checks, configuration drift detection, validation gates, rollback, failure recovery, operational observability.
  • Document runbooks, APIs, interfaces, and support handoffs.
Minimum Qualifications
  • Bachelors + 12 years of related experience, or Masters + 8 years of related or equivalent related work experience.
  • Experience with Linux systems and AI/GPU cluster architecture knowledge.
  • Coding experience using Python and automation/API development
  • Prior experience with PXE, DHCP, Kubernetes, Slurm, BIOS/firmware, networking, and storage integration.
  • Experience troubleshooting distributed provisioning failures and system dependencies.
Preferred Qualifications
  • Familiarity with REST/Redfish and infrastructure-as-code concepts.1024-GPU-class lab operations, Supermicro systems, Cisco UCS, NVIDIA platforms, Vast storage, and Cisco switching.
  • NVIDIA BCM, Cisco network automation, storage automation, and GPU server platforms such as Supermicro and Cisco UCS.
  • Experience automating multi-plane/ToR network designs and large-scale cluster lifecycle operations.
  • Familiarity with CI/CD, configuration management, logging, and telemetry systems.
Why Cisco?

At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.

Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere.

We are Cisco, and our power starts with you.

Get your free, confidential resume review.

or drag and drop your file here.