Experience: 5-7+ Years
Domain: Azure Cloud Infrastructure | Software Engineering | Diagnostics | Quality Engineering | Live-Site Operations
Ideal Candidate Profile:
5-7+ years | C#/Python/PowerShell | Azure/Cloud Infrastructure | Distributed Systems | Live-Site & RCA | Hardware Diagnostics | Telemetry & Observability | Git/CI/CD | AI/HPC preferred
Pan India - remote role
Experience- 4+ works
Skills - C# .net, Azure Infrastructure
Role Overview
We are looking for an experienced Azure Software Design Engineer with 5-7 years of relevant experience in software engineering, cloud infrastructure, distributed systems, or production engineering. The role focuses on developing and supporting Azure infrastructure software, hardware diagnostics, SKU qualification, telemetry/observability, and live-site operations. The ideal candidate will have strong debugging and root-cause analysis skills and be comfortable troubleshooting complex software, hardware, networking, and infrastructure issues in large-scale production environments.
- Develop and maintain backend services, diagnostics, and infrastructure tools using C#, .NET, Python, PowerShell, C++, or Rust.
- Execute and monitor hardware SKU qualification and host networking pipelines, investigate failures, and rerun failed scenarios.
- Perform live-site incident triage, debugging, fault isolation, RCA, mitigation, and recovery.
- Analyze flaky tests and recurring failures and drive improvements in test coverage, reliability, scale, and execution efficiency.
- Monitor and troubleshoot HDTest and hardware diagnostic setup/execution failures.
- Develop and support hardware health monitoring, fault attribution, diagnostics, certification, remediation, and recovery workflows.
- Enhance telemetry, metrics, observability, monitoring, dashboards, and operational documentation.
- Support hardware-adjacent systems involving BMC, firmware, server management, provisioning, inventory, deployment, and device operations.
- Develop and maintain CI/CD pipelines using Azure DevOps, GitHub Actions, or equivalent technologies.
- Manage source code through Git, branching, pull requests, code reviews, and release workflows.
- Develop automation using C#, Python, and PowerShell to improve operational efficiency.
- Collaborate with software, hardware, networking, quality, PM, and datacenter operations teams.
Required Qualifications
- 5-7 years of relevant professional experience in software engineering, cloud infrastructure, systems engineering, SRE, production engineering, or infrastructure diagnostics.
- Strong software engineering fundamentals in one or more of C#, C++, Rust, Python, or PowerShell.
- Proven experience in debugging, troubleshooting, fault isolation, and root-cause analysis of distributed production systems.
- Experience supporting large-scale cloud infrastructure, datacenter platforms, hardware management, or distributed systems.
- Hands-on experience with live-site operations, incident management, triage, mitigation, and recovery workflows.
- Experience with telemetry, logging, diagnostics, monitoring, observability, and fault-attribution systems.
- Experience with hardware-adjacent systems such as BMC, firmware, server diagnostics, inventory, provisioning, or deployment.
- Strong knowledge of Git and CI/CD workflows.
- Ability to independently investigate complex problems in ambiguous environments and drive them through resolution.
- Strong communication and cross-functional collaboration skills.
Preferred Qualifications
- Experience with distributed systems, orchestration engines, stateful workflows, and scalable control planes.
- Experience with Linux-based development and diagnostics.
- Experience with hardware diagnostics, health monitoring, certification, remediation, or recovery solutions.
- Knowledge of servers, networking, storage, rack-level architecture, and accelerator infrastructure.
- Experience with fleet management, firmware deployment, provisioning, inventory, or device operations.
- Experience with Infrastructure as Code using Terraform, ARM templates, or Bicep.
- Experience with AI/HPC infrastructure, GPU clusters, Bare Metal Instances, or hyperscale compute environments.
- Knowledge of RDMA, InfiniBand, and accelerator-focused infrastructure is an advantage.
- Understanding of reliability engineering, MTTR optimization, resiliency, and operational excellence.
Azure Infrastructure Focus
Experience in the following areas is highly desirable:
- Hardware health monitoring and fault attribution
- AI/HPC SKU onboarding and qualification
- GB300, VR200, MI455x, Maia300, or similar accelerator platforms
- Bare Metal Instance diagnostics and validation
- Device recovery and automated remediation
- Inventory, provisioning, firmware deployment, and certification
- Integration with Titan, SCHIE, OneMOS, APlat, and Azure infrastructure services
Preferred Certifications
- Microsoft Azure Developer Associate
- Microsoft DevOps Engineer Expert
- Microsoft Azure Solutions Architect Expert
- Relevant cloud, Linux, networking, or infrastructure certifications