Azure Software Design Engineer – Infrastructure, Quality & Live-Site

Aptly Technology Corp.

India

Remote

INR 1,500,000 - 2,800,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Apty Technology Corp. is seeking an experienced Azure Software Design Engineer with 5-7 years of relevant software engineering, cloud infrastructure, and live-site operations expertise.

The role focuses on Azure infrastructure software, hardware diagnostics, telemetry, RCA, and fault attribution in large-scale production environments. You'll develop backend services and diagnostics using C#, Python, PowerShell, C++, or Rust, and own SKU qualification, incident triage, and recovery workflows

Qualifications

  • 5-7 years of relevant professional experience in software engineering, cloud infrastructure, systems engineering, SRE, production engineering, or infrastructure diagnostics.
  • Strong software engineering fundamentals in C#, C++, Python, or PowerShell.
  • Proven experience debugging, troubleshooting, fault isolation, and root-cause analysis of distributed production systems.
  • Experience with live-site operations, incident management, triage, mitigation, and recovery workflows.
  • Experience with telemetry, logging, diagnostics, monitoring, observability, and fault-attribution systems.
  • Experience with Git and CI/CD workflows.
  • Strong communication and cross-functional collaboration skills.

Responsibilities

  • Develop and maintain backend services, diagnostics, and infrastructure tools using C#, .NET, Python, PowerShell, C++, or Rust.
  • Execute and monitor hardware SKU qualification and host networking pipelines, investigate failures, and rerun failed scenarios.
  • Perform live-site incident triage, debugging, fault isolation, RCA, mitigation, and recovery.
  • Analyze flaky tests and recurring failures and drive improvements in test coverage, reliability, scale, and execution efficiency.
  • Monitor and troubleshoot HDTest and hardware diagnostic setup/execution failures.
  • Develop and support hardware health monitoring, fault attribution, diagnostics, certification, remediation, and recovery workflows.
  • Enhance telemetry, metrics, observability, monitoring, dashboards, and operational documentation.
  • Support hardware-adjacent systems involving BMC, firmware, server management, provisioning, inventory, deployment, and device operations.
  • Develop and maintain CI/CD pipelines using Azure DevOps, GitHub Actions, or equivalent technologies.
  • Manage source code through Git, branching, pull requests, code reviews, and release workflows.
  • Develop automation using C#, Python, and PowerShell to improve operational efficiency.
  • Collaborate with software, hardware, networking, quality, PM, and datacenter operations teams.

Skills

C#
Python
PowerShell
C++
Rust

Tools

Azure DevOps
GitHub Actions

Job description

Experience: 5-7+ Years
Domain: Azure Cloud Infrastructure | Software Engineering | Diagnostics | Quality Engineering | Live-Site Operations

Ideal Candidate Profile:
5-7+ years | C#/Python/PowerShell | Azure/Cloud Infrastructure | Distributed Systems | Live-Site & RCA | Hardware Diagnostics | Telemetry & Observability | Git/CI/CD | AI/HPC preferred

Pan India - remote role
Experience- 4+ works
Skills - C# .net, Azure Infrastructure

Role Overview

We are looking for an experienced Azure Software Design Engineer with 5-7 years of relevant experience in software engineering, cloud infrastructure, distributed systems, or production engineering. The role focuses on developing and supporting Azure infrastructure software, hardware diagnostics, SKU qualification, telemetry/observability, and live-site operations. The ideal candidate will have strong debugging and root-cause analysis skills and be comfortable troubleshooting complex software, hardware, networking, and infrastructure issues in large-scale production environments.

  • Develop and maintain backend services, diagnostics, and infrastructure tools using C#, .NET, Python, PowerShell, C++, or Rust.
  • Execute and monitor hardware SKU qualification and host networking pipelines, investigate failures, and rerun failed scenarios.
  • Perform live-site incident triage, debugging, fault isolation, RCA, mitigation, and recovery.
  • Analyze flaky tests and recurring failures and drive improvements in test coverage, reliability, scale, and execution efficiency.
  • Monitor and troubleshoot HDTest and hardware diagnostic setup/execution failures.
  • Develop and support hardware health monitoring, fault attribution, diagnostics, certification, remediation, and recovery workflows.
  • Enhance telemetry, metrics, observability, monitoring, dashboards, and operational documentation.
  • Support hardware-adjacent systems involving BMC, firmware, server management, provisioning, inventory, deployment, and device operations.
  • Develop and maintain CI/CD pipelines using Azure DevOps, GitHub Actions, or equivalent technologies.
  • Manage source code through Git, branching, pull requests, code reviews, and release workflows.
  • Develop automation using C#, Python, and PowerShell to improve operational efficiency.
  • Collaborate with software, hardware, networking, quality, PM, and datacenter operations teams.
Required Qualifications
  • 5-7 years of relevant professional experience in software engineering, cloud infrastructure, systems engineering, SRE, production engineering, or infrastructure diagnostics.
  • Strong software engineering fundamentals in one or more of C#, C++, Rust, Python, or PowerShell.
  • Proven experience in debugging, troubleshooting, fault isolation, and root-cause analysis of distributed production systems.
  • Experience supporting large-scale cloud infrastructure, datacenter platforms, hardware management, or distributed systems.
  • Hands-on experience with live-site operations, incident management, triage, mitigation, and recovery workflows.
  • Experience with telemetry, logging, diagnostics, monitoring, observability, and fault-attribution systems.
  • Experience with hardware-adjacent systems such as BMC, firmware, server diagnostics, inventory, provisioning, or deployment.
  • Strong knowledge of Git and CI/CD workflows.
  • Ability to independently investigate complex problems in ambiguous environments and drive them through resolution.
  • Strong communication and cross-functional collaboration skills.
Preferred Qualifications
  • Experience with distributed systems, orchestration engines, stateful workflows, and scalable control planes.
  • Experience with Linux-based development and diagnostics.
  • Experience with hardware diagnostics, health monitoring, certification, remediation, or recovery solutions.
  • Knowledge of servers, networking, storage, rack-level architecture, and accelerator infrastructure.
  • Experience with fleet management, firmware deployment, provisioning, inventory, or device operations.
  • Experience with Infrastructure as Code using Terraform, ARM templates, or Bicep.
  • Experience with AI/HPC infrastructure, GPU clusters, Bare Metal Instances, or hyperscale compute environments.
  • Knowledge of RDMA, InfiniBand, and accelerator-focused infrastructure is an advantage.
  • Understanding of reliability engineering, MTTR optimization, resiliency, and operational excellence.
Azure Infrastructure Focus

Experience in the following areas is highly desirable:

  • Hardware health monitoring and fault attribution
  • AI/HPC SKU onboarding and qualification
  • GB300, VR200, MI455x, Maia300, or similar accelerator platforms
  • Bare Metal Instance diagnostics and validation
  • Device recovery and automated remediation
  • Inventory, provisioning, firmware deployment, and certification
  • Integration with Titan, SCHIE, OneMOS, APlat, and Azure infrastructure services
Preferred Certifications
  • Microsoft Azure Developer Associate
  • Microsoft DevOps Engineer Expert
  • Microsoft Azure Solutions Architect Expert
  • Relevant cloud, Linux, networking, or infrastructure certifications
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Finthrive • Gurugram District

On-site
INR 1,800,000 - 2,400,000
Azure Cloud Infrastructure Developer
Azure Cloud Infrastructure Developer

Weekday (YC W21) • Delhi

On-site
INR 1,000,000 - 1,500,000
Azure Infrastructure Lead
Azure Infrastructure Lead

Intech Systems Pvt. Ltd. • Ahmedabad

On-site
INR 1,500,000 - 2,500,000
Azure Cloud Developer
Azure Cloud Developer

VME Vhire Solutions • India

Remote
INR 1,500,000 - 2,300,000
Azure Platform Engineer
Azure Platform Engineer

Tata Consultancy Services • Chennai District, Bengaluru

On-site
INR 2,600,000 - 3,800,000
AI Infrastructure Architect
AI Infrastructure Architect

Accenture • Hyderabad

On-site
INR 1,400,000 - 2,000,000
Azure Platform Site Reliability Engineer
Azure Platform Site Reliability Engineer

Foss United • India

On-site
INR 2,800,000 - 4,600,000
Azure Infrastructure Engineer
Azure Infrastructure Engineer

Infosys Limited • Telangana

On-site
INR 1,800,000 - 3,000,000
Azure Infrastructure Engineer
Azure Infrastructure Engineer

Infosys • Hyderabad

On-site
INR 1,200,000 - 2,200,000
Senior Site Reliability Engineering (Azure Cloud)
Senior Site Reliability Engineering (Azure Cloud)

Cvent • Bengaluru

On-site
INR 3,500,000 - 6,000,000