We are seeking a highly skilled and experienced Senior DevOps Engineer to join McKesson Specialty Health Technology Products LLC. This role is heavily focused on application support, production stability, and operational excellence in a fast‑paced healthcare environment. The ideal candidate will play a critical role in supporting and optimizing CI/CD pipelines, cloud infrastructure, and application platforms, ensuring high availability, reliability, and performance across multiple environments. You will work closely with Development, QA, Database, and SRE teams to support the full software development lifecycle (SDLC), production operations, and release management. This role requires strong experience in production support, incident management, and proactive monitoring, along with a clear understanding of the criticality of healthcare production systems.
Key Responsibilities
- Application & Production Support
- Provide L2/L3 support for applications running in Azure AKS Kubernetes clusters.
- Actively monitor production systems using Dynatrace, Azure Monitor, and other observability tools.
- Handle on‑call support, respond to incidents, perform troubleshooting, and drive root cause analysis (RCA).
- Ensure system stability, performance, and quick resolution of production issues.
- CI/CD & Release Management
- Design, maintain, and support CI/CD pipelines using GitHub Actions, Jenkins, and Azure DevOps.
- Support release deployments across environments (DEV, QA, PROD).
- Collaborate with teams to improve release cycles, deployment reliability, and rollback strategies.
- Ensure smooth coordination during releases with Dev, QA, and DB teams.
- Cloud & Platform Operations
- Manage and support Azure cloud infrastructure, including AKS, networking, and storage.
- Work with Kubernetes (AKS) for application deployment, scaling, and troubleshooting.
- Support service mesh (Istio) configurations and traffic management.
- Automation & Infrastructure
- Implement automation using scripting (Python, Bash, PowerShell, Terraform) to improve operational efficiency.
- Support Infrastructure as Code (IaC) practices where applicable.
- Automate repetitive tasks through Agentic AI.
- Monitoring, Logging & Security
- Utilize tools such as Dynatrace, Azure Monitor, and logging platforms to ensure observability.
- Integrate and support security and code quality tools such as Wiz CLI and SonarQube.
- Proactively identify potential risks and ensure compliance with enterprise security standards.
- Cross‑Team Collaboration
- Work closely with Development, QA, Database, and SRE teams to resolve issues and improve system performance.
- Support troubleshooting across application layers including APIs, services, and databases.
- Participate in troubleshooting war rooms and critical incident calls.
- ServiceNow & Operational Excellence
- Manage and track incidents, changes, and service requests using ServiceNow.
- Ensure proper documentation, ticket updates, and adherence to SLA/OLA expectations.
- Contribute to runbooks, SOPs, and knowledge base documentation.
- Continuous Improvement & Mentorship
- Mentor junior engineers and improve overall DevOps and operational practices.
- Identify automation and optimization opportunities to reduce manual effort.
- Drive adoption of DevOps best practices with a focus on reliability and supportability.
Qualifications
- Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience).
- 3+ years of experience in DevOps, SRE, or application support roles.
- Strong experience in production support and incident management.
- Hands‑on experience with:
- CI/CD tools: GitHub Actions, Jenkins, Azure DevOps
- Cloud: Azure (preferred)
- Containers & Orchestration: Docker, Kubernetes (AKS)
- Service Mesh: Istio
- Monitoring: Dynatrace, Azure Monitor
- Security/Quality Tools: Wiz CLI, SonarQube
- ITSM: ServiceNow
- Strong scripting skills (Python, Bash).
- Experience supporting full SDLC, releases, and multi‑environment deployments.
- Ability to troubleshoot complex issues across applications, infrastructure, and databases.
- Strong understanding of production criticality, uptime requirements, and healthcare systems sensitivity.
- Excellent communication skills and ability to work across multiple teams.
- Willingness to participate in on‑call rotation and support critical systems.