MANAGER, SRE & PLATFORM ENGINEERING

Armor Defense Inc.

Plano (TX)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Armor Defense Inc. seeks a hands‑on Manager for SRE & Platform Engineering to lead reliability for production infrastructure, including managed security services, enterprise cloud, and MDR platforms.

You will direct incident response, define reliability strategy, and develop engineering talent in a hybrid Plano, Texas and remote environment. This role combines people management with direct operational accountability, overseeing monitoring, automation, audits, and capacity planning across

Qualifications

  • 8+ years in SRE/DevOps or Infrastructure Engineering in production environments.
  • 2+ years in leading or managing engineering teams.
  • Hands-on with Kubernetes and cloud infrastructure (AWS preferred; Azure/GCP acceptable).
  • Experience with IaC (Terraform or equivalent) and CI/CD/GitOps.

Responsibilities

  • Lead and manage the SRE and infrastructure engineering team and set technical direction.
  • Own on-call coverage, incident response, escalation, and blameless postmortems.
  • Automate infrastructure operations across CI/CD, IaC, and self-healing systems.
  • Oversee infrastructure spanning VMware/Proxmox private cloud, AWS/Azure/GCP, and hybrid envs.
  • Own monitoring, alerting, observability, dashboards, and runbooks.
  • Coordinate migrations, platform transitions, and capacity planning with leadership.

Skills

SRE/DevOps leadership
Kubernetes
Cloud infrastructure
Automation & IaC
CI/CD / GitOps
Scripting (Python, Bash, PowerShell)
Observability (Datadog, Prometheus,/in

Education

Bachelor's degree in Computer Science or equivalent

Tools

VMware vSphere
Proxmox
NSX-T
Zerto
Rubrik

Job description

Locations
  • Plano, Texas: Hybrid schedule with mandatory in-office three days on Tuesday, Wednesday, and Thursday.
  • Pune, India: Hybrid schedule with mandatory in-office four days on Monday, Tuesday, Wednesday, and Thursday.
SUMMARY

The Manager, SRE & Platform Engineering reports to the Director of Platform Engineering and provides technical leadership for Armor’s SRE and infrastructure engineering team. This position is responsible for the operational reliability, availability, and performance of Armor’s production infrastructure, including managed security services, enterprise cloud, and MDR platforms. The role exercises independent judgment and discretion in directing incident response, establishing reliability strategy, making infrastructure architecture decisions, and developing engineering talent. This is a hands‑on technical leadership position that combines people management with direct operational accountability.

ESSENTIAL DUTIES AND RESPONSIBILITIES
  • Lead, hire, develop, and manage the SRE and infrastructure engineering team, including setting technical direction, conducting performance evaluations, and building engineering capability.
  • Own operational coverage including on-call rotations, incident command, escalation procedures, and blameless postmortem processes to drive continuous improvement and reduce incident recurrence.
  • Drive automation of infrastructure operations across CI/CD pipelines, infrastructure‑as‑code, self‑healing systems, and automated remediation to systematically reduce manual operational burden.
  • Manage and improve infrastructure spanning VMware/Proxmox private cloud, public cloud platforms (AWS, Azure, GCP), and hybrid environments.
  • Own monitoring, alerting, and observability across the production environment, including designing actionable dashboards, meaningful alert thresholds, and operational runbooks.
  • Plan and execute infrastructure migrations, platform transitions, and hardware refresh programs in coordination with the Director of Platform Engineering, ensuring minimal customer disruption.
  • Maintain compliance and audit readiness (PCI‑DSS, HIPAA, SOC 2) as an integrated operational discipline across all infrastructure operations.
  • Define and track SLIs/SLOs that reflect customer impact. Systematically identify and reduce operational toil.
  • Implement AI‑assisted operations including automated triage, root cause analysis, predictive alerting, and intelligent escalation to improve mean time to resolution.
  • Coordinate with the Director of Platform Engineering and Product Engineering teams on release readiness, production handoffs, change management, and capacity planning.
REQUIRED SKILLS
  • 8+ years of experience in SRE, DevOps, or Infrastructure Engineering in production environments, including 2+ years leading or managing engineering teams.
  • Hands‑on production experience with Kubernetes and container orchestration, cloud infrastructure (AWS preferred; Azure or GCP acceptable), Infrastructure as Code (Terraform or equivalent), and CI/CD/GitOps practices.
  • Strong automation and observability experience using scripting (Python, Bash, or PowerShell) and monitoring platforms such as Datadog, Prometheus, Grafana, or equivalent.
PREFERRED SKILLS
  • VMware vSphere, Proxmox, NSX‑T, Zerto, and Rubrik (expected to learn within first six months).
  • Advanced networking (DNS, VPNs, firewalls, load balancing).
  • Git proficiency and AI‑assisted development tools (GitHub Copilot, Claude Code, or similar).
  • Bachelor's degree in Computer Science, Information Technology, or equivalent experience.
  • Experience operating in regulated environments with compliance and audit responsibilities (PCI‑DSS, HIPAA, SOC 2, or similar), including supporting infrastructure through audit readiness and compliance activities.
WORK ENVIRONMENT

The work environment characteristics described here are representative of those an employee encounters while performing the essential functions of this job. The noise level in the work environment is usually low to moderate. The work environment can be either in an office setting or remotely from anywhere.

Equal opportunity employer

It is the policy of the company to comply with all employment laws and to afford equal employment opportunity to individuals in all aspects of employment, including in selection for job opportunities, without regard to race, color, religion, sex, national origin, age, disability, genetic information, veteran status, or any other consideration protected by federal, state or local laws.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II
Site Reliability Engineer II

Kastle Systems • Orlando (FL)

On-site
USD 100,000 - 120,000
Equal Opportunity Employer
Innovative Workplace Culture
Product Manager
Product Manager

Armor Defense • Plano (TX)

Hybrid
USD 110,000 - 170,000
Site Reliability Engineer II
Site Reliability Engineer II

Kastle Systems • Raleigh (NC)

On-site
USD 100,000 - 130,000
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

Calabrio • United States

On-site
USD 90,000 - 130,000
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

R&D • United States

On-site
USD 90,000 - 140,000
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Palo Alto Networks • United States

On-site
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

Fortress Information Security, LLC • Patuxent Highland (MD)

Hybrid
USD 160,000 - 180,000
Medical, dental, and vision plans
401(k) match
Flexible Paid Time Off
+1
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Platform Engineer
Platform Engineer

Relativity Space • Long Beach (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision coverage
401(k)
Generous parental leave
+1
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

HTEC Group • United States

Hybrid
USD 90,000 - 130,000