Director of Infrastructure & Reliability

PracticeSuite, Inc.

Tampa (FL)

On-site

USD 160,000 - 230,000

Full time

22 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

PracticeSuite, Inc. is seeking a founding head of Infrastructure & Site Reliability Engineering to establish the SRE function, set standards, and hire the first engineer. You’ll define the operating model and roadmap from day one for a healthcare-focused platform.

You’ll lead incident response, design AWS-based cloud infrastructure (accounts, VPC, IAM), implement CI/CD, observability, backups, and cost controls while shaping reliability, scalability and performance.

Qualifications

  • 3+ years leading platform and/or infrastructure teams.
  • Led an SRE, platform, or infra team at director/manager level and hands-on in production.
  • Experience with multi-server architecture for redundancy, scale, performance.
  • Production experience with Linux, Oracle and Java stack.
  • Hands-on AWS core services design including cost, scaling, and resilience.
  • Familiarity with hypervisor concepts (KVM/VMware) and Proxmox is a plus.
  • Experience implementing monitoring/observability and CI/CD.
  • Experience with IAM, encryption, patching, vulnerability management; HIPAA/SOC 2/HITRUST/ISO 27001 preferred.
  • Experience collaborating with production teams.

Responsibilities

  • Stand up SRE as a function: incident command, US-hours on-call coverage, post-incident reviews, and SLOs
  • Set infrastructure direction across current servers and new solutions including vendor recommendations
  • Drive Cloud Infrastructure: Design AWS accounts, VPC, IAM, and data-store shape; own cost, scaling, and resilience trade-offs
  • Own CI/CD and operational foundations: environment and domain management, observability, secrets management, access control, and backup/restore evidence
  • Define the operating split between Server Admin and SRE

Skills

SRE leadership
Cloud architecture
AWS core services
Linux administration
CI/CD implementation
Observability/Monitoring
IAM / Access control
Security compliance
Production systems design
Cross-functional leadership

Tools

Proxmox
VMware
KVM

Job description

Department: Engineering & System Operations

About the Role

Opportunity to lead Infrastructure and site reliability engineering with authority around DevOps as well. You'll take a platform that's been reliably running critical healthcare operations and evolve it into a modern, resilient, forward-looking infrastructure organization with the autonomy to set the technical standards and the roadmap from day one.

This is a founding leadership role: you'll make your first hire, help that person build their own team, and establish the operating model that the entire infrastructure function will run on for years to come. If you've always wanted to build an SRE function the right way with executive backing, room to define what "good" looks like, and a real chance to leave your mark on how a company operates this is that opportunity.

Responsibilities
  • Stand up SRE as a function: incident command, US-hours on-call coverage, post-incident reviews, and SLOs
  • Set infrastructure direction: across current servers, and new solutions including vendor recommendations
  • Drive Cloud Infrastructure: Design and implement AWS accounts, VPC, IAM, and data-store shape. Own cost, scaling, and resilience trade-offs.
  • Own CI/CD and operational foundations: environment and domain management, observability, secrets management, access control, and backup/restore evidence
  • Define the operating split between Server Admin (running the current estate) and SRE (reliability, toil reduction, and the new platform)
Requirements
  • 3+ years leading platform and/or infrastructure teams
  • Built or led an SRE, platform, or infrastructure team at Director or Manager level, and still hands-on: you have recently designed and implemented production systems, not only managed people who did
  • Experience working on horizontal multi-server architecture to deliver redundancy, scale, reliability and performance
  • Production experience in Linux, Oracle and Java tech stack
  • Hands-on AWS core services (EC2, VPC, IAM, S3, RDS, CloudWatch): you can design accounts, networks, and IAM, and reason about cost, scaling, and resilience.
  • Solid hypervisor/VM concepts (KVM, VMware, or similar); Proxmox familiarity a plus
  • You have implemented (not only selected) monitoring/observability and CI/CD
  • Hands-on access control/IAM, encryption, patching, vulnerability management; HIPAA, SOC 2, HITRUST, or ISO 27001 strongly preferred
  • Experience leading or closely partnering with production teams based
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Staffing Science • Arizona

On-site
USD 180,000 - 240,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Practice by Numbers • United States

On-site
USD 120,000 - 160,000
High ownership and autonomy
Strong engineering culture
Impactful work on healthcare infrastructure
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GoGuardian • El Segundo (CA)

Hybrid
USD 180,000 - 240,000
Director Platform Engineering - SRE / Observability
Director Platform Engineering - SRE / Observability

Request Technology, LLC • Chicago (IL)

On-site
USD 180,000 - 240,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Good co India • United States

Remote
USD 120,000 - 160,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Calance • United States

On-site
USD 150,000 - 200,000
Director of Site Reliability Engineering & Service Enablement
Director of Site Reliability Engineering & Service Enablement

ServiceNow • Santa Clara (CA)

On-site
USD 260,000 - 360,000
Generous family leave
Matched donations
Annual learning stipends
+3
Lead SRE
Lead SRE

JPMorgan Chase & Co. • Plano (TX)

On-site
USD 150,000 - 190,000