Senior/Staff Site Reliability Engineer - Data Center

PathAI

Boston

Hybrid

PHP 9,187,000 - 14,133,000

Full time

12 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

PathAI is seeking a Senior/Staff Site Reliability Engineer to design, build, and operate our on-prem/cloud data center for a rapid-growing ML team. You will implement SRE best practices, manage hybrid cloud integration, and ensure high availability with automation, monitoring, and secure NIST/ISO-compliant environments.

This hybrid role supports Boston-area data centers and occasional travel, with relocation not offered. Join a team transforming pathology with AI-powered healthcare solutions.

Qualifications

  • BS in Computer Science, Computer Engineering, Electrical Engineering, Software Engineering or related field.
  • 8 years of experience in physical hardware/facilities, networking, automation or related areas.
  • Experience with modern datacenter network designs and operating across network layers.
  • Administered physical hardware stacks in production (iDRAC/IPMI/Nvidia UFM/Juniper).
  • Experience with virtualization, containerization, or orchestration (EKS-Anywhere/ClusterAPI/KVM).
  • Expertise in storage solutions for high-performance workloads (Quobyte, S3, FSx, EFS).
  • Proficient with automation tools; automate using scripting and configuration management (Ansible/RedFish).
  • Built monitoring with Datadog, Grafana, Prometheus.
  • Operational background managing critical production systems, incident response, scaling, and high-growth environments.
  • Willingness to travel to onsite Data Center locations.

Responsibilities

  • Advance operations by implementing SRE best practices focused on users, monitoring, and automation.
  • Design, build, and operate our data center for our ML team.
  • Build secure on-premises environments meeting NIST/ISO standards.
  • Integrate on-prem data centers with cloud infrastructure for hybrid cloud.
  • Improve reliability and resilience via root-cause analysis and design reviews.
  • Participate in platform on-call rotations and assist incident response.

Skills

SRE experience
Data center operations
Networking
Automation scripting
Observability tooling
Incident response
Travel to on-site data centers

Education

BS in Computer Science/Engineering

Tools

iDRAC/IPMI
NVIDIA UFM
Juniper networks
KVM
EKS-Anywhere/ClusterAPI
Ansible/RedFish
Datadog/Grafana/Prometheus
Quobyte/S3/FSx/EFS

Job description

Our team is passionate about solving big challenges in healthcare and transforming the field of pathology with artificial intelligence.

Senior/Staff Site Reliability Engineer - Data Center

PathAI's mission is to improve patient outcomes with AI-powered pathology.

PathAI is transforming traditional pathology methods into powerful, new technologies. These innovations in pathology can help accelerate drug development, improve confidence in the accuracy of diagnosis, and get life-saving therapies to patients more quickly. At PathAI, you'll work with a diverse and talented team of people, who are dedicated to solving complex problems and making a huge impact.

We are expanding our team and recruiting for a skilled Senior/Staff Site Reliability Engineer focused on designing, building, and operating our on-prem/cloud environment.

The Opportunity
  • You will advance the state of our operations by implementing SRE best practices - focusing on users, monitoring, and automation.
  • You will design, build and operate our data center to support our rapidly growing Machine Learning team.
  • You will build highly-secure on-premises environments handling NIST/ISO standards.
  • You will integrate on-premises datacenter environments with existing cloud infrastructure to create a seamless hybrid cloud environment.
  • You will improve the reliability and resilience of our infrastructure through root-cause analysis and reviewing gaps in designs, and implementations of our infrastructure.
  • You will participate in platform on-call rotations and assist with urgent incident response.
Who You Are: (Required)
  • You have a BS in Computer Science, Computer Engineering, Electrical Engineering, Software Engineering or closely related technical field.
  • You have 8 years experience working in physical hardware/facilities, networking, automation or other relevant areas.
  • You have demonstrated experience with modern datacenter network designs and comfort operating across network layers.
  • You’ve administered physical hardware stacks in production settings (iDRAC/IPMI/Nvidia UFM/Juniper Systems).
  • You have demonstrated experience and opinions on virtualization, containerization, or container orchestration platforms. (EKS-Anywhere/ClusterAPI/KVM).
  • You have strong expertise in storage solutions and optimizing them for high-performance workloads (e.g., Quobyte, S3, FSx, EFS).
  • You are highly proficient with automation tools; you eliminate toil by automating everything through scripting, configuration management tools (Ansible/RedFish).
  • You’ve built monitoring infrastructure with modern observability tools (Datadog/Grafana/Prometheus).
  • You have a proven operational background managing critical production systems, with extensive experience in incident response, infrastructure scaling, and navigating high-growth challenges.
  • You have the ability to travel to onsite Datacenter location(s) as needed.

Preferred:

  • You have outstanding interpersonal, verbal, and written communication and influencing skills: have built and cultivated important relationships both inside and outside of the organization and externally; have proven abilities to influence internal partners and stakeholders, thought leaders, national advocacy organizations, national standard-setting bodies, and other relevant external parties.
  • You have strong analytical and critical thinking skills with attention to detail; you have the ability to manage multiple projects and drive results in a fast-paced environment; you have a collaborative mindset with demonstrated leadership capabilities..

This is a hybrid position based in Boston (preferred), NYC, Indianapolis. (A remote option may be considered for an exceptional candidate.)
Relocation benefits are not available for this position.

The expected salary range for this position based on the primary location Boston, MA is $146,250 - $225,000. Actual pay will be determined based on experience, qualifications, geographic location, and other job-related factors permitted by law.

PathAI is an equal opportunity employer. It is our policy and practice to employ, promote, and otherwise treat any and all employees and applicants on the basis of merit, qualifications, and competence. The company's policy prohibits unlawful discrimination, including but not limited to, discrimination on the basis of Protected Veteran status, individuals with disabilities status, and consistent with all federal, state, or local laws.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Center SRE: Hybrid Cloud & On-Prem Ops
Senior Data Center SRE: Hybrid Cloud & On-Prem Ops

PathAI • Boston

Hybrid
PHP 9,187,000 - 14,133,000
Assistant Controller
Assistant Controller

PathAI • Boston

Hybrid
PHP 8,150,000 - 14,086,000
Not Overtime Eligible
Hybrid position based in Boston, MA
Relocation benefits not available
Senior Software Engineer
Senior Software Engineer

Phia LLC • Boston

Hybrid
PHP 4,478,000 - 9,035,000
Medical, Dental & Vision benefits
401K retirement plan
Disability insurance
+5
Technical Program Manager, Science
Technical Program Manager, Science

Basecamp Research • Boston

On-site
PHP 9,864,000 - 13,564,000
Senior Clinical Data Engineer (LATAM)
Senior Clinical Data Engineer (LATAM)

Precision Medicine Group • Mexico

On-site
PHP 7,524,000 - 11,285,000
Full-Stack Software Engineer: ML Focus
Full-Stack Software Engineer: ML Focus

Boston Bioprocess • Hinoba-an

On-site
PHP 938,000 - 1,876,000
Health insurance
Vision insurance
Dental insurance
+1
Senior Software Engineer (Full Stack)
Senior Software Engineer (Full Stack)

Prudentia Sciences • Boston

Hybrid
PHP 9,404,000 - 14,420,000
Impact
Ownership
Growth
+2
Principal Data Engineer
Principal Data Engineer

CodaMetrix • Boston

Hybrid
PHP 10,822,000 - 12,369,000
Health Insurance
401(k) plan
Paid Time Off
+4
Senior Software Engineer - Enterprise AI
Senior Software Engineer - Enterprise AI

3M HEALTHCARE • Boston

Hybrid
PHP 8,620,000 - 11,700,000
Senior Manager, Engineering
Senior Manager, Engineering

Prudentia Sciences • Boston

Hybrid
PHP 11,314,000 - 16,342,000