Principal Site Reliability Engineer (SRE)

Symmetrio

United States

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Paid Time Off (Vacation, Sick & Public Holidays)

Job summary

Symmetrio is seeking a Principal Site Reliability Engineer to ensure the reliability and performance of a healthcare technology platform supporting providers across the United States. The ideal candidate will have extensive experience with AWS, cloud infrastructure, and production operations.

The role involves troubleshooting complex issues, leading incident responses, and collaborating with development teams to enhance operational excellence. A bachelor's degree and 6+ years of relevant experience are preferred.

The position offers a comprehensive benefits package, including health care, retirement plans, and paid time off.

Qualifications

  • 6+ years of hands-on experience supporting AWS environments.
  • 4+ years of experience with web applications, preferably Python/Django.
  • Experience leading production incidents and root cause analysis.

Responsibilities

  • Serve as primary technical owner for production reliability.
  • Investigate and resolve complex application and infrastructure issues.
  • Lead production incident response and coordinate cross-functional teams.

Skills

AWS-based production environments
Application troubleshooting
Production operations leadership
Cloud infrastructure expertise
Technical problem-solving

Education

Bachelor’s degree in Computer Science or related field

Tools

Terraform
AWS networking technologies
Datadog
Kubernetes

Job description

Symmetrio is recruiting a Principal Site Reliability Engineer (SRE) for our customer, a rapidly growing healthcare technology organization focused on advanced healthcare technology solutions.

This individual will play a critical role in ensuring the reliability, scalability, security, and performance of a mission-critical SaaS platform supporting healthcare providers across the United States. The ideal candidate will possess a unique blend of cloud infrastructure expertise, application troubleshooting experience, production operations leadership, and customer-facing technical problem-solving skills.

The ideal candidate will be equally comfortable investigating application-level issues, troubleshooting AWS networking and infrastructure, leading production incident response efforts, and collaborating with development teams to improve operational excellence.

Responsibilities
  • Serve as the primary technical owner for production reliability across U.S. customer environments.
  • Investigate and resolve complex issues spanning web applications, APIs, backend services, data pipelines, cloud infrastructure, and customer integrations.
  • Lead production incident response efforts, coordinating cross-functional teams to restore service and minimize customer impact.
  • Perform root cause analysis and drive corrective actions that improve long-term system stability and resilience.
  • Partner with software engineering and platform teams to identify recurring reliability risks and implement sustainable solutions.
  • Design, configure, and validate secure customer connectivity solutions including Site-to-Site VPNs, Transit Gateway integrations, routing configurations, and secure network paths.
  • Support customer onboarding initiatives by troubleshooting connectivity challenges and ensuring consistent implementation processes.
  • Enhance platform observability through improvements in monitoring, logging, alerting, tracing, and operational dashboards.
  • Contribute to CI/CD, infrastructure automation, and deployment processes that improve release safety and operational consistency.
  • Develop operational tooling that supports incident response, troubleshooting, onboarding, and system monitoring activities.
  • Collaborate with engineering leadership to improve cloud architecture, scalability, security, and operational readiness.
  • Partner with customer-facing teams to communicate technical issues, remediation plans, and reliability improvements in a clear and effective manner.
  • Support compliance, security, and risk management initiatives within highly regulated healthcare environments.
  • 6+ years of hands‑on experience supporting and managing AWS-based production environments.
  • 4+ years of experience supporting web applications and backend services (Python/Django experience strongly preferred).
  • Experience with AWS networking technologies including VPCs, Site-to-Site VPNs, Transit Gateways, routing, NAT gateways, and security groups.
  • Strong experience with Terraform and infrastructure‑as‑code deployment practices.
  • Experience with containerized environments including ECS, Fargate, Kubernetes, or similar technologies.
  • Experience building and supporting CI/CD pipelines and release automation processes.
  • Familiarity with monitoring and observability platforms such as Datadog, CloudWatch, Sentry, Grafana, or similar tools.
  • Experience leading production incidents, outage management, and root cause analysis initiatives.
  • Exposure to Windows Server environments, Active Directory, Kerberos, and enterprise infrastructure concepts is preferred.
  • Healthcare technology, healthcare SaaS, clinical software, or other regulated industry experience is highly preferred.
  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related technical field preferred.
  • Health Care Plan (Medical, Dental & Vision)
  • Retirement Plan (401k, IRA)
  • Paid Time Off (Vacation, Sick & Public Holidays)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal SRE for Healthcare SaaS - Incident Leader
Principal SRE for Healthcare SaaS - Incident Leader

Symmetrio • United States

Hybrid
USD 120,000 - 160,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Paid Time Off (Vacation, Sick & Public Holidays)
SRE Leader
SRE Leader

Kontakt Micro-Location Sp. Z.o.o. • New York (NY)

Hybrid
USD 180,000 - 260,000
Equity in a high-growth company
Health, dental, and vision coverage
401k
+3
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Kontakt.io • New York (NY)

Hybrid
USD 200,000 - 250,000
Hybrid work 3 days/week in NYC office.
Equity in a high-growth company
Health, dental, vision insurance
+1
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Abbott • Sunnyvale (CA)

On-site
USD 90,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Oracle • Frankfort (KY)

On-site
USD 120,000 - 160,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Practice by Numbers • United States

On-site
USD 120,000 - 160,000
High ownership and autonomy
Strong engineering culture
Impactful work on healthcare infrastructure