Platform Reliability Engineer: Scale & Resilience

Squarepoint

Houston (TX)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Squarepoint in Houston, Texas is seeking a Platform Reliability Specialist responsible for ensuring stability and performance of platform services. This role involves enhancing reliability practices through software engineering and operational ownership, collaborating with developers and infrastructure teams.

Candidates should have a minimum of 4 years in relevant roles, with expertise in system administration, Python, and observability systems. The position is geared towards making significant improvements to platform operations and reliability standards.

Qualifications

  • 4+ years in SRE, Production Engineering, or Reliability Engineering roles with direct ownership of production systems.
  • Experience with system administration and troubleshooting (Linux, Bash, containers).
  • Software development experience with Python, version control (Git), and CI/CD systems.
  • Hands-on experience with observability systems including metrics, tracing, log pipelines, and alert design.
  • Demonstrated experience running systems at scale, including performance tuning, HA/DR architectures, and resilience engineering.

Responsibilities

  • Own and improve day‑to‑day platform operations by streamlining workflows.
  • Work with service owners to improve resilience and performance.
  • Build and maintain platform tools and automation.
  • Capture and share reliability knowledge through documentation.
  • Help define and evolve reliability standards across the platform.

Skills

SRE
Production Engineering
Reliability Engineering
Python
Git
CI/CD
Linux
Observability systems

Tools

Prometheus
Grafana
ELK
Kubernetes
Ansible
Terraform

Job description

Squarepoint in Houston, Texas is seeking a Platform Reliability Specialist responsible for ensuring stability and performance of platform services. This role involves enhancing reliability practices through software engineering and operational ownership, collaborating with developers and infrastructure teams.

Candidates should have a minimum of 4 years in relevant roles, with expertise in system administration, Python, and observability systems. The position is geared towards making significant improvements to platform operations and reliability standards.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Engineer - Reliability
Platform Engineer - Reliability

Squarepoint • Houston (TX)

On-site
USD 100,000 - 130,000
Platform Reliability Architect
Platform Reliability Architect

Grow Therapy • New York (NY)

Hybrid
USD 182,000 - 250,000
Comprehensive Health Coverage
Parental Leave
401(k) program
+4
Senior Platform Reliability Engineer: Scale & Observability
Senior Platform Reliability Engineer: Scale & Observability

Grow Therapy • Seattle (WA)

Hybrid
USD 182,000 - 250,000
Comprehensive health coverage
401(k) program
Flexible PTO and paid holidays
+2
Platform Reliability Architect
Platform Reliability Architect

Transformcap • San Francisco (CA)

Hybrid
USD 182,000 - 250,000
Comprehensive Health Coverage
Parental Leave & Family Support
401(k) program
+3
Senior Systems Engineer: Enterprise Platform Reliability
Senior Systems Engineer: Enterprise Platform Reliability

SquareDomain • Austin (TX)

On-site
USD 90,000 - 120,000
Senior Platform Reliability Engineer - Remote
Senior Platform Reliability Engineer - Remote

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000
Senior Platform Reliability Engineer | Scale Observability
Senior Platform Reliability Engineer | Scale Observability

Grow Therapy • San Francisco (CA)

Hybrid
USD 182,000 - 250,000
Comprehensive Health Coverage
Parental Leave & Family Support
401(k) program
+3
Senior Site Reliability Engineer — Platform & Observability
Senior Site Reliability Engineer — Platform & Observability

Jobtailor • North Carolina

On-site
USD 180,000 - 240,000
Platform Reliability & Automation Engineer II
Platform Reliability & Automation Engineer II

Expand Energy • Oklahoma City (OK)

On-site
USD 90,000 - 120,000
Senior Platform SRE: Scale & Reliability
Senior Platform SRE: Scale & Reliability

United States Digital Space LLC • United States

Hybrid
USD 103,000 - 162,000
Health insurance
Vacation and RTT
Mental health and coaching
+7