Platform Reliability Engineer: Scale & Resilience

Squarepoint

Houston (TX)

On-site

USD 100,000 - 130,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Squarepoint in Houston, Texas is seeking a Platform Reliability Specialist responsible for ensuring stability and performance of platform services. This role involves enhancing reliability practices through software engineering and operational ownership, collaborating with developers and infrastructure teams.

Candidates should have a minimum of 4 years in relevant roles, with expertise in system administration, Python, and observability systems. The position is geared towards making significant improvements to platform operations and reliability standards.

Qualifications

  • 4+ years in SRE, Production Engineering, or Reliability Engineering roles with direct ownership of production systems.
  • Experience with system administration and troubleshooting (Linux, Bash, containers).
  • Software development experience with Python, version control (Git), and CI/CD systems.
  • Hands-on experience with observability systems including metrics, tracing, log pipelines, and alert design.
  • Demonstrated experience running systems at scale, including performance tuning, HA/DR architectures, and resilience engineering.

Responsibilities

  • Own and improve day‑to‑day platform operations by streamlining workflows.
  • Work with service owners to improve resilience and performance.
  • Build and maintain platform tools and automation.
  • Capture and share reliability knowledge through documentation.
  • Help define and evolve reliability standards across the platform.

Skills

SRE
Production Engineering
Reliability Engineering
Python
Git
CI/CD
Linux
Observability systems

Tools

Prometheus
Grafana
ELK
Kubernetes
Ansible
Terraform

Job description

Squarepoint in Houston, Texas is seeking a Platform Reliability Specialist responsible for ensuring stability and performance of platform services. This role involves enhancing reliability practices through software engineering and operational ownership, collaborating with developers and infrastructure teams.

Candidates should have a minimum of 4 years in relevant roles, with expertise in system administration, Python, and observability systems. The position is geared towards making significant improvements to platform operations and reliability standards.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform Engineer - Reliability
Platform Engineer - Reliability

Squarepoint • Houston (TX)

On-site
USD 100,000 - 130,000
Platform Engineer II: Scalable Infra & CI/CD Reliability
Platform Engineer II: Scalable Infra & CI/CD Reliability

Octopus • Houston (TX), Northern (KY)

Hybrid
USD 120,000 - 160,000
Senior Systems Engineer: Enterprise Platform Reliability
Senior Systems Engineer: Enterprise Platform Reliability

SquareDomain • Austin (TX)

On-site
USD 90,000 - 120,000
Remote Senior Platform Reliability Engineer (SRE)
Remote Senior Platform Reliability Engineer (SRE)

Bright Vision Technologies • Plymouth (MN)

Remote
USD 100,000 - 150,000
Health insurance
Platform Reliability Engineer - TechOps & Automation
Platform Reliability Engineer - TechOps & Automation

Selby Jennings • Chicago (IL)

On-site
USD 115,000 - 165,000
Senior Platform Reliability & Automation Engineer
Senior Platform Reliability & Automation Engineer

Selby Jennings • New York (NY)

On-site
USD 140,000 - 180,000
Reliability Platform Software Engineer
Reliability Platform Software Engineer

SIG Susquehanna • Pennsylvania

On-site
USD 80,000 - 100,000
Platform Reliability Engineer: Cloud, Kubernetes & Automation
Platform Reliability Engineer: Cloud, Kubernetes & Automation

Nasdaq • New York (NY)

Hybrid
USD 100,000 - 153,000
401(k) matching
Employee Stock Purchase Program
Parental leave
+2
Senior Engineer, Observability & Platform Reliability
Senior Engineer, Observability & Platform Reliability

LPL Financial • Austin (TX)

On-site
USD 102,000 - 169,000
401K matching
Health benefits
Employee stock options
+3
Platform Reliability Engineer
Platform Reliability Engineer

Amplemarket • United States

Remote
USD 120,000 - 170,000
Health Insurance
Stock Options
Annual Company Trip
+2