Platform Engineer - Reliability

Squarepoint

Houston (TX)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Squarepoint in Houston, Texas is seeking a Platform Reliability Specialist responsible for ensuring stability and performance of platform services. This role involves enhancing reliability practices through software engineering and operational ownership, collaborating with developers and infrastructure teams.

Candidates should have a minimum of 4 years in relevant roles, with expertise in system administration, Python, and observability systems. The position is geared towards making significant improvements to platform operations and reliability standards.

Qualifications

  • 4+ years in SRE, Production Engineering, or Reliability Engineering roles with direct ownership of production systems.
  • Experience with system administration and troubleshooting (Linux, Bash, containers).
  • Software development experience with Python, version control (Git), and CI/CD systems.
  • Hands-on experience with observability systems including metrics, tracing, log pipelines, and alert design.
  • Demonstrated experience running systems at scale, including performance tuning, HA/DR architectures, and resilience engineering.

Responsibilities

  • Own and improve day‑to‑day platform operations by streamlining workflows.
  • Work with service owners to improve resilience and performance.
  • Build and maintain platform tools and automation.
  • Capture and share reliability knowledge through documentation.
  • Help define and evolve reliability standards across the platform.

Skills

SRE
Production Engineering
Reliability Engineering
Python
Git
CI/CD
Linux
Observability systems

Tools

Prometheus
Grafana
ELK
Kubernetes
Ansible
Terraform

Job description

As a Platform Reliability Specialist at Squarepoint, you will play a critical role in ensuring the stability, performance, and day to day reliability of the shared platform services. You will work with a diverse group of stakeholders, including developers, researchers, and infrastructure teams, to maintain highly reliable systems and drive proactive improvements.

You will be responsible for reducing operational toil, improving response and learning from production issues, and evolving our reliability practices. This role blends software engineering, platform ownership, operational ownership, and long‑term architectural thinking to enhance our production systems. While you may have deep expertise in one or more areas, you will contribute across the platform.

Key areas include:
  • Operations & Toil Reduction: Own and improve day‑to‑day platform operations by streamlining workflows and enhancing on‑call ergonomics through better automations and runbooks
  • Reliability Engineering & Hardening: Work with service owners to apply engineering principles to improve resilience and performance: harden critical services against degradation and outages.
  • Tooling & Automation: Build and maintain platform tools, automation, and GitOps workflows that make it easy for teams to deploy, operate, and observe their services with minimal friction and operational overhead.
  • Knowledge & Standards: Capture and share reliability knowledge through documentation, runbooks, and post‑incident reviews. Help define and evolve reliability standards and best practices across the platform.
Required qualifications
  • 4+ years in SRE, Production Engineering, or Reliability Engineering roles with direct ownership of production systems.
  • Experience with system administration and troubleshooting (Linux, Bash, containers).
  • Software development experience with Python, version control (Git), and CI/CD systems.
  • Hands‑on experience with observability systems including metrics, tracing, log pipelines, and alert design.
  • Demonstrated experience running systems at scale, including performance tuning, HA/DR architectures, and resilience engineering.
Nice to have
  • Expertise in a modern observability stack (e.g., Prometheus, Grafana, ELK, VictoriaMetrics).
  • Experience operating enterprise platform software such as Kubernetes clusters, GitLab at scale, or Slurm environments.
  • Familiarity with messaging systems (Kafka/RabbitMQ), service discovery (Consul), and databases (PostgreSQL, ClickHouse, Redis).
  • Experience authoring runbooks, running failure/chaos experiments, and participating in DR exercises.
  • Infrastructure automation and configuration management experience (e.g., Ansible, Terraform, Puppet).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Reliability Engineer: Scale & Resilience
Platform Reliability Engineer: Scale & Resilience

Squarepoint • Houston (TX)

On-site
USD 100,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Analyticpartners • Denver (CO)

On-site
USD 110,000 - 140,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
Platform Engineer
Platform Engineer

Synergy • Chicago (IL)

On-site
USD 100,000 - 150,000
Platform Engineer
Platform Engineer

Relativity Space • Long Beach (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision coverage
401(k)
Generous parental leave
+1
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000