Site Reliability Engineer

Longbridge

Dallas (TX)

On-site

USD 120,000 - 170,000

Full time

15 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Longbridge is seeking a hands-on Site Reliability Engineer to design, scale, and safeguard the reliability of its next-generation financial platforms as part of a global expansion. You will partner with product and engineering teams across the globe to ensure highly available, secure systems from day one.

In this fintech/tech startup, you’ll build automation, lead incident response, and influence the technical direction while shaping the reliability foundation for our U.S.

Qualifications

  • 5+ years of experience in SRE, DevOps, or production engineering.
  • Strong background in AWS or GCP/Azure and container orchestration (Docker, Kubernetes).
  • Proficiency in Python or Go for automation and tooling.
  • Solid Linux administration and experience with CI/CD pipelines.
  • Proven incident-management and troubleshooting of distributed systems.
  • Strong collaboration and communication across remote/global teams.
  • Experience in fast-moving fintech/tech startup environment.
  • Proficiency in Mandarin and English for international collaboration.

Responsibilities

  • Own system reliability: Design, implement, and operate highly available, secure distributed systems to meet uptime and performance targets.
  • Build automation at scale: Monitoring, alerting, and infrastructure-as-code (Terraform, Ansible, Helm).
  • Partner globally: Work with development teams from design through deployment, ensuring reliability is built in from day one.
  • Lead incident response: Drive on-call processes, root-cause analysis, and reduce MTTR and failure recurrence.
  • Future-proof our stack: Evaluate and adopt modern cloud-native technologies (Kubernetes, Prometheus, AWS/GCP).
  • Stress-test and safeguard: Lead disaster recovery, chaos testing, and capacity planning for wealth management services.

Skills

SRE/DevOps
Cloud architecture
Incident management
Automation scripting
CI/CD
Team collaboration
Fintech startup experience
English/Mandarin bilingual

Tools

Docker
Kubernetes
Terraform
Ansible
Helm
AWS
GCP/Azure

Job description

Job Description

Longbridge is a new-generation, AI-driven online brokerage on a mission to make investing smarter, simpler, and more accessible for everyone. Headquartered in Singapore, we are redefining the investment journey by connecting the stages of "Discovery → Learning → Trading." With our proprietary AI assistant, Longbridge AI, and a cloud-native infrastructure, we provide retail investors with institutional-grade insights and a seamless global trading network. At Longbridge, you won’t just be working for a brokerage; you’ll be building the future of financial infrastructure.

As part of our global expansion, we’re looking for a hands‑on Site Reliability Engineer (SRE) to design, scale, and safeguard the reliability of our next‑generation financial platforms. This is a high‑impact role where you’ll partner closely with product and engineering teams across the globe.

  • Own system reliability: Design, implement, and operate highly available, secure distributed systems to meet strict uptime and performance targets.
  • Build automation at scale: Develop and enforce best practices in monitoring, alerting, and infrastructure-as-code (e.g., Terraform, Ansible, Helm).
  • Partner globally: Work with development teams from design through deployment, ensuring reliability and resiliency are built in from day one.
  • Lead incident response: Drive on-call processes, conduct root-cause analysis, and continuously reduce MTTR and failure recurrence.
  • Future-proof our stack: Evaluate and adopt modern cloud-native technologies (e.g., Kubernetes, Prometheus, AWS/GCP) to keep systems secure and scalable.
  • Stress-test and safeguard: Lead disaster recovery, chaos testing, and capacity planning for critical wealth management services.
What We’re Looking For
  • 5+ years of experience in SRE, DevOps, or production engineering roles.
  • Strong background in AWS (or GCP/Azure) and container orchestration (Docker, Kubernetes).
  • Proficiency in at least one programming language (Python, Go, or similar) for automation and tooling.
  • Solid Linux administration skills and experience with CI/CD pipelines.
  • Proven ability in incident management and troubleshooting distributed systems.
  • Strong collaboration and communication skills across remote/global teams.
  • Comfortable working in a fast-moving fintech/tech startup environment.
  • Proficiency in Mandarin and English at the business communication level for international team collaboration.
Why Join Us
  • Shape the reliability foundation of our U.S. product launch in wealth tech.
  • Opportunity to build systems from the ground up and influence technical direction.
  • Competitive compensation package and growth opportunities.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Longbridge Singapore • New York (NY)

On-site
USD 140,000 - 190,000
Competitive compensation
Growth opportunities
Site Reliability Engineer
Site Reliability Engineer

Longbridge • New York (NY)

On-site
USD 130,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Longbridge Singapore • Dallas (TX)

On-site
USD 120,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Longbridge Securities • Town of Texas (WI)

On-site
USD 100,000 - 130,000
Competitive compensation package
Growth opportunities
Site Reliability Engineer (Remote - United States)
Site Reliability Engineer (Remote - United States)

Longbridge Singapore • California (MO)

On-site
USD 90,000 - 120,000
Medical insurance
Vision insurance
401(k)
+1
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas

The Goldman Sachs Group • Dallas (TX)

On-site
USD 110,000 - 140,000
Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas
Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas

The Goldman Sachs Group • Dallas (TX)

On-site
USD 180,000 - 280,000
Senior Site Reliability Engineer - FinTech Infra at Scale
Senior Site Reliability Engineer - FinTech Infra at Scale

Longbridge Singapore • Dallas (TX)

On-site
USD 120,000 - 190,000
Senior SRE — Fintech Reliability & Cloud Automation
Senior SRE — Fintech Reliability & Cloud Automation

Longbridge Singapore • New York (NY)

On-site
USD 140,000 - 190,000
Competitive compensation
Growth opportunities
Site Reliability Engineer
Site Reliability Engineer

Stott and May • New York (NY)

On-site
USD 120,000 - 140,000