SRE & DevOps Engineer — Platform Reliability & Automation

BTIG

San Francisco (CA)

On-site

USD 160,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

BTIG seeks a DevOps/SRE to join our technology team to improve developer velocity through production support, infrastructure evolution, and strong operational continuity. You will own critical platform systems and drive reliability, scalable services, and customer‑facing support within a financial services environment.

You will work across the full stack, gaining hands‑on exposure to monitoring, incident response, and automation as production health directly impacts trading operations.

Qualifications

  • 3–5+ years of experience in a DevOps, SRE, platform engineering, or production support role.
  • Strong Linux/Unix systems administration fundamentals.
  • Hands‑on experience with containerization and container orchestration.
  • Experience with database administration (relational and/or document stores).
  • Familiarity with reverse proxy configuration, TLS/certificate management, and identity systems.
  • Experience with monitoring, logging, and alerting tools in a production environment.
  • Comfort working directly with end users or customers in a support capacity.
  • Strong troubleshooting and diagnostic skills across application and infrastructure layers.
  • Scripting and automation proficiency (Bash, Python, PowerShell, or similar).
  • Self‑directed; comfortable operating with autonomy in a small, high‑output team.
  • Strong communicator – can translate between technical and non‑technical stakeholders.
  • Ownership mentality – treats production health as a personal responsibility.
  • Intellectually curious; eager to learn new domains and expand scope.
  • Must be authorized to work full time in the U.S.; BTIG does not offer sponsorship for work visas of any type.

Responsibilities

  • Serve as the primary point of contact for production application escalations – triage, diagnosis, and resolution across application components and services.
  • Monitor application and infrastructure health; investigate anomalies and remediate without pulling developers off feature work.
  • Develop and maintain runbooks, escalation procedures, and operational knowledge base documentation.
  • Identify recurring issues and collaborate with developers to drive root‑cause fixes.
  • Own incident response and post‑incident review processes.
  • Manage, maintain, and improve the team's infrastructure stack: reverse proxy & traffic management, IAM, secret/config management, certificate management, databases, OTEL/metrics/logs, container orchestration, event streaming.
  • Automate provisioning, deployment, and configuration management.
  • Plan and execute upgrades, patches, security hardening, and capacity management.
  • Evolve infrastructure toward greater reliability, scalability, and developer self‑service.
  • Own the observability stack – monitor, alert, dashboard, and centralized logging to detect and diagnose production issues quickly.
  • Work with developers to ensure applications emit useful metrics, logs, and trace context.
  • Serve as backup to customer support during PTO/peak periods.
  • Develop familiarity with customer workflows and resolution paths.
  • Contribute to support documentation and self‑service tooling to reduce support burden.

Skills

Linux/Unix
Containerization
Container orchestration
Database administration
TLS/certificate mgmt
Monitoring/Logging/Alerting
Scripting (Bash/Python/PowerShell)
Troubleshooting
Autonomy
Communication

Job description

BTIG seeks a DevOps/SRE to join our technology team to improve developer velocity through production support, infrastructure evolution, and strong operational continuity. You will own critical platform systems and drive reliability, scalable services, and customer‑facing support within a financial services environment.

You will work across the full stack, gaining hands‑on exposure to monitoring, incident response, and automation as production health directly impacts trading operations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE/DevOps Engineer for Global Trading Platform
SRE/DevOps Engineer for Global Trading Platform

Goliath Partners • United States

On-site
USD 130,000 - 170,000
Systems Engineer, SRE & Production Reliability
Systems Engineer, SRE & Production Reliability

Hunter Bond • New York (NY)

On-site
USD 425,000 - 575,000
Beautiful offices
SRE BizOps: CI/CD, Observability & Fintech Ops
SRE BizOps: CI/CD, Observability & Fintech Ops

TechDigital Group • St. Louis (MO)

On-site
USD 80,000 - 130,000
Senior SRE Engineer - Trading Platform Reliability
Senior SRE Engineer - Trading Platform Reliability

Deutsche Bank • Cary (NC)

Hybrid
USD 91,000 - 153,000
Hybrid work model
Generous vacation
ERGs
+2
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas

Socket.dev • Dallas (TX)

On-site
USD 120,000 - 180,000
Senior Platform Reliability Engineer - Remote
Senior Platform Reliability Engineer - Remote

Bright Vision Technologies • Redmond (WA)

On-site
USD 100,000 - 150,000
Real-Time SRE: Cloud Infra & Automation
Real-Time SRE: Cloud Infra & Automation

OP Recruiting • Chicago (IL)

On-site
USD 120,000 - 180,000
Medical insurance
401(k) retirement plan
Paid time off
Site Reliability Engineer
Site Reliability Engineer

Goliath Partners • United States

On-site
USD 130,000 - 170,000
Technology, DevOps/Site Reliability Engineer
Technology, DevOps/Site Reliability Engineer

BTIG • San Francisco (CA)

On-site
USD 160,000 - 200,000
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas

Goldman Sachs • Dallas (TX)

On-site
USD 120,000 - 160,000
None