Site Reliability Engineer

Portage Ventures GP Inc.

New Jersey

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive Salary & Stock Options
Health Benefits
New Hire Home-Office Setup: One-time USD $500
Monthly Stipend: USD $150

Job summary

Diagram is looking for a Site Reliability Engineer to maintain and improve our brokerage platform's reliability and observability. The role encompasses managing cloud infrastructure and databases, especially PostgreSQL, while mentoring fellow engineers on best practices.

The successful candidate will have a passion for production operations, Kubernetes, and incident response, contributing to shaping our infrastructure as we scale.

Qualifications

  • 4+ years in SRE, DevOps, Platform/Infrastructure, or backend engineering.
  • Experience operating production services on Kubernetes.
  • Solid PostgreSQL production knowledge including query plans and migrations.
  • Understanding of cloud networking fundamentals.
  • Comfortable with a modern observability stack.
  • Calm under pressure with incident response skills.
  • Proficient in Go or Python and strong communication skills.

Responsibilities

  • Operate production environments, handle incident response.
  • Define and refine SLIs/SLOs and manage error budgets.
  • Improve observability across metrics, logs, and traces.
  • Ship infrastructure in a GitOps workflow.
  • Manage PostgreSQL including performance tuning and migrations.
  • Mentor engineers on reliability and database fundamentals.

Skills

Production operations ownership
Hands-on Kubernetes experience
PostgreSQL knowledge
Cloud networking fundamentals
Incident response experience
Proficiency in Go or Python

Job description

Who We Are:

Alpaca is a US-headquartered self-clearing broker‑dealer and brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more. Our recent Series D funding round brought our total investment to over $320 million, fueling our ambitious vision.

Amongst our subsidiaries, Alpaca is a licensed financial services company, serving hundreds of financial institutions across 40 countries with our institutional‑grade APIs. This includes broker‑dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, totalling over 9 million brokerage accounts.

Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to achieve our mission of opening financial services to everyone on the planet. We're deeply committed to open‑source contributions and fostering a vibrant community, continuously enhancing our award‑winning, developer‑friendly API and the robust infrastructure behind it.

Alpaca is proudly backed by top‑tier global investors, including Portage Ventures, Spark Capital, Tribe Capital, Social Leverage, Horizons Ventures, Unbound, SBI Group, Derayah Financial, Elefund, and Y Combinator.

Our Team Members:

We're a dynamic team of 380+ globally distributed members who thrive working from our favorite places around the world, with teammates spanning the USA, Canada, Japan, Hungary, Nigeria, Brazil, the UK, and beyond! We're searching for passionate individuals eager to contribute to Alpaca's rapid growth. If you align with our core values—Stay Curious, Have Empathy, and Be Accountable—and are ready to make a significant impact, we encourage you to apply.

Your Role:

As a Site Reliability Engineer at Alpaca, you'll help keep our brokerage platform reliable, observable, and operable as we grow – working across our cloud infrastructure, Kubernetes platform, observability stack, messaging layer, and data layer. We're especially interested in candidates with strong PostgreSQL fundamentals who'd like to grow into deeper ownership of our database reliability posture: PostgreSQL sits on the trading‑critical path, and we want this person to spend a meaningful share of their time leveling it up while still being a well‑rounded SRE the rest of the week.

Things You Get To Do
  • Operate production day-to-day - oncall, incident response, postmortems, and the follow‑ups that actually close the loop.
  • Own reliability practice - define and refine SLIs/SLOs and error budgets, and help product teams live within them.
  • Strengthen our observability across metrics, logs, traces, and alerting.
  • Ship infrastructure through code in a GitOps workflow - cloud resources and Kubernetes workloads alike.
  • Look after PostgreSQL: performance tuning, schema and migration review, online migrations on large tables, HA/DR, and CDC pipelines.
  • Mentor engineers on reliability and database fundamentals through code review, design review, and pairing.
Who You Are (must-haves)
  • 4+ years in SRE, DevOps, Platform/Infrastructure, or backend engineering with significant production operations ownership.
  • Hands‑on experience operating production services on Kubernetes, and shipping infrastructure as code in a GitOps workflow.
  • Solid working knowledge of PostgreSQL in production — query plans, pg_stat_*, indexing and schema trade‑offs, and what a safe online migration looks like on a non‑trivial table.
  • Cloud networking fundamentals (VPCs, routing, L4/L7 load balancing, DNS, TLS) and comfort debugging cross‑service connectivity.
  • Comfortable with a modern observability stack and proficient with Linux at the operator level.
  • Practiced in incident response - calm under pressure, structured debugging, postmortems that drive change.
  • At least working proficiency in Go or Python, plus strong written and verbal communication.
  • Genuine interest in databases and in growing your PostgreSQL/DBA expertise.
Who You Might Be (Nice-to-Haves)
  • Deeper PostgreSQL experience: large clusters at OLTP load, online migrations on big tables, HA/DR ownership, connection pooling at scale, or change‑data‑capture pipelines.
  • Experience with typed SQL access layers in Go (e.g. pgx, gorm, sqlc).
  • Production experience with messaging systems at scale (e.g. RabbitMQ, Kafka, Redpanda).
  • Security & compliance experience in a regulated environment (SOC 2, secrets management, audit logging).
  • Familiarity with trading, brokerage, or other regulated fintech domains.
How We Take Care of You
  • Competitive Salary & Stock Options
  • Health Benefits
  • New Hire Home‑Office Setup: One‑time USD $500
  • Monthly Stipend: USD $150 per month via a Brex Card

Alpaca is proud to be an equal opportunity workplace dedicated to pursuing and hiring a diverse workforce.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alpaca • New York (NY)

On-site
USD 140,000 - 210,000
Competitive Salary & Stock Options
New Hire Home-Office Setup: USD $500
Monthly Stipend: USD $150 via Brex
Senior SRE: PostgreSQL & Cloud Reliability Lead
Senior SRE: PostgreSQL & Cloud Reliability Lead

Portage Ventures GP Inc. • New Jersey

On-site
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Doist • New York (NY)

On-site
USD 150,000 - 210,000
Health Benefits
Stock Options
Home-Office Setup
Senior Software Engineer New Markets
Senior Software Engineer New Markets

Social Leverage • Northern (KY)

Hybrid
USD 150,000 - 210,000
Stock options
Home-office setup
Monthly stipend
Senior Software Engineer - Market Data
Senior Software Engineer - Market Data

Portage Ventures GP Inc. • New Jersey

On-site
USD 150,000 - 210,000
Stock options
Health benefits
Home-office setup
+1
Senior Data Scientist
Senior Data Scientist

Alpaca • New York (NY)

Hybrid
USD 80,000 - 110,000
Health Benefits
New Hire Home-Office Setup: One-time USD $500
Monthly Stipend: USD $150 per month
Senior Software Engineer - Trading
Senior Software Engineer - Trading

Alpaca • New York (NY)

On-site
USD 180,000 - 280,000
Stock Options
New Hire Home-Office Setup
Monthly Stipend
Senior Fullstack Engineer - Operational Automations
Senior Fullstack Engineer - Operational Automations

Alpaca • New York (NY)

On-site
USD 140,000 - 210,000
Stock options
Health benefits
One-time home-office setup $500
+1
Senior Software Engineer, Quality Engineering
Senior Software Engineer, Quality Engineering

Portage Ventures GP Inc. • San Mateo (CA)

On-site
USD 180,000 - 240,000
Stock options
Health benefits
Home-office setup (USD 500)
+1
Senior Software Engineer - Clearing
Senior Software Engineer - Clearing

Alpaca • Indiana (PA)

On-site
USD 130,000 - 190,000
Stock options
Health benefits
New Hire Home-Office Setup
+1