Staff Software Engineer, Infrastructure

Stream

United States

Hybrid

USD 159,000 - 215,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Stream is seeking a Staff Software Engineer to help rebuild the platform infrastructure. The role focuses on moving from AWS to GCP, consolidating Postgres shards, and running Kubernetes clusters at scale for billions of API requests.

You will own large parts of the infrastructure, write production Go and Python, and guide post-migration tuning and capacity planning within a small, senior team. Amsterdam remote-friendly with hybrid work options in Europe.

Qualifications

  • 5+ years in infrastructure, platform, DevOps or SRE engineering, with depth in infrastructure over apps.
  • Production coding experience in Go or Python; scripting-only backgrounds are not a fit.
  • Kubernetes at meaningful production scale with past operations and migrations.

Responsibilities

  • Design, build and operate infrastructure for real-time systems with millions of connections and billions of API requests.
  • Drive Kubernetes end-to-end: cluster architecture, workload design, migrations.
  • Re-architect workloads for AWS to GCP migration focusing on cost and performance.
  • Own cloud cost and efficiency work; measure spend and utilization, show improvements.
  • Write production Go and Python for internal services, tooling and automation.
  • Lead post-migration tuning and capacity planning to align architecture with production behavior.
  • Collaborate with backend, video and moderation teams on system design and reliability.
  • Participate in on-call, incident response and root cause analysis to deliver durable fixes.

Skills

Go
Python
Kubernetes
DevOps

Tools

GCP
PostgreSQL
Redis
Terraform
Prometheus

Job description

STAFF SOFTWARE ENGINEER, INFRASTRUCTURE


THE ROLE

We are hiring a Staff Software Engineer to help rebuild the platform underneath Stream. Over the next year the infrastructure team is moving from AWS to GCP, moving onto Kubernetes, and relocating 35 to 40 Postgres shards off managed RDS to self-hosted, while the platform keeps serving billions of API requests a month. You will own parts of that outright.


This is a small, senior team without the support structures of a large organisation. You will write code most of the time and make infrastructure calls on your own. Success looks like systems that scale predictably under load, cloud spend that falls per unit of traffic, and migrations that land without incident.


This is a full-time job opening based in Amsterdam (3 days hybrid) or remote in Europe/EU.


ABOUT STREAM

Stream powers real-time Chat https://getstream.io/chat/, Video https://getstream.io/video/, Activity Feeds https://getstream.io/activity-feeds/, and AI Moderation https://getstream.io/moderation/ for billions of end-users across thousands of apps, from Strava and Bumble to eBay and Patreon. Our platform processes billions of API requests per month and supports applications with millions of concurrent users, while delivering highly reliable, low-latency services and a great developer experience.


WHAT YOU WILL DO


  • Design, build and operate infrastructure for real-time systems carrying millions of concurrent connections and billions of monthly API requests.

  • Drive Kubernetes end to end: cluster architecture, workload design and the migration of existing services. You will be designing clusters, not operating someone else's.

  • Re-architect workloads as part of the AWS to GCP migration, for cost and performance rather than a lift and shift.

  • Own cloud cost and efficiency work: find the levers, measure them against real spend and utilisation data, and show what moved.

  • Write production Go and Python: internal services, platform tooling and automation that change how product and SDK engineers deploy, observe and debug.

  • Lead post-migration tuning and capacity planning, closing the loop between the architecture you chose and what production actually does.

  • Work with backend, video and moderation engineers on system design, reliability targets and tradeoffs that cross service boundaries.

  • Take part in on-call, incident response and root cause analysis, and turn what you find into durable fixes.


WHAT WE ARE LOOKING FOR


  • 5+ years in infrastructure, platform, DevOps or SRE engineering, with clear depth in infrastructure over application development.

  • A software engineering background. You have built systems, not only configured them. Production coding experience in Go or Python. Scripting-only backgrounds are not a fit.

  • Kubernetes at meaningful production scale, past operations: you have driven cluster strategy, designed workloads, or led a migration, and you have tuned what came out the other side for cost and efficiency.

  • Cloud cost or efficiency optimisation you personally led on AWS or GCP, with an outcome you can put a number on. FinOps practice is a plus.

  • Direct experience running high-scale, high-load production systems.

  • Strong cloud fundamentals across networking, compute, storage and IAM, and the habit of asking why a system behaves the way it does instead of accepting the default.

  • Comfortable in a small team: leading a project and reviewing a PR in the same week.

  • AI tooling already in your engineering workflow. Applied use, not familiarity.


BONUS POINTS


  • Both AWS and GCP, and migration experience between providers.

  • PostgreSQL at scale: sharding, replication strategy, partitioning tradeoffs, ideally self-hosted.

  • Real-time systems: WebSockets, WebRTC, streaming or other persistent-connection workloads.

  • The wider stack: CockroachDB, Redis, Terraform, and a Prometheus-based observability stack.

  • Open source contributions to infrastructure or platform tooling.

  • Writing or talks on cloud, platform or distributed systems.

  • Formal FinOps practice, or owning cloud commitment and reservation strategy.

  • Work on developer-facing API or SDK products.


OUR STACK


  • Go, gRPC, RocksDB, Python

  • PostgreSQL, RabbitMQ

  • GCP

  • Grafana, Prometheus, ELK (Elasticsearch and Kibana)

  • Jaeger and Tempo for distributed tracing, Datadog

  • Redis, Memcached

  • Claude Code, Cursor


YOU WILL THRIVE HERE IF


  • You want infrastructure problems at a scale most engineers never touch, and the autonomy to own them.

  • You ship fast and learn fast, including when it is hectic.

  • You are self-directed and comfortable working with a globally distributed team across time zones.


YOU PROBABLY WILL NOT IF


  • You want tightly scoped tickets and step-by-step direction.

  • You need a calm, highly predictable environment.

  • You would rather wait for a defined …

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Full-Stack Engineer: Infrastructure & Distributed Systems
Senior Full-Stack Engineer: Infrastructure & Distributed Systems

3M HEALTHCARE • San Francisco (CA)

On-site
USD 140,000 - 180,000
Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal
Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal

Servicenow • San Diego (CA)

Hybrid
USD 180,000 - 230,000
Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal
Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal

Servicenow • Kirkland (WA)

Hybrid
USD 190,000 - 270,000
Lead Data Engineer
Lead Data Engineer

Stream • Boulder (CO)

Hybrid
USD 150,000 - 180,000
Paid time off 19+ days
Hybrid work (3 days in office)
Health insurance (employee)
+5
Senior Devops Engineer/lead
Senior Devops Engineer/lead

Scalence • Morristown (NJ)

Hybrid
USD 140,000 - 190,000
Senior Software Engineer, Infrastructure
Senior Software Engineer, Infrastructure

vsco39 • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior DevOps Engineer
Senior DevOps Engineer

Decca Consulting LLC • United States

Remote
USD 140,000 - 210,000
Founding Engineer - Platform
Founding Engineer - Platform

uRun • San Francisco (CA)

On-site
USD 140,000 - 180,000
Competitive salary and equity
Full health, dental, and vision coverage
401(k) retirement savings
+4
Platform Engineer - Site Reliability
Platform Engineer - Site Reliability

Auto Hauler Exchange • Rochester (MI)

On-site
USD 120,000 - 160,000
Founding Infrastructure Engineer
Founding Infrastructure Engineer

Matterhaul Inc. • San Francisco (CA)

On-site
USD 200,000 - 260,000
Equity options
Hardware + AI coding budget
Real office in SF