Senior/Staff Software Engineer, Distributed Systems

Hedra Inc.

San Francisco (CA)

On-site

USD 150,000 - 260,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive compensation and equity
401k
Healthcare (Silver PPO Medical, Vision
Lunch and snacks at the office

Job summary

Hedra Inc. in San Francisco is seeking a Senior or Staff Software Engineer with deep expertise in building and operating distributed production systems.

You will work on Hedra’s inference infrastructure, designing and operating systems that schedule, route, and manage compute-intensive workloads across diverse resources. The role emphasizes reliability, performance, and cost efficiency, requiring you to reason from first principles and drive thoughtful tradeoffs.

Qualifications

  • 5+ years of software engineering experience, with significant experience building distributed backend or infrastructure systems.
  • A track record of designing and operating high-availability production services at meaningful scale.
  • Strong distributed systems fundamentals, including concurrency, queues, retries, failure recovery, consistency, backpressure, capacity, and load behavior.
  • Experience debugging complex production systems across multiple stack layers.
  • Strong judgment around architectural tradeoffs, particularly reliability, performance, complexity, and operational cost.
  • Experience with CI/CD, comprehensive testing and validation strategies, monitoring, alerting, and production operations.
  • Strong programming fundamentals and experience in systems-oriented backend languages; Python experience is helpful.
  • Experience using agentic coding tools and workflows beyond basic prompting.
  • Comfort operating in underspecified environments with independent decision-making.
  • Ability to communicate technical decisions clearly and collaborate across teams.

Responsibilities

  • Design, build, and operate distributed systems that power Hedra’s inference infrastructure.
  • Build systems for scheduling, routing, and managing compute-intensive workloads across heterogeneous resources.
  • Improve throughput, latency, reliability, and resource utilization across our serving infrastructure.
  • Design systems that remain predictable and resilient under load, partial failures, changing capacity, and unpredictable workloads.
  • Own production systems end to end, including deployment, CI/CD, testing and validation, observability, monitoring, alerting, debugging, and incident response.
  • Identify architectural bottlenecks and failure modes and drive solutions.
  • Work across infrastructure, model serving, APIs, and developer-facing systems as the platform evolves.
  • Use modern agentic coding workflows as part of how you design, build, debug, and ship software.
  • Help shape the technical direction of a small engineering organization where individual engineers have substantial ownership.

Skills

Distributed systems
Backend engineering
CI/CD
Python
Performance optimization
Observability
Cloud infrastructure
System design

Tools

Kubernetes

Job description

About Hedra

Hedra is the platform, models, and infrastructure for visual intelligence.

We build the systems that make large-scale visual inference fast, reliable, and accessible to developers. Our work spans model serving, compute infrastructure, scheduling and routing, APIs, and the developer platform that sits on top of it.

We’re a small, highly technical team in San Francisco, backed by a16z and other leading investors. Engineers at Hedra work across boundaries, own systems end to end, and have significant influence over both what we build and how we build it.

The Role

We’re looking for a Senior or Staff Software Engineer with deep experience building and operating distributed production systems.

You’ll work on the infrastructure underlying Hedra’s inference platform: systems that schedule and route compute, serve models efficiently, handle high-throughput workloads, and remain reliable as both traffic and the number of models we support grow.

The problems are often ambiguous and don’t have obvious answers. We’re looking for someone who can reason from first principles, identify bottlenecks and failure modes before they become problems, and make thoughtful tradeoffs across performance, reliability, complexity, and cost.

You do not need to come from an AI company or already be an expert in model inference. We care much more about depth in distributed systems and your ability to apply that experience to a new problem space.

What You’ll Do
  • Design, build, and operate distributed systems that power Hedra’s inference infrastructure.

  • Build systems for scheduling, routing, and managing compute-intensive workloads across heterogeneous resources.

  • Improve throughput, latency, reliability, and resource utilization across our serving infrastructure.

  • Design systems that remain predictable and resilient under load, partial failures, changing capacity, and unpredictable workloads.

  • Own production systems end to end, including deployment, CI/CD, testing and validation, observability, monitoring, alerting, debugging, and incident response.

  • Identify architectural bottlenecks and failure modes and drive solutions rather than waiting for problems to be fully specified.

  • Work across infrastructure, model serving, APIs, and developer-facing systems as the platform evolves.

  • Use modern agentic coding workflows as part of how you design, build, debug, and ship software.

  • Help shape the technical direction of a small engineering organization where individual engineers have substantial ownership.

What We’re Looking For
  • 5+ years of software engineering experience, with significant experience building distributed backend or infrastructure systems.

  • A track record of designing and operating high-availability production services at meaningful scale.

  • Strong distributed systems fundamentals, including experience reasoning about concurrency, queues, retries, failure recovery, consistency, backpressure, capacity, and system behavior under load.

  • Experience debugging complex production systems across multiple layers of the stack.

  • Strong judgment around architectural tradeoffs, particularly reliability, performance, complexity, and operational cost.

  • Experience with CI/CD, comprehensive testing and validation strategies, monitoring, alerting, and production operations.

  • Strong programming fundamentals and experience working in systems-oriented backend languages. Python experience is helpful given our stack.

  • Experience using agentic coding tools and workflows beyond basic prompting.

  • Comfort operating in an environment where problems are often underspecified and engineers are expected to independently determine the right approach.

  • Ability to communicate technical decisions clearly and collaborate with engineers across infrastructure, research, and product.

Nice to Have
  • Experience with inference or model-serving infrastructure.

  • GPU or accelerator infrastructure.

  • Compute scheduling, orchestration, or resource management.

  • High-throughput or low-latency systems.

  • Kubernetes or other cluster orchestration systems.

  • Performance optimization and profiling.

  • Developer infrastructure, APIs, SDKs, or platform engineering.

  • Experience operating infrastructure across cloud and/or bare-metal environments.

Benefits
  • Competitive compensation and equity

  • 401k

  • Healthcare (Silver PPO Medical, Vision, Dental)

  • Lunch and snacks at the office

This role is based in San Francisco, and we work together in person five days a week.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Customer Success Engineer
Customer Success Engineer

Hedra, Inc • San Francisco (CA)

On-site
USD 140,000 - 190,000
Competitive compensation and equity
Healthcare (medical, vision, dental)
401K
+1
Customer Success Engineer
Customer Success Engineer

Hedra • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive compensation and equity
Healthcare (Silver PPO Medical, Vision
401K
+1
Account Executive
Account Executive

Hedra, Inc • San Francisco (CA)

On-site
USD 120,000 - 190,000
Competitive compensation and equity
Healthcare (Silver PPO)
401K
+1
Account Executive
Account Executive

Hedra • San Francisco (CA)

On-site
USD 140,000 - 220,000
Competitive compensation and equity
Healthcare (Silver PPO)
401K
+1
Head of Marketing
Head of Marketing

Hedra • San Francisco (CA)

On-site
USD 170,000 - 250,000
Equity
Healthcare
401K
+1
Inference Optimization Engineer
Inference Optimization Engineer

Speedrun Talent Network • San Francisco (CA)

On-site
USD 150,000 - 230,000
Competitive compensation and equity
401k
Healthcare (Silver PPO Medical, Vision
+1
Inference Optimization Engineer
Inference Optimization Engineer

US Health Partners, LLC • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation and equity
401k
Healthcare (Silver PPO Medical, Vision
+1
Senior Distributed Systems Engineer for Inference Platform
Senior Distributed Systems Engineer for Inference Platform

Hedra Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Competitive compensation and equity
401k
Healthcare (Silver PPO Medical, Vision
+1
Inference Optimization Engineer
Inference Optimization Engineer

AI Chopping Block, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Equity
401k
Healthcare
+1
Product Marketing Lead
Product Marketing Lead

hedra • San Francisco (CA)

On-site
USD 120,000 - 170,000
Equity
Healthcare
401K
+1