Senior Distributed Systems Engineer for Inference Platform

Hedra Inc.

San Francisco (CA)

On-site

USD 150,000 - 260,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive compensation and equity
401k
Healthcare (Silver PPO Medical, Vision
Lunch and snacks at the office

Job summary

Hedra Inc. in San Francisco is seeking a Senior or Staff Software Engineer with deep expertise in building and operating distributed production systems.

You will work on Hedra’s inference infrastructure, designing and operating systems that schedule, route, and manage compute-intensive workloads across diverse resources. The role emphasizes reliability, performance, and cost efficiency, requiring you to reason from first principles and drive thoughtful tradeoffs.

Qualifications

  • 5+ years of software engineering experience, with significant experience building distributed backend or infrastructure systems.
  • A track record of designing and operating high-availability production services at meaningful scale.
  • Strong distributed systems fundamentals, including concurrency, queues, retries, failure recovery, consistency, backpressure, capacity, and load behavior.
  • Experience debugging complex production systems across multiple stack layers.
  • Strong judgment around architectural tradeoffs, particularly reliability, performance, complexity, and operational cost.
  • Experience with CI/CD, comprehensive testing and validation strategies, monitoring, alerting, and production operations.
  • Strong programming fundamentals and experience in systems-oriented backend languages; Python experience is helpful.
  • Experience using agentic coding tools and workflows beyond basic prompting.
  • Comfort operating in underspecified environments with independent decision-making.
  • Ability to communicate technical decisions clearly and collaborate across teams.

Responsibilities

  • Design, build, and operate distributed systems that power Hedra’s inference infrastructure.
  • Build systems for scheduling, routing, and managing compute-intensive workloads across heterogeneous resources.
  • Improve throughput, latency, reliability, and resource utilization across our serving infrastructure.
  • Design systems that remain predictable and resilient under load, partial failures, changing capacity, and unpredictable workloads.
  • Own production systems end to end, including deployment, CI/CD, testing and validation, observability, monitoring, alerting, debugging, and incident response.
  • Identify architectural bottlenecks and failure modes and drive solutions.
  • Work across infrastructure, model serving, APIs, and developer-facing systems as the platform evolves.
  • Use modern agentic coding workflows as part of how you design, build, debug, and ship software.
  • Help shape the technical direction of a small engineering organization where individual engineers have substantial ownership.

Skills

Distributed systems
Backend engineering
CI/CD
Python
Performance optimization
Observability
Cloud infrastructure
System design

Tools

Kubernetes

Job description

Hedra Inc. in San Francisco is seeking a Senior or Staff Software Engineer with deep expertise in building and operating distributed production systems.

You will work on Hedra’s inference infrastructure, designing and operating systems that schedule, route, and manage compute-intensive workloads across diverse resources. The role emphasizes reliability, performance, and cost efficiency, requiring you to reason from first principles and drive thoughtful tradeoffs.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior/Staff Software Engineer, Distributed Systems
Senior/Staff Software Engineer, Distributed Systems

Hedra Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Competitive compensation and equity
401k
Healthcare (Silver PPO Medical, Vision
+1
Fast, Scalable Vision Inference Engineer
Fast, Scalable Vision Inference Engineer

US Health Partners, LLC • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation and equity
401k
Healthcare (Silver PPO Medical, Vision
+1
Senior Distributed Systems Engineer — AI Inference Platform
Senior Distributed Systems Engineer — AI Inference Platform

Baseten • San Francisco (CA)

On-site
USD 180,000 - 240,000
Equity
Comprehensive medical for US
Flexible PTO
+4
Senior Distributed Systems Engineer, Infra & Data Platforms
Senior Distributed Systems Engineer, Infra & Data Platforms

LinkedIn • Mountain View (CA)

Hybrid
USD 144,000 - 236,000
Senior Platform Engineer - Inference Infra, Remote
Senior Platform Engineer - Inference Infra, Remote

Subconscious Systems Technologies, Inc. • Boston (MA)

On-site
USD 180,000 - 230,000
Equity ownership
Health coverage
Unlimited PTO
+1
Lead Node Systems Engineer — Frontier Inference HPC
Lead Node Systems Engineer — Frontier Inference HPC

Etched • San Jose (CA)

On-site
USD 200,000 - 260,000
Medical coverage
Dental coverage
Vision coverage
+4
Senior Inference Engineer, AI Infrastructure & Production
Senior Inference Engineer, AI Infrastructure & Production

Hamilton Barnes • United States

On-site
USD 225,000 - 275,000
Full Benefits
Staff Inference Engineer — Production-Scale AI Platform
Staff Inference Engineer — Production-Scale AI Platform

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan with company match
Paid holidays
Staff Software Engineer — Real-Time Inference Systems
Staff Software Engineer — Real-Time Inference Systems

Cerebras • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Inference Performance Engineer for Visual AI
Inference Performance Engineer for Visual AI

AI Chopping Block, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Equity
401k
Healthcare
+1