Senior Software Engineer (AI Inference & Runtime Platform) at AZX

Matcha

Northern (KY)

Hybrid

USD 150,000 - 190,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance with dependents
Bonus eligibility
Equity
Flexible paid time off
Fully remote culture

Job summary

AZX in the United States seeks an experienced platform engineer to own the layer where AI work happens: the machines, isolation boundary, and the models running on them. You’ll write Rust in the morning, a Kubernetes controller after lunch, and a FastAPI control-plane endpoint before you go home — building the fork engine, the guest agent, and the multi-substrate model lifecycle.

We’re looking for individuals who’ve built this stack, operated it at scale, or both.

Qualifications

  • 5+ years of shipping production systems in a systems language.
  • Experience with Rust, Go, C/C++, or Zig and async runtimes.
  • Experience with Kubernetes operators, autoscaling, and node lifecycle.
  • Experience deploying open-weight LLM serving infrastructure (vLLM/SGLang).
  • Ability to threat-model isolation boundaries and secure sandbox environments.

Responsibilities

  • Manage the serving tier for open-weight models, including deployment, config, and SLOs.
  • Administer the Kubernetes layer for inference and sandbox workloads, including GPUs.
  • Own the data plane: hosted vector stores and graph stores with backup and restore.
  • Oversee sandbox runtime and host control plane: lifecycle, exec, teardown, metering, isolation threat model.
  • Direct the FastAPI control-plane services, IaC tooling, and dashboards.
  • Run the layer the backend services team builds and read their dashboards.
  • Maintain OSS posture: releases, docs, and reproducible builds.

Skills

Rust
Kubernetes
Security threat modeling
LLM serving
Distributed systems
Open-source tooling

Tools

vLLM/SGLang
Kubebuilder/CRDs
Terraform/OpenTofu
Bicep
KEDA
Karpenter
OpenTelemetry

Job description

About AZX

Our mission is to accelerate positive impact in critical industries through AI transformation. We specialize in physics-informed ML and enterprise AI solutions that directly address climate and sustainability challenges.

We’re growing quickly and already work with category-leaders in real estate (CBRE), energy (LevelTen Energy), logistics (Flexe) and utilities.

We’re a public benefit corporation, founded in 2024, and have been profitable from inception.

We work on challenges in clean energy, decarbonization, climate risk, energy systems, and global economics. We’re building our company for long-term success and aim to create the ultimate place to work for those passionate about AI and making a positive impact.

About This Role:

You will be responsible for owning the layer where AI work physically happens: the machines, the isolation boundary, and the models running on them. You'll write Rust in the morning, a Kubernetes controller after lunch, and a FastAPI control-plane endpoint before you go home — building the fork engine, the guest agent, and the multi-substrate model lifecycle. We’re looking for individuals who've built this class of stack (an inference-serving or serverless-GPU platform), operated it hard at scale, or ideally both.

Responsibilities:
  • Manage the serving tier for open-weight models: engine deployment and configuration (vLLM/SGLang), cold-start strategy, per-model SLOs, and upgrade/canary discipline.
  • Administer the Kubernetes layer for inference and sandbox workloads: operators and CRDs, autoscaling (KEDA/Karpenter-class), GPU scheduling and sharing, and node lifecycle.
  • Own the stateful data plane: hosted vector stores for semantic memory and graph stores for knowledge graphs — deployed, backed up, scaled, and recovered, with restores that are tested rather than hoped for.
  • Oversee the sandbox runtime and its host-side control plane: lifecycle, exec, snapshot/fork, teardown, metering, and the threat model of the isolation boundary.
  • Direct the FastAPI control-plane services, Terraform/OpenTofu, Bicep, and the dashboards
  • Run the layer the backend services team builds on, expect to debug into their services, and expect them to read your dashboards.
  • Manage the open-source posture: build to OSS standards and release as it matures, with reviewed PRs, real docs, and reproducible builds.
Core Qualifications
  • 5+ years of shipping production systems in a systems language. Rust is the house language, but polyglots are welcome — deep Go, C/C++, or Zig with genuine appetite for Rust counts. Async runtimes, memory-safety discipline, and debugging at the syscall boundary should be familiar territory.
  • Operated Kubernetes workloads that other people depended on — controllers or operators, scheduling, autoscaling, node lifecycle. You've been paged, and the experience changed how you build.
  • Strong ability to threat-model isolation boundaries (namespaces, cgroups, seccomp, hypervisors), including identifying what an untrusted guest could observe, forge, or exhaust, and applying security best practices for agentic execution — least privilege, no credentials in the sandbox, audit trails, and human approval on write actions.
  • Hands‑on experience deploying or operating open-weight LLM serving infrastructure (vLLM/SGLang or similar), including packaging models into reliable, metered production endpoints.
  • Performance discipline in distributed systems: you measure before you optimize, and you can tell the story of a latency you killed with the numbers attached.
  • Practical depth in some of our core stack — Rust (tokio), Python/FastAPI, Kubernetes operators (controller-runtime/Kubebuilder/CRDs), KEDA, Karpenter, GPU device plugins/DRA — with a genuine willingness to research your way into the rest.
  • Familiarity with isolation technology (Firecracker, Kata, gVisor, or comparable), secrets management and egress control (Vault/KMS-class), and hosting stateful systems (vector stores like pgvector/Qdrant, graph stores like Neo4j) with backup and failover discipline.
  • Comfort operating across cloud and GPU substrates — AWS/Azure/GCP plus managed GPU clouds — using infrastructure-as-code (Terraform/OpenTofu, Bicep) and observability tooling (OpenTelemetry).
Why AZX!
  • Be part of a fast-growing, profitable, mission-driven company with industry-leading clients tackling the massive opportunity of AI transformation in critical industries.
  • Competitive early-stage startup compensation (based on capabilities, experience, and location)
  • Bonus eligibility
  • Health insurance with meaningful coverage for dependents
  • Flexible paid time off
  • Equity
  • Fully remote culture with a cluster of teammates in Seattle
Additional Information:
  • Must be able to travel 2x/year for company summits
  • Applicants must be currently authorized to work in the United States on a full-time basis
  • We are unable to sponsor or take over sponsorship of employment visas at this time
Next Steps:

If this job sounds like a great fit but you don’t check ALL of these qualification boxes, we’d still love to hear from you!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer (Energy & Utilities) at AZX
Senior ML Engineer (Energy & Utilities) at AZX

Matcha • Northern (KY)

Hybrid
USD 150,000 - 210,000
Bonus eligibility
Health insurance
Equity
+1
Senior Software Engineer (Formal Methods & Agentic Systems) at AZX
Senior Software Engineer (Formal Methods & Agentic Systems) at AZX

Matcha • Northern (KY)

Hybrid
USD 150,000 - 230,000
Health insurance
Equity
Fully remote culture
+2
Senior Product Engineer (AI, Full-Stack) at AZX
Senior Product Engineer (AI, Full-Stack) at AZX

Matcha • Northern (KY)

Hybrid
USD 120,000 - 180,000
Health insurance
Flexible PTO
Equity
+2
Senior ML Engineer (Client Solutions) at AZX
Senior ML Engineer (Client Solutions) at AZX

Matcha • Northern (KY)

Hybrid
USD 140,000 - 200,000
Bonus eligibility
Health insurance
Flexible paid time off
+2
Senior Product Engineer (AI, Full-Stack)
Senior Product Engineer (AI, Full-Stack)

AZX • Seattle (WA)

On-site
USD 140,000 - 210,000
Health insurance
Equity
Fully remote culture
+2
Senior AI Inference & Runtime Platform Engineer
Senior AI Inference & Runtime Platform Engineer

AZX • Seattle (WA), Northern (KY)

Hybrid
USD 150,000 - 190,000
Health insurance
Equity
Fully remote culture with Seattle team
+2
Senior AI Inference & Runtime Platform Engineer
Senior AI Inference & Runtime Platform Engineer

Matcha • Northern (KY)

Hybrid
USD 150,000 - 190,000
Health insurance with dependents
Bonus eligibility
Equity
+2
Senior AI Inference & Runtime Platform Engineer (Remote)
Senior AI Inference & Runtime Platform Engineer (Remote)

Uncover • Seattle (WA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Health insurance
Bonus eligibility
Equity
+1
Solutions Engagement Manager (AI/ML)
Solutions Engagement Manager (AI/ML)

AZX • Seattle (WA)

Hybrid
USD 130,000 - 200,000
Principal AI/Ops Engineer
Principal AI/Ops Engineer

RxSense • United States

On-site
USD 160,000 - 210,000