Founding AI Infrastructure Engineer

Harden

San Francisco (CA)

Hybrid

USD 150,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity

Job summary

Harden is seeking a founding AI Infrastructure Engineer to own the foundations for AI-powered workflows, from prototype to production. You will design distributed systems, build scalable Kubernetes-based infrastructure, and shape security-sensitive runtimes in customer environments.

You’ll work across data/model pipelines, observability, and API surfaces, delivering fast, reliable, and auditable platforms that integrate with customers' CI/CD, SIEM, IAM, and cloud ecosystems in a high-ownership

Qualifications

  • Required: strong production distributed systems experience.

Responsibilities

  • Design and build production services for AI/ML workloads with clear SLOs.
  • Own Kubernetes deployments end-to-end: rollout, autoscaling, networking.

Skills

Distributed systems
Kubernetes
Python
Go/Rust/C++
AWS/GCP/Azure
Performance engineering
API design
Security-aware infra

Tools

Kafka
Datadog
CI/CD tooling
Terraform

Job description

# Founding AI Infrastructure EngineerRemote (US) or Hybrid (SF Bay Area)Full-time$150,000–$250,000 + equity## About UsHarden is the enterprise security and governance platform for AI-built apps. Every AI-generated application ships with the same gaps — hardcoded secrets, missing auth, uncontrolled egress, no audit trail. Harden closes them automatically: build-time SAST and SCA scanning with auto-remediation, a one-page approval report that gates CI/CD, and runtime enforcement that covers egress controls, prompt-injection defense, AI cost caps, and tamper-evident audit logs — all deployed inside the customer's own cloud.We're a venture-backed startup based in the SF Bay Area. The founding team combines AI/ML research leadership at Google DeepMind and Amazon with product leadership at Verkada and AWS. Enterprise customers are already in production and we're bringing on exceptional engineers to build the infrastructure that makes enterprise AI adoption safe and scalable.## About the RoleYou'll own the engineering foundations for taking AI-powered workflows (including agentic systems) from prototype to reliable production. At Harden, that means building the infrastructure that analyzes, transforms, and governs AI-built applications in customer-controlled environments. This is a high-ownership, hands-on role spanning distributed systems, ML systems, Kubernetes/cloud infrastructure, performance engineering, and enterprise-ready platform design.You'll build core platform capabilities (execution/runtime, data/model pipelines, APIs, observability, scaling) and ensure the system is fast, resilient, secure, and easy to integrate into a customer's existing stack, including VPCs, CI/CD pipelines, SIEM, identity, ticketing, and compliance workflows.## What You'll Do* •Design and build production services for AI/ML workloads with clear SLOs (latency, throughput, availability), including synchronous request paths and safe fallbacks for Harden's scanning, remediation, and runtime enforcement systems* •Build a sandboxed execution runtime for running customer-provided or semi-trusted logic safely at scale (isolation boundaries, cold-start mitigation, warm pools, resource governance), including generated-code analysis environments* •Build and operate large-scale data + embedding/model pipelines (batch processing, feature generation, training data preparation, serving-friendly formats) for findings, dependency graphs, approval reports, telemetry, and audit trails* •Architect event-driven systems using Kafka-style streaming for ingestion, replayability, and decoupling offline pipelines from latency-sensitive online services* •Own Kubernetes deployments end-to-end: rollouts (blue/green, rolling), autoscaling, networking (ingress, gateways, service-to-service), resource tuning, and on-call grade debugging (OOM, crash loops, throttling)* •Build platform-grade observability: metrics, logs, traces, dashboards, alerting, and incident runbooks; instrument application-level profiling for memory/GC and performance regressions* •Implement multi-tenant API controls: API key management, quotas, rate limiting (token/leaky bucket), request scheduling/fairness, backpressure, and retry strategies with jitter* •Drive performance optimizations across the stack (parallelization, serialization/I/O bottlenecks, caching/batching), and rewrite hotspots when needed* •Build enterprise-ready integrations-first product surfaces (fit into Datadog/ServiceNow/identity/logging workflows instead of "one more dashboard"), including CI/CD, SIEM, and customer-cloud workflows* •Partner with product and customers to translate ambiguous requirements into robust, developer-friendly platform primitives for secure AI-built application deployment## What We're Looking For* •Strong experience building production distributed systems (service design, reliability patterns, scaling strategies, failure modes)* •Proven track record in performance engineering (profiling, concurrency/parallelism, bottleneck elimination, cost/perf tradeoffs)* •Deep hands-on experience with Kubernetes in production (deployments, networking, autoscaling, debugging, observability)* •Solid cloud fundamentals (AWS/GCP/Azure): compute lifecycle automation, networking, IAM basics, cost controls, and operational tooling* •Experience designing low-latency APIs and making pragmatic tradeoffs between latency, availability, and consistency* •Familiarity with streaming/event systems (Kafka/event sourcing) and building pipelines that support replay and auditability* •Strong programming skills (Python plus at least one systems language like Go/Rust/C++ is a plus)* •Comfortable operating in a fast-moving startup environment: clear communication, high ownership, and good engineering judgment* •4+ years experience in backend/platform/infra roles (or equivalent depth)* •Comfort with security-sensitive infrastructure such as sandboxed execution, tenant isolation, secrets handling, or customer-controlled cloud deployments## Nice to Have* •Experience building sandboxed runtimes or isolation layers (microVMs, gVisor, containers, secure execution boundaries)* •Built large-scale embedding/recommendation or retrieval pipelines and served them in production* •Experience with multi-tenant platform concerns: noisy neighbor mitigation, quota enforcement, fairness scheduling, per-tenant observability* •Strong opinions on enterprise integrations and "platform adoption" mechanics (connectors-first, workflow-native design)* •Experience implementing safe progressive delivery for ML-backed systems (shadowing, canarying, rollback automation, regression gating)* •Experience with infrastructure-as-code, progressive delivery, CI/CD, and production readiness practices from scratch* •Security platform, DevSecOps, or enterprise SaaS infrastructure experience
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding AI Engineer
Founding AI Engineer

Harden • San Francisco (CA)

Hybrid
USD 120,000 - 250,000
Founding Enterprise Security Engineer
Founding Enterprise Security Engineer

Harden • United States

Hybrid
USD 150,000 - 250,000
Pioneering Enterprise Security Engineer for AI/DevSecOps
Pioneering Enterprise Security Engineer for AI/DevSecOps

Harden • United States

On-site
Software Engineer, Applied AI $170,000 - $230,000
Software Engineer, Applied AI $170,000 - $230,000

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 170,000 - 230,000
Company-paid health coverage for you &
Meaningful early-stage equity
Founding AI Infrastructure Engineer — Remote/Hybrid, Equity
Founding AI Infrastructure Engineer — Remote/Hybrid, Equity

Harden • San Francisco (CA)

Hybrid
USD 150,000 - 250,000
Equity
Founding Member of Technical Staff (MTS)
Founding Member of Technical Staff (MTS)

VizopsAI • San Francisco (CA)

On-site
USD 150,000 - 220,000
Platform Engineer - AI Infrastructure
Platform Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 180,000 - 250,000
Significant freedom and ownership in project development
Work on challenging problems related to ultra-low latency
Join a high-growth environment
Staff Software Engineer (AI Infrastructure)
Staff Software Engineer (AI Infrastructure)

DeepRec.ai • Palo Alto (CA)

On-site
USD 180,000 - 320,000
Senior AI Platform SWE - PlayerZero
Senior AI Platform SWE - PlayerZero

HireOTS • Atlanta (GA)

On-site
USD 120,000 - 180,000
Software Development Manager - Replay
Software Development Manager - Replay

lyric.ai • United States

On-site
USD 120,000 - 160,000