Infrastructure Engineer

Pursuit Talent Advisory

Austin (TX)

On-site

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Pursuit Talent Advisory is partnering with an early growth-stage company building AI systems for industrial environments. The engineer will own on-prem deployment, run reliability on constrained hardware, and ensure repeatable releases onto customer hardware with VPN constraints.

Expect to write application code, contribute to integration with cloud dev/staging, and own CI/CD, GPU inference, and hardened security practices. Work is hands-on and autonomous in a small, AI-native team.

Qualifications

  • Strong Kubernetes fundamentals including StatefulSets, storage, networking, ingress, and debugging.
  • Infrastructure as code in production (Terraform or equivalent).
  • Docker with multi-service builds and registry workflows.
  • CI/CD pipelines ownership and maintenance.
  • Experience deploying AI workloads on GPUs with private hardware.
  • Linux/ Bash proficiency and ability to read/write Python.

Responsibilities

  • Own the on-prem deployment path and ensure reliability on single-node and small multi-node setups.
  • Own GitOps with declarative reconciliation, image pinning, and rollback boundaries.
  • Own cloud dev/staging infrastructure as code and promote through environments.
  • Own CI/CD pipelines for multi-service builds and environments that teardown automatically.
  • Own GPU inference layer with model/quantization choices and latency targets.
  • Improve operability with runbooks, telemetry, and secure separation of components.
  • Reduce toil by automating deployment steps and delegating to agents.

Skills

Kubernetes fundamentals
Terraform
Docker
CI/CD ownership
Linux Bash
Python
nvidia-smi readouts
On-prem hardware

Tools

GitOps (Flux/Argo CD)
GPU inference server (vLLM/TGI/TensorRT-LLM)

Job description

We recently partnered with an early growth-stage company that builds AI systems for industrial environments. The platform runs on customer-owned hardware, on-site, without a cloud dependency at runtime. It takes in live data from plant equipment and business systems, turns it into a structured view of the operation, and puts an AI layer on top of it for the people running the floor.

The team is small, flat, and AI-native. ICs own their work end to end: scope, build, ship, verify. Autonomy is the default.

THE ROLE

You own how the platform gets deployed, runs, and stays up, both on constrained on-premise hardware with unreliable connectivity and in the cloud environments used to develop and validate it. This is not a "keep the CI green" job: the deployment target is physical hardware at a customer site, and getting a release onto it repeatably is a genuine engineering problem.

You'll be one of the first few engineers on the team. Expect to write application code too.

WHY THIS IS INTERESTING

Most infrastructure roles are cloud roles. This one isn't. The software has to run on a box someone can walk up to and unplug, in a building with a spotty VPN, next to machines that cost more than the company. That constraint makes almost every decision (deployment, secrets, observability, rollback) more interesting than the cloud version of the same problem.

WHAT YOU'LL DO
  • Own the on-premise deployment path. Lightweight Kubernetes on single-node and small multi-node topologies, templated and layered manifests, staged bring-up of a full stack from bare hardware, and clean recovery after hard power loss.
  • Own GitOps. Declarative reconciliation from a Git source of truth, CI-driven image pinning, and revert-based rollback, with a clear, enforced boundary around what the reconciler is and isn't allowed to manage.
  • Own the cloud dev and staging estate. Infrastructure as code for compute, managed databases, container registry, and storage. Enforce the promotion flow from branch to PR to static checks to a real staging environment before anything reaches a shared environment.
  • Own CI/CD. Multi-service image builds, manifest generation, migration gating, and episodic environments that come up and tear down without a human babysitting them.
  • Own the GPU inference layer. The company runs its own models on customer hardware rather than calling a provider API. You'll own that: driver and device-plugin prep, model and quantization choices against whatever hardware you're given, and serving configuration (batching, KV cache, context length, parallelism, memory utilization) tuned so a fixed box meets latency targets. Customers buy one appliance, so utilization is a hard budget, not a cost-optimization exercise.
  • Make the appliance operable. Secrets that don't live in Git, telemetry and model-call tracing a support engineer can actually read, and runbooks for the failure modes that will happen at 2 am on a factory floor.
  • Harden it. Industrial networks are a real trust boundary, and some components need elevated access to do their job. Keep the isolation between privileged components, application services, and the AI runtime intact as the system grows.
  • Reduce toil. If a deployment step needs a person, automate it or delegate it to an agent.
WHAT WE'RE LOOKING FOR
Required:
  • Strong Kubernetes fundamentals, not just kubectl apply, but StatefulSets, storage, networking, ingress, and debugging a cluster that's misbehaving. Bare-metal or single-node experience (k3s, RKE2, MicroK8s) counts more here than managed EKS/AKS.
  • Infrastructure as code in production (Terraform or equivalent), with real state management and multi-environment discipline.
  • Docker beyond the basics: multi-service builds, image size and layer hygiene, registry workflows.
  • CI/CD ownership: you've built and maintained pipelines, not just consumed them.
  • Hosting AI workloads on GPUs. You've run model inference on your own hardware: an inference server (vLLM, TGI, TensorRT-LLM, or similar) behind a real workload, on GPUs you were responsible for. You can reason about VRAM budgets, batching and concurrency, KV cache, quantization tradeoffs, and where throughput vs. latency actually breaks. You know how to read nvidia-smi and a profile and say why a GPU is underutilized.
  • Comfortable in Linux and Bash, and able to read and write Python. Backend services are written in Python; you'll be in the code.
  • You debug from first principles and you write down what you learn.
Nice to have:
  • GitOps at scale (Flux or Argo CD).
  • Edge, air-gapped, or on-prem deployments: anywhere you couldn't assume a cloud control plane.
  • Industrial / OT exposure: PLCs, industrial protocols, MQTT, or manufacturing environments generally.
  • Multi-tenant or multi-model GPU sharing (MIG, time-slicing, MPS), and GPU scheduling in Kubernetes.
  • Operating Postgres and time-series databases: migrations, backup/restore, retention.
  • Observability: OpenTelemetry, or LLM tracing and evaluation tooling.
HOW THE TEAM WORKS
  • 1-week sprints. Plan Monday, review end of week.
  • Async-first. Write it down over scheduling a meeting. Daily standup is short and time-boxed.
  • AI-native. The team leans on AI tooling across the whole workflow. Mechanical work gets automated or handed to an agent.
  • Bias to ship. Small, frequent changes over big-bang releases.
  • Multiple hats. Lean startup. Everyone stretches beyond their title.

This role reports to the CTO and works alongside the engineering ICs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Engineer - AI, Kubernetes & Edge Systems
Infrastructure Engineer - AI, Kubernetes & Edge Systems

UMATR • Austin (TX)

On-site
USD 120,000 - 190,000
Principal Product Engineer, Cloud Platform
Principal Product Engineer, Cloud Platform

Verdigris Technologies Inc • Palo Alto (CA)

On-site
USD 130,000 - 180,000
Member of Technical Staff, Infrastructure
Member of Technical Staff, Infrastructure

Psi • Boston (MA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Meaningful equity
Competitive compensation
Benefits
Infrastructure Engineer — Seed-Stage AI Lab
Infrastructure Engineer — Seed-Stage AI Lab

Aionia • San Francisco (CA)

On-site
USD 185,000 - 235,000
Visa Sponsorship
Equity Options
Principal Product Engineer, Cloud Platform
Principal Product Engineer, Cloud Platform

Verdigris Technologies • Palo Alto (CA)

On-site
USD 190,000 - 270,000
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Madrona Venture Labs • United States

Hybrid
USD 180,000 - 260,000
Head of Infrastructure
Head of Infrastructure

General Compute Inc. • New York (NY)

On-site
USD 120,000 - 150,000
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Member of Technical Staff (Software Engineer, Inference & Training Platform)

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Software Engineer - Infrastructure
Software Engineer - Infrastructure

Emergentlabsinc • San Francisco (CA)

On-site
USD 110,000 - 150,000
401(k)
Health, dental, and vision insurance
Unlimited Paid Time Off
+1
Senior / Lead Infrastructure & Operations Engineer
Senior / Lead Infrastructure & Operations Engineer

Austin Werner • Boston (MA)

On-site
USD 120,000 - 150,000