Member of Technical Staff — Infrastructure

Observable Intuition, Inc.

New York, Northern (NY, KY)

Hybrid

USD 120,000 - 200,000

Full time

9 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Observable Intuition, Inc. seeks a founding Infrastructure Engineer to define and own the production inference platform behind our data-driven AI layer.

You will collaborate with the founding team to build core systems from the ground up, make foundational architectural decisions, and shape the engineering culture across deployments. Your work will span managed cloud, customer-owned infrastructure, and fully air-gapped environments, ensuring performance, reliability, and security at scale.

Qualifications

  • Experience operating machine-learning or similarly compute-intensive distributed systems in production.
  • Strong experience with cloud infrastructure, containers, orchestration, and observability.
  • Experience designing multi-tenant platforms with rigorous isolation and security boundaries.
  • A record of owning production systems from architecture through operation.
  • Familiarity with Infrastructure as Code (IaC) and tools such as Terraform.
  • Comfort working hands-on with substantial autonomy and without an established blueprint.

Responsibilities

  • Own our production inference platform across model serving, orchestration, deployment, observability, and operations.
  • Optimize workloads for batching, caching, scheduling, and routing while balancing latency, throughput, quality, and cost.
  • Build reliable infrastructure for high-volume data processing, including backfills and safe reprocessing.
  • Design a portable platform that runs across managed cloud, customer-owned infrastructure, and air-gapped environments.
  • Ensure multi-tenancy with strong per-customer isolation across data, compute, identity, and operations.
  • Own the model lifecycle, including versioning, evaluation, rollout, monitoring, and rollback.
  • Define data retention, recovery, governance, and secure deletion across deployments.
  • Collaborate with founders to shape technical strategy and infrastructure roadmap.

Skills

Distributed systems
Cloud infrastructure
Containers & orchestration
Observability
Multi-tenant design
Production ownership
IaC (Terraform)
Autonomy

Tools

Kubernetes
Terraform
Triton
vLLM
TensorRT
Ray

Job description

Member of Technical Staff — Infrastructure

San Francisco / New York · $120,000 – $200,000 + equity

AI has advanced by expanding what machines can represent.

Deep learning enabled models to learn structure directly from raw data. Transformers extended that capability further, giving systems access to the vast corpus of human knowledge encoded in language.

But language is not experience.

Human expertise is not formed by reading descriptions of judgment: it is formed through action, feedback, and consequence. It comes from operating in the world, making decisions under uncertainty, and learning what actually matters when reality pushes back.

Today’s models can ingest the artifacts of that experience, but they cannot observe the experience itself. Collective Intuition is building the missing layer: infrastructure that makes real-world human experience observable, structured, and learnable by AI systems.

We work with some of the world’s largest enterprises, where judgment is exercised continuously through decisions, exceptions, approvals, failures, and outcomes. We transform this fragmented operational reality into structured, provenance-rich data that AI systems can learn from.

Language gave machines access to what humanity has said about the world. Collective Intuition gives them access to what actually happened when people acted within it.

The role

As our Founding Infrastructure Engineer, you’ll define and own the production inference platform behind this new layer of intelligence.

You’ll work alongside a founding team who published in Nature, whose research was funded by Google DeepMind, and with experience building and deploying AI systems in Fortune 500s. This is a deeply hands-on role: you’ll build core systems from the ground up, make foundational architectural decisions, and help shape the engineering culture of the company.

The platform must operate wherever the world’s most demanding enterprises need it: from our managed cloud to customer-controlled infrastructure and fully air-gapped environments. You’ll design the architecture that makes this possible without fragmentation, while preserving performance, reliability, and security across all deployments.

What you’ll do
  • Own our production inference platform across model serving, orchestration, deployment, observability, and operations.
  • Optimize how workloads are batched, cached, scheduled, and routed while balancing latency, throughput, quality, and cost.
  • Build reliable infrastructure for continuously processing high-volume data, including backfills and safe reprocessing.
  • Design one portable platform that runs consistently across managed cloud, customer-owned infrastructure, and fully air-gapped environments.
  • Make multi-tenancy a foundational property of the system, with strong per-customer isolation across data, compute, identity, and operations.
  • Own the model lifecycle, including versioning, evaluation, rollout, monitoring, and rollback.
  • Define how data is retained, recovered, governed, and securely deleted across every deployment model.
  • Work directly with the founders to shape our technical strategy and infrastructure roadmap.
What we’re looking for
  • Experience operating machine-learning or similarly compute-intensive distributed systems in production.
  • Strong experience with cloud infrastructure, containers, orchestration, and observability.
  • Experience designing multi-tenant platforms with rigorous isolation and security boundaries.
  • A record of owning production systems from architecture through operation.
  • Familiarity with Infrastructure as Code (IaC) and tools such as Terraform.
  • Comfort working hands-on with substantial autonomy and without an established blueprint.
Useful
  • Experience operating and optimizing GPU workloads.
  • Experience with Kubernetes, Triton, vLLM, TensorRT, Ray, or comparable technologies.
  • Experience deploying into customer-controlled, private-cloud, or fully air-gapped environments.
  • Experience with enterprise security, data governance, or production ML evaluation.
  • Experience as a founding or early infrastructure engineer.

The next frontier in AI is the representation of experience itself.

Real work is a difficult learning environment: state is distributed, actions are often implicit, and outcomes arrive long after decisions are made. Turning this into something machines can learn from requires rethinking inference, distributed systems, data infrastructure, security, and long-horizon reasoning all at once.

That is the problem we are solving.

You’ll build infrastructure that allows models to learn from real human decisions and their consequences, while operating within the strict security and deployment constraints of the world’s largest enterprises. The challenge is not only to make the system scale; it is to make a single intelligent platform work across thousands of isolated customers, private environments, and fully disconnected networks.

There is no established blueprint for this.

You’ll have significant influence over the architecture, the product, and the company we build around it.

We are not building another interface to existing models. We are building the foundation from which the next generation of intelligent systems can learn.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff — Research
Member of Technical Staff — Research

Observable Intuition, Inc. • New York (NY), Northern (KY)

On-site
USD 120,000 - 250,000
Equity
Founding Engineer - ML Infrastructure
Founding Engineer - ML Infrastructure

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k)
Paid time off
+2
Machine Learning Engineer, Platform
Machine Learning Engineer, Platform

Brain Co. • San Francisco (CA)

On-site
USD 150,000 - 210,000
Machine Learning Engineer, Platform NY
Machine Learning Engineer, Platform NY

Brain Co. • New York (NY)

On-site
USD 140,000 - 210,000
Member of Technical Staff - Training Platform
Member of Technical Staff - Training Platform

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Remote or SF office
Visa sponsorship
Relocation support
+3
Member of Technical Staff, Infrastructure
Member of Technical Staff, Infrastructure

Intent Lab Inc • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Research Engineer - RL Infrastructure
Research Engineer - RL Infrastructure

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 350,000
Remote or SF office work option
Visa sponsorship & relocation
Quarterly team offsites
Founding Engineer, AI Infrastructure
Founding Engineer, AI Infrastructure

Piris Labs • San Francisco (CA)

On-site
USD 100,000 - 200,000
Equity
401(k)
Health insurance
+1
Member of Technical Staff - Compute Platform
Member of Technical Staff - Compute Platform

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 230,000
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2