AI Lab Tech Engineer

Commergence

Fremont (CA)

Hybrid

USD 120,000 - 180,000

Part time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Commergence in Fremont, CA seeks an AI Lab Tech Engineer to own the execution layer for RL environments, focusing on sandboxing, fast and cost-efficient execution, and snapshot/restore capabilities to support hours- or days-long runs.

You will collaborate with research, data teams, and enterprise customers to build reliable on-prem infrastructure for deploying RL environments with strong performance, scale, and observability.

Qualifications

  • Experience building production systems or research infrastructure at scale.
  • Strong knowledge of containers, namespaces, cgroups, VMs, and sandboxing.
  • Experience optimizing performance and cost at scale.
  • Proficiency in Python; systems programming in Rust/Go/C is a plus.
  • Experience deploying models on on-prem hardware.

Responsibilities

  • Design and own the sandboxing and execution layer for environments.
  • Develop failure-detection and rollback mechanisms for rollouts.
  • Scale environment rollouts across multi-node setups.
  • Own performance metrics: throughput, latency, cost per rollout at scale.
  • Build framework for specifying, packaging, and deploying RL environments.
  • Collaborate with research and enterprise customers.
  • Produce tooling and documentation for internal teams.

Skills

Python
Rust
Go
Systems programming
Distributed systems
Containerization
Performance optimization
Cloud platforms

Job description

AI Lab Tech Engineer

Location-Fremont, CA (Local Onsite/Hybrid work) Job Type-Long Term Contract

We're looking for an Infrastructure Engineer to own the execution layer beneath our RL environments: the systems that let an agent operate inside a realistic, multi-tool world coherently for hours or days.

That means sandboxing and isolation you can trust, execution that's fast and cheap enough to run at training scale, and the ability to snapshot, restore, inspect, and branch a running environment instead of treating every rollout as one-shot. You'll build the platform that makes all of this possible.

You’ll work closely with our research and data teams, and directly with frontier labs and enterprise customers, to turn environment designs into infrastructure that runs reliably in production.

What You'll Do
  • Environment Execution & Sandboxing:
    • Design and own the sandboxing and execution layer that environments run inside. Build systems to snapshot and restore environment state (disk, process, and where relevant memory and accelerator state) so runs can be paused, resumed, inspected, and branched rather than executed once.
    • Develop the machinery to detect failure modes early in a rollout (reward hacks, infra faults, fairness issues) and to revert to a known-good state, patch, and continue.
    • Extend execution to long-horizon and multi-node environments, where an agent operates across many tools and services over hours or days.
  • Performance & Scale
    • Own the performance characteristics of the platform: throughput, latency, and cost-per-rollout at scale.
    • Drive utilization and scheduling so we can run far more environment rollouts per dollar without sacrificing reliability.
    • Profile and remove bottlenecks across the stack, from container startup to environment teardown.
    • Build the observability that lets us understand what's happening inside thousands of concurrent, long-running rollouts.
  • Environment Platform
    • Build and maintain the framework for specifying, packaging, and deploying RL environments which is used by both humans and agents authoring environments internally.
    • Create the tooling that lets researchers and environment authors debug a specific failure across hundreds of long agent traces.
    • Deploy large / small models on on-prem hardware
  • Collaboration & Production Excellence
    • Scale prototypes into production systems with reproducible workflows and high engineering standards.
    • Write the documentation and tools that let internal teams and external users build on the platform.
What We're Looking For
  • Systems & Infrastructure
    • Strong track record building production systems or research infrastructure at scale: distributed systems, execution engines, container/sandboxing infrastructure, or similar.
    • Deep comfort with the systems layer: containers and isolation (e.g. namespaces, cgroups, VMs, gVisor/Firecracker-style sandboxing), filesystems, process and state management.
    • Experience making systems fast and cheap --- profiling, scheduling, resource utilization, and cost optimization at scale.
    • Proficiency with cloud platforms (GCP, AWS) and distributed computing.
    • Strong engineering fundamentals and a systematic approach to testing, validation, and reliability.
    • Experience of deploying large / small models on on-prem hardware
  • Execution & Ownership
    • Comfort operating in ambiguity.
    • Strong Python skills; comfort in a systems language (Rust, Go, or C ) is a plus.
    • Ability to use modern tools such as Claude Code effectively.
  • Collaboration & Communication
    • Excellent communication skills for working with research teams and enterprise customers.
    • Ability to translate between research needs and infrastructure requirements.
    • Comfortable presenting technical work to diverse audiences.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Lab Tech Engineer
AI Lab Tech Engineer

Commergence • Colorado

Hybrid
USD 150,000 - 210,000
Backend Engineer
Backend Engineer

bespokelabs • Mountain View (CA)

On-site
USD 120,000 - 160,000
Health coverage
Competitive salary and equity
Opportunity to work with leading AI research labs
Backend Engineer
Backend Engineer

Bespoke Labs • Mountain View (CA)

On-site
USD 100,000 - 140,000
Health coverage
Opportunity to work with leading AI research labs
Staff ML Engineer, Agent Training & Environments
Staff ML Engineer, Agent Training & Environments

EngineersOfAI • San Francisco (CA)

On-site
USD 170,000 - 260,000
RL Sandbox & Infra Engineer - Scalable Run-Time Platform
RL Sandbox & Infra Engineer - Scalable Run-Time Platform

Commergence • Fremont (CA)

Hybrid
USD 120,000 - 180,000
Research Engineer
Research Engineer

Bespoke Labs • Mountain View (CA)

On-site
USD 120,000 - 140,000
Health coverage
Opportunity to work with leading AI labs
Competitive salary and equity
RL Environment Infrastructure Engineer — Sandbox & Scale
RL Environment Infrastructure Engineer — Sandbox & Scale

Commergence • Colorado

Hybrid
USD 150,000 - 210,000
RL Environments Engineer
RL Environments Engineer

Bespoke-Labs • Mountain View (CA)

On-site
USD 250,000 - 300,000
Health, dental, and vision
401(k)
Daily onsite lunch
+3
Software Engineer, RL Training Infra
Software Engineer, RL Training Infra

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
RL Environments Engineer
RL Environments Engineer

Bespoke Labs Inc. • Mountain View (CA), Northern (KY)

Hybrid
USD 250,000 - 300,000
Health, dental, and vision coverage
401(k)
Daily onsite lunch provided
+2