AI Lab Tech Engineer

Innova software Services Inc

Fremont (CA)

Hybrid

USD 120,000 - 180,000

Full time

12 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Innova software Services Inc. is seeking an AI Lab Tech Engineer to own the execution layer beneath reinforcement learning environments. You will build scalable, secure sandboxing and high-performance execution for production workloads, collaborating with research, data teams, and enterprise clients.

The role emphasizes long-horizon, multi-node scaling, cost-aware optimization, and robust observability. A hybrid work model (2 days onsite, 3 days remote) is supported in Fremont, CA.

Qualifications

  • Production distributed systems experience.
  • Containerization and isolation tech (namespaces, cgroups, VMs, gVisor, Firecracker).
  • Proficiency in Python; experience with Rust, Go, or C++.
  • Cloud & on-prem deployments on AWS, GCP and enterprise hardware.
  • Performance profiling, scheduling, and cost optimization.
  • Strong communication with technical stakeholders and enterprise clients.

Responsibilities

  • Environment Execution & Sandboxing: design/maintain isolation and execution layer for RL environments; snapshot/restore state to pause, resume, inspect, branch rollouts.
  • Fault Detection & Recovery: automate failure detection, revert to known-good states, patch and resume execution.
  • Long-Horizon & Multi-Node Scaling: extend execution to multi-node environments across tools/services for hours/days.
  • Performance & Cost Optimization: drive throughput/latency, manage scheduling and usage to maximize rollouts per dollar.
  • Observability & Profiling: identify bottlenecks from container init to teardown; build deep observability for thousands of runs.
  • Environment Platform Engineering: build frameworks/tools to package/deploy RL environments; provide debugging tooling.
  • Model Deployment: deploy/manage AI models on on-premises hardware.
  • Production Standards: scale prototypes into production-grade systems with tests, validation, and docs.

Skills

Systems & Infrastruktur
Low-Level Systems
Languages & Tools
Cloud & On-Premises
Performance Engineering
Soft Skills

Tools

Claude Code

Job description

Job Title:

AI Lab Tech Engineer

Location:

Fremont, CA (Hybrid: 2 days Onsite / 3 days Remote)

Job Type:

Long-Term Contract

Role Overview

We are seeking an Infrastructure Engineer to own the execution layer beneath our Reinforcement Learning (RL) environments: the core systems that enable AI agents to operate coherently inside realistic, multi-tool worlds for extended periods.

This position focuses on hard systems engineering applied to AI infrastructure. As agent task horizons expand, our training environments must maintain stability, speed, and cost-efficiency. You will build and scale the platform that provides secure sandboxing, high-performance execution, and state-management capabilities (snapshotting, restoring, inspecting, and branching running environments) for production workloads. You will collaborate with research and data teams, frontier labs, and enterprise clients to translate complex environment requirements into resilient systems.

Key Responsibilities
  • Environment Execution & Sandboxing: Design and maintain the isolation and execution layer for RL environments. Build systems to snapshot and restore environment state (disk, process, memory, and accelerator state) to enable pausing, resuming, inspecting, and branching agent rollouts.
  • Fault Detection & Recovery: Implement automated machinery to identify failure modes early (infra faults, reward hacks, fairness issues), revert to known-good states, patch, and resume execution.
  • Long-Horizon & Multi-Node Scaling: Extend execution capabilities to handle complex, multi-node environments where agents interact across multiple tools and services over hours or days.
  • Performance & Cost Optimization: Drive throughput, latency optimization, and cost-per-rollout efficiency. Manage scheduling and resource utilization to maximize environment rollouts per dollar without compromising system reliability.
  • Observability & Profiling: Profile bottlenecks from container initialization to teardown. Develop deep observability into thousands of concurrent, long-running agent rollouts.
  • Environment Platform Engineering: Build frameworks and tooling for specifying, packaging, and deploying RL environments used by internal researchers and agents. Provide debugging tooling to trace failures across complex agent runs.
  • Model Deployment: Deploy and manage small and large AI models on on-premises hardware.
  • Production Standards: Scale prototypes into production-grade systems with high testing, validation, and documentation standards.
Required Qualifications & Skills
  • Systems & Infrastructure: Proven background in building production distributed systems, execution engines, or container/sandboxing infrastructure at scale.
  • Low-Level Systems Knowledge: Deep technical experience with containerization and isolation tech (namespaces, cgroups, VMs, gVisor, Firecracker), filesystems, and process/state management.
  • Languages & Tools: High proficiency in Python. Experience with systems languages (Rust, Go, or C++) and modern development tools (e.g., Claude Code) is strongly preferred.
  • Cloud & On-Premises: Practical experience with major cloud platforms (AWS, GCP) alongside hands-on experience deploying models on on-premises hardware.
  • Performance Engineering: Demonstrated track record in profiling, workload scheduling, resource management, and cost optimization for compute-heavy workloads.
  • Soft Skills: Strong communication skills to effectively translate research requirements into production infrastructure and engage directly with technical stakeholders and enterprise clients.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Lab Tech Engineer
AI Lab Tech Engineer

Commergence • Colorado

Hybrid
USD 150,000 - 210,000
AI Lab Tech Engineer
AI Lab Tech Engineer

Commergence • Fremont (CA)

Hybrid
USD 120,000 - 180,000
Backend Engineer
Backend Engineer

Bespoke Labs • Mountain View (CA)

On-site
USD 100,000 - 140,000
Health coverage
Opportunity to work with leading AI research labs
Backend Engineer
Backend Engineer

bespokelabs • Mountain View (CA)

On-site
USD 120,000 - 160,000
Health coverage
Competitive salary and equity
Opportunity to work with leading AI research labs
Research Engineer, RL Environments and Infrastructure
Research Engineer, RL Environments and Infrastructure

Hyphen Connect • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
RL AI Infrastructure Engineer
RL AI Infrastructure Engineer

Innova software Services Inc • Fremont (CA)

Hybrid
USD 120,000 - 180,000
Research Engineer – RL Infrastructure & Agent Environments
Research Engineer – RL Infrastructure & Agent Environments

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 150,000 - 190,000
AI Engineer
AI Engineer

Teserac, Inc. • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Health Care Plan
Paid Time Off
Free Food
+2
Hybrid AI RL Infrastructure Engineer — Sandbox & Scale
Hybrid AI RL Infrastructure Engineer — Sandbox & Scale

Momento USA • Fremont (CA)

Hybrid
USD 120,000 - 155,000
RL Environment Software Engineer
RL Environment Software Engineer

talentpluto • San Francisco (CA)

On-site
USD 180,000 - 220,000