Software Engineer, Infrastructure & Platform

10a Labs

United States

Remote

USD 110,000 - 160,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Fully remote, U.S.-based
Performance-based annual bonus
Professional development support (con-
Leadership training

Job summary

10a Labs is seeking a Software Engineer, Infrastructure & Platform, to craft the sandboxed environments and backend systems powering advanced AI evaluations. You will enable tools interaction, secure code execution, and multi-step workflows across distributed services.

The role emphasizes building scalable, observable, and secure infrastructure for autonomous model behavior, agentic workflows, and risk evaluation, with a focus on security and reliability in a remote US-based setting.

Qualifications

  • 3–5+ years of professional software engineering experience in backend/infrastructure
  • Strong Python production-grade software development skills
  • Experience designing and operating backend services or distributed systems
  • Hands-on experience with Docker, Kubernetes, VMs or similar container/orchestration tech
  • Experience with cloud platforms such as GCP or AWS and Linux networking/security
  • Familiarity with infrastructure-as-code/automation tools like Terraform
  • Ability to build reproducible, observable, scalable, and secure systems
  • Interest in AI safety, agentic workflows, or model evaluations is a plus

Responsibilities

  • Design and build sandboxed evaluation environments for safe code execution and tool-use
  • Develop backend services and infrastructure for large-scale AI evaluations
  • Create agent scaffolding and evaluation harnesses with robust context management and retries
  • Provision and orchestrate isolated environments using Docker, Kubernetes, VMs and cloud infra
  • Implement secure networking, secrets management and resource isolation
  • Build APIs and automation to enable efficient evaluations by analysts and engineers
  • Increase reliability through logging, observability, snapshotting and automated tests
  • Support thousands of evaluation tasks with scalable, reproducible workflows
  • Collaborate with analysts and domain experts to translate evaluation ideas into systems
  • Investigate failures across the stack to distinguish model vs infra issues

Skills

Python
Backend systems
Distributed systems
Docker
Kubernetes
Cloud infrastructure
Security
Observability
Linux

Tools

Docker
Kubernetes
Terraform
GCP
AWS

Job description

Software Engineer, Infrastructure & Platform

Remote

About 10a Labs: 10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. Our adversarial red teaming, model evaluations, and intelligence collection enable engineering, safety, and security teams to stay ahead of evolving threats and deploy AI systems safely.

Software Engineer, Infrastructure & Platform
About the Role

We are seeking a Software Engineer, Infrastructure & Platform to build the systems and infrastructure that power advanced AI evaluations, including evaluations focused on autonomous model behavior, agentic systems, and loss-of-control risks.

This is a hands-on engineering role at the intersection of backend systems, infrastructure, and AI. You will build secure and reproducible environments where frontier models can interact with tools, execute code, complete complex tasks, and operate across realistic multi-step workflows, including machine learning research and engineering.

The ideal candidate has strong backend and infrastructure fundamentals with attention to security, enjoys debugging complex distributed systems, and is excited to apply those skills to difficult problems in AI safety and evaluation.

What You'll Do
  • Design and build sandboxed evaluation environments where AI models can safely execute code, use tools, interact with services, and complete complex tasks.
  • Build backend services and infrastructure supporting large-scale, repeatable AI and agentic evaluations.
  • Develop agent scaffolding and evaluation harnesses, including tool-use loops, context management, retries, state management, token budgets, and multi-agent or subagent workflows.
  • Build systems for provisioning and orchestrating isolated environments using technologies such as Docker, Kubernetes, VMs, and cloud infrastructure.
  • Design secure approaches to networking, permissions, secrets, credentials, and resource isolation for model-driven environments.
  • Develop APIs, internal tools, and automation that allow analysts, engineers, and subject-matter experts to create and run evaluations efficiently.
  • Improve the reliability and reproducibility of evaluations through logging, observability, snapshotting, debugging tools, and automated testing.
  • Build systems capable of running thousands of evaluation tasks reliably and capturing the artifacts and telemetry needed to understand model behavior.
  • Partner with analysts, red teamers, and domain experts to translate complex evaluation ideas into robust technical systems.
  • Investigate failures across the evaluation stack and distinguish between model limitations and infrastructure, harness, or environment failures.
What We're Looking For
  • 3–5+ years of professional software engineering experience, particularly in backend, infrastructure, platform, SRE, or distributed systems engineering.
  • Strong programming skills in Python and experience building production-quality software.
  • Experience designing and operating backend services or distributed systems.
  • Hands‑on experience with Docker, Kubernetes, virtual machines, or other container/orchestration technologies.
  • Experience working with GCP, AWS, or similar cloud infrastructure.
  • Strong understanding of Linux systems, networking, authentication, permissions, and infrastructure security.
  • Experience with infrastructure-as-code or automation tools such as Terraform.
  • Strong debugging skills and comfort diagnosing failures across application, infrastructure, and networking layers, especially in agentic loops.
  • Ability to build systems that are reproducible, observable, scalable, and secure.
  • Comfort working on ambiguous technical problems where the architecture and requirements may evolve quickly.
  • Interest in AI systems, agentic workflows, AI security, or model evaluations. Prior professional AI experience is helpful but not required.
Nice to Have
  • Experience building developer platforms, CI/CD systems, test infrastructure, sandboxes, or epistemic compute environments.
  • Experience with agent frameworks, LLM APIs, tool‑calling systems, or AI evaluation infrastructure.
  • Experience designing secure execution environments for untrusted or semi‑trusted code.
  • Background in SRE, platform engineering, cloud infrastructure, cybersecurity, or developer tooling.
  • Experience with distributed task execution, queues, workflow orchestration, or large-scale automated testing.
  • Familiarity with AI safety, adversarial testing, model evaluations, or autonomous‑agent systems.
  • Familiarity with agentic AI fundamentals, including common harnesses, Model Context Protocol, agent benchmarks, and security risks to AI agents.
  • Salary Range: $110K–$160K, depending on experience and location
  • Bonus: Performance‑based annual bonus
  • Professional Development: Support for conferences, continuing education, or leadership training
  • Work Environment: Fully remote, U.S.-based
  • Time Off: Generous PTO and paid holiday schedule
  • Retirement: 401(k) plan
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Infrastructure & Platform
Software Engineer, Infrastructure & Platform

Jobgether • United States

Remote
USD 110,000 - 160,000
Competitive salary
Bonus
Fully remote
+5
Machine Learning Engineer
Machine Learning Engineer

Jobzhr • San Francisco (CA), Northern (KY)

On-site
USD 130,000 - 200,000
Performance-based annual bonus
Conferences/continuing education
Fully remote (U.S.-based)
+2
Security Engineer
Security Engineer

10a Labs • United States

Remote
USD 105,000 - 125,000
Health, dental, and vision coverage
Generous PTO and paid holidays
401(k) plan
+1
Remote AI Infrastructure & Platform Engineer
Remote AI Infrastructure & Platform Engineer

10a Labs • United States

Remote
USD 110,000 - 160,000
Fully remote, U.S.-based
Performance-based annual bonus
Professional development support (con-
+1
Principal AI/ML Engineer - AI Safety & Evaluation
Principal AI/ML Engineer - AI Safety & Evaluation

A10 Networks, Inc. • San Jose (CA)

On-site
USD 225,000 - 245,000
Head of AI Red Teaming
Head of AI Red Teaming

Trajectory Labs, PBC • Berkeley (CA), Northern (KY)

Hybrid
USD 250,000 - 400,000
Equity
Health coverage
401(k)
+1
Forward Deployed Engineer - Language Models
Forward Deployed Engineer - Language Models

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation including **e
Member of Technical Staff (Language Model Evaluations)
Member of Technical Staff (Language Model Evaluations)

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Senior AI Forward Deployed Engineer
Senior AI Forward Deployed Engineer

Handshake • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Equity in a fast-growing company
401(k) match
Paid parental leave
+2
Backend Software Engineer (Evals)
Backend Software Engineer (Evals)

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 385,000