Distributed Systems ML Infrastructure Engineer

openteams

Washington

Hybrid

USD 145,000 - 250,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenTeams is building an open, owner-controlled AI platform. We seek a Distributed Systems ML Infrastructure Engineer to design core services for workflow orchestration, data ingestion, and model serving across cloud and on-prem environments.

You will implement API-first services, ensure reliability at scale, and operate governance and logging for security-critical workloads. Applicants must have US citizenship and eligibility for clearance.

Qualifications

  • U.S. citizenship and eligibility to obtain and maintain a U.S. security clearance.
  • 6+ years of experience in distributed systems, platform or infrastructure engineering.
  • Experience with Kubernetes and containerized workloads on a major cloud platform.
  • Experience with managed Kubernetes services (e.g., Amazon EKS).
  • Experience supporting ML workloads in production and model serving.
  • Proficiency in Python, Go, or similar languages.
  • Experience with infrastructure-as-code and deployment tools (Terraform, Helm).
  • Ability to document interfaces, procedures, and validation results.
  • Bachelor's degree in computer science or engineering.

Responsibilities

  • Design and implement platform services for workflow orchestration and data ingestion with APIs.
  • Operate model gateway and serving services with policy enforcement and auditing.
  • Maintain cloud-agnostic provider abstractions for deployment on AWS or other environments.
  • Deploy releases into classified and controlled production environments and perform validations.
  • Ensure environment parity and validate platform capacity under load.
  • Document interfaces and architectural decisions clearly.

Skills

Kubernetes experience
Python or Go
API design
Infrastructure as code
Production ML workloads
Distributed systems
Security clearance eligibility

Education

Bachelor's degree in computer science or engineering

Tools

Terraform
Helm
AWS or cloud platforms

Job description

Who We Are

Every organization runs on intelligence: years of accumulated knowledge, decisions, and context. As AI takes on more of that work, companies face a choice: rent that intelligence from vendors who keep the data, the context, and the results, or own it.

OpenTeams exists to make ownership possible.

Founded by Travis Oliphant, creator of NumPy and SciPy, and built by people with deep roots across the open-source ecosystem, including NumPy, SciPy, PyTorch, and Jupyter, we help enterprises and governments build AI they control, govern, and evolve themselves.

If that sounds like your kind of work, we'd like to meet you.

Distributed Systems ML Infrastructure Engineer

Location: Washington, DC; Denver, CO; or Colorado Springs, CO preferred (hybrid). Highly qualified candidates outside these locations may also be considered for unclassified work.

Work Authorization: U.S. citizenship required

Clearance: An active TS/SCI clearance with CI polygraph is strongly preferred. Candidates without an active clearance may be considered for unclassified work but must be eligible to obtain and maintain a U.S. security clearance.

Salary Range: $145,000-$250,000 USD, dependent on experience level and location

About the Role

We're looking for a Distributed Systems and ML Infrastructure Engineer to build the core services of a containerized, API-first AI platform. This is a role for someone who wants to build the thing itself, not integrate someone else's.

You design and implement the services the platform runs on - workflow orchestration, data ingestion, results management, model serving, policy enforcement, usage accounting, audit logging. Those services have to hold up across cloud, dedicated, isolated, and limited-connectivity deployments, which means portability and operability are design constraints from the first commit rather than problems handed to someone downstream.

Development happens primarily on unrestricted infrastructure with an open-source toolchain. Engineers with the right access also carry releases into controlled production environments, integrate data sources there, and validate the platform in place - so there's a path to seeing your work through to where it actually runs.

This position is contingent upon contract award. Travel of up to 15% may be required, primarily to Government facilities and between company locations. Unclassified work may be performed remotely, while classified promotion and validation activities require onsite work in an accredited facility and the appropriate security clearance.

Key Responsibilities
  • Design and implement platform services for workflow orchestration, data ingest, and results management, exposed through documented APIs with no proprietary front end
  • Implement and operate model gateway and serving services that route invocations to approved managed model services with policy enforcement, usage accounting, and audit logging
  • Maintain a documented provider abstraction so the platform runs on AWS-native managed services where appropriate while remaining deployable across other cloud and dedicated environments
  • Deploy platform releases into classified host environments, perform data source integration, and execute validation procedures on a recurring promotion cadence
  • Verify environment parity after each promotion
  • Size and validate the platform against documented workload models, and verify capacity and performance by load test
  • Constrain platform dependencies to services confirmed available in the target environments, and gate any development-only dependency behind feature flags
  • Reproduce high-side defects on the low side through sanitized feedback paths and fix them where the full toolchain is available
Required Skills & Experience
  • U.S. citizenship and eligibility to obtain and maintain a U.S. security clearance
  • 6+ years of experience in distributed systems, platform engineering, infrastructure engineering, or a related software engineering role
  • Production experience operating Kubernetes and containerized workloads on a major cloud platform
  • Experience with managed Kubernetes services such as Amazon EKS or an equivalent platform
  • Experience supporting machine learning workloads in production, such as model serving, GPU scheduling, or large-scale data and evaluation pipelines
  • Experience designing, building, or operating distributed services that support reliability, scalability, and performance requirements
  • Proficiency in Python, Go, or a comparable programming language
  • Experience with infrastructure-as-code and deployment tools such as Terraform, Helm, or equivalent technologies
  • Experience designing API-first services and implementing documented interface specifications
  • Experience testing platform capacity and performance against expected workload requirements
  • Ability to document technical interfaces, deployment procedures, architectural decisions, and validation results
  • Bachelor's degree in computer science, engineering, or a related fi
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Full-Stack Platform Engineer
Full-Stack Platform Engineer

openteams • Washington

Hybrid
USD 145,000 - 250,000
Distributed Systems ML Infrastructure Engineer
Distributed Systems ML Infrastructure Engineer

OpenTeams • Washington, Denver (CO), Colorado Springs (CO)

Hybrid
USD 145,000 - 250,000
401(k) Match – Up to 5%
Unlimited PTO
Fully Remote Setup – up to $3,000
+3
Distributed Systems ML Infrastructure Engineer New Washington, DC Metro; Denver, CO Metro; or Colorado Springs, CO - Hybrid/Remote
Distributed Systems ML Infrastructure Engineer New Washington, DC Metro; Denver, CO Metro; or Colorado Springs, CO - Hybrid/Remote

OpenTeams • Colorado

Hybrid
USD 145,000 - 250,000
401(k) Match
Unlimited PTO
Fully Remote Setup
+3
Site Reliability Engineer / DevSecOps Engineer
Site Reliability Engineer / DevSecOps Engineer

openteams • Washington

Hybrid
USD 145,000 - 250,000
Technical Delivery Lead
Technical Delivery Lead

openteams • Washington, Denver (CO)

Hybrid
USD 145,000 - 250,000
Sr. Platform Engineer, ML Infrastructure
Sr. Platform Engineer, ML Infrastructure

Insilico Search Partners • Cambridge (MA)

On-site
USD 140,000 - 210,000
Senior Principal Platform Engineer
Senior Principal Platform Engineer

Clarity Innovations • Jessup (MD)

On-site
USD 130,000 - 150,000
Platform Architect (AI/ML Infrastructure, GCP-focused)
Platform Architect (AI/ML Infrastructure, GCP-focused)

Wizdaa • United States

Remote
MXN 2,042,000 - 3,062,000
SME Platform Engineer
SME Platform Engineer

General Dynamics Information Technology • Arlington (VA)

On-site
USD 150,000 - 190,000
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Perplexity • California (MO)

On-site
USD 140,000 - 190,000