Senior Systems and Platform Engineer

Maxisiq

Bethesda (AR)

On-site

USD 180,000 - 260,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Vexterra Group in Bethesda, MD seeks a platform engineer with deep OS, hardware, GPU, and networking expertise to design and optimize Kubernetes clusters powering enterprise AI for mission customers. You will implement IaC, secure pipelines, and maintain documentation while ensuring compliance with federal standards.

Preferred candidates have DoD 8570 IAT certifications, experience with NVIDIA GPU data centers, and AWS infrastructure.

Qualifications

  • Bachelor’s degree with extensive relevant experience or higher degree and matching years.
  • 5+ years in Platform/Systems Engineering.
  • Strong Linux and Kubernetes administration experience.
  • Experience with Docker containers across SDLC.
  • AI/ML workflow orchestration experience (Argo, Airflow, Kubeflow) is a plus.
  • Must meet DoD 8570.11-IAT Level II certification requirements; higher certs acceptable.

Responsibilities

  • Design, configure, and maintain enterprise Kubernetes platforms.
  • Develop and manage IaC using Terraform, Salt, Ansible, Bash, and Python.
  • Collaborate to design secure, automated CI/CD pipelines.
  • Troubleshoot complex issues across cloud, network, and platform layers.
  • Maintain technical docs and ensure compliance with federal standards.

Skills

Linux
Kubernetes
Docker
AI/ML workflows
REST APIs
IAT II/III certs
Security Clearance
IaC (Terraform/Salt/Ansible)
CI/CD pipelines
NVIDIA GPU data centers
GPU tooling (Base Command, Run:AI)
Cloud networking

Education

Bachelor’s degree

Tools

Terraform
Salt
Ansible
Bash
Python
GitLab CI/CD

Job description

Job Description

Our Partner company, Vexterra Group, is looking for a highly skilled platform engineer with deep expertise in operating systems, hardware, GPU, and high-speed networking. In this role, you will design, develop, and optimize Kubernetes clusters that power enterprise AI for the mission customers. The work location is in Bethesda at the Intelligence Community Campus. Primary Responsibilities Kubernetes Cluster Engineering: Design, configure, and maintain enterprise Kubernetes platforms. Collaborate with a multidisciplinary team to define and optimize Kubernetes architecture, ensuring they meet performance, efficiency, and feature requirements. Infrastructure as Code: Develop and manage Infrastructure as Code (IaC) using tools such as Terraform, Salt, Ansible, Bash, Python or similar frameworks. Collaborate with development teams to design and implement secure, automated, and repeatable pipelines (e.g., GitLab CI/CD). Troubleshot complex systems issues across cloud, network, and platform layers. Compliance & Documentation: Maintain technical documentation, architectural specifications, and Linux best practices. Support ATO (Authority to Operate) and ensure compliance with federal security standards.

Qualifications
Basic Qualifications
  • Requires a Bachelor’s degree and 10+ years of relevant experience, or Masters degree with 8+ years of experience. Additional years of experience may be considered in lieu of a degree
  • 5+ years in Platform Engineering or System Engineering experience.
  • Strong expertise with Linux distributions. (RHEL, Ubuntu, Oracle Linux, and Rocky).
  • Experience administering Kubernetes clusters, including deploying, scaling, and maintaining containerized workloads.
  • Hands-on experience creating, managing, and troubleshooting Docker containers and container images throughout the software development lifecycle.
  • Experience with Kubernetes cluster management and AI/ML workflow orchestration (Argo, Airflow, and Kubeflow).
  • Strong track record with consuming, and troubleshooting RESTful APIs for platform integration and automation.
  • Excellent problem-solving skills and the ability to collaborate within a team.
  • Candidate must, at a minimum, meet DoD 8570.11- IAT Level II certification requirements (currently Security+ CE, CCNA-Security, GICSP, GSEC, or SSCP along with an appropriate computing environment (CE) certification). An IAT Level III certification would also be acceptable (CASP+, CCNP Security, CISA, CISSP, GCED, GCIH, CCSP).

Security Clearance: TS/SCI with CI Poly is required for position or a TS/SCI and willingness to obtain a Poly.

Preferred Qualifications
  • Experience in managing NVIDIA GPU data center platforms. (DGX, HGX, H200, H100, 200, B300, L40S).
  • Experience with NVIDIA enterprise tools such as Base Command Manager, Run:AI, Nvidia AI Enterprise.
  • Knowledge of enterprise server components (storage/network controllers, HBA, SSDs).
  • Familiarity with GPU virtualization and cloud computing.
  • Experience developing and deploying infrastructure in AWS.
  • Knowledge of distributed resource scheduling systems. (Slurm , LSF, Open MPI,etc
Additional Information

All your information will be kept confidential according to EEO guidelines.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Systems and Platform Engineer
Senior Systems and Platform Engineer

MAXISIQ, Inc. • Bethesda (MD)

On-site
USD 170,000 - 210,000
Systems and Platform Engineer - TS/SCI
Systems and Platform Engineer - TS/SCI

VMD Corp • Bethesda (MD)

On-site
USD 150,000 - 190,000
Systems and Platform Engineer - TS/SCI
Systems and Platform Engineer - TS/SCI

Xcelerate Solutions • Bethesda (MD)

Hybrid
USD 140,000 - 190,000
Competitive salary
Health benefits
Systems Engineer – HPC & GPU Infrastructure
Systems Engineer – HPC & GPU Infrastructure

MAXISIQ, Inc. • Bethesda (MD)

On-site
USD 170,000 - 210,000
Systems Engineer – HPC & GPU Infrastructure
Systems Engineer – HPC & GPU Infrastructure

FiveInsights • Bethesda (MD)

On-site
USD 170,000 - 210,000
GPU Systems Engineer 4
GPU Systems Engineer 4

Base-2 Solutions • Bethesda (MD)

On-site
USD 180,000 - 230,000
GPU Systems Engineer 4
GPU Systems Engineer 4

RPMGlobal • Bethesda (MD), Northern (KY)

Hybrid
USD 180,000 - 240,000
GPU Systems Engineer 3
GPU Systems Engineer 3

Base-2 Solutions, LLC • Bethesda (MD)

On-site
USD 150,000 - 210,000
GPU Systems Engineer 4
GPU Systems Engineer 4

Base-2 Solutions, LLC • Bethesda (MD)

On-site
USD 190,000 - 230,000
Senior Platform Engineer — AI/ML Kubernetes & Infra
Senior Platform Engineer — AI/ML Kubernetes & Infra

Maxisiq • Bethesda (AR)

On-site
USD 180,000 - 260,000