Job Title : Platform Engineer
Note : US Citizen or Green Card candidates
Duration : 160 hours
Location: St. Louis, MO
100% Onsite
Role Overview
Highly skilled AI Infrastructure Engineer to design, build, and operate scalable GPU-enabled Kubernetes platforms for AI/ML workloads.
Job Title : Platform Engineer
Note : US Citizen or Green Card candidates
Duration : 160 hours
Location: St. Louis, MO
100% Onsite
Role Overview
Highly skilled AI Infrastructure Engineer to design, build, and operate scalable GPU-enabled Kubernetes platforms for AI/ML workloads.
Must have skills:
Kubernetes, Linux, Terraform, GPUs and NVIDIA stack
Required Qualifications
- 3-8+ years' experience
- Strong Kubernetes knowledge
- Experience with GPUs and NVIDIA stack
- Linux (Ubuntu) expertise
- Experience with Terraform
Preferred Qualifications
- Longhorn or Ceph experience
- Canonical ecosystem (MAAS, Juju)
- AI/ML tools like Kubeflow
- Certifications (CKA, NVIDIA)
Soft Skills
- Strong problem-solving and troubleshooting mindset
- Ability to collaborate with cross-functional teams (ML engineers, data scientists)
- Clear communication and documentation skills
- Passion for automation and platform scalability
Key Responsibilities
- Design and manage Kubernetes clusters
- Build GPU-enabled infrastructure
- Deploy Longhorn storage
- Automate infrastructure using Terraform
- Monitor systems using Prometheus and Grafana
- Knowledge Transfer & Client Enablement
- Provide structured knowledge transfer (KT) sessions to client teams on all core platform components, including:
- Kubernetes architecture, operations, and troubleshooting
- GPU infrastructure (NVIDIA stack, scheduling, resource optimization)
- Longhorn storage management and performance tuning
- Canonical ecosystem tools (MAAS, Juju, Charmed Kubernetes)
- Develop and deliver technical documentation, runbooks, and training materials to support ongoing operations
- Conduct hands‑on workshops and guided sessions to enable client teams to independently manage and scale the platform
- Act as a technical advisor, helping client stakeholders understand best practices in:
- Cloud‑native infrastructure
- AI/ML platform operations
- Reliability, performance, and cost optimization
- Ensure smooth handoff of production systems with full operational readiness and support knowledge