Get more replies from employers
Send a job-specific resume in minutes.
MRE Consulting in Houston seeks a HPC AI Systems Administrator to architect a secure, scalable on-prem HPC compute platform, enabling your development teams to fine-tune and deploy production ML models. You will lead deployment, manage multi-GPU hardware, and oversee Linux, GPU drivers, CUDA, NCCL, containers, and orchestration tools, ensuring performance and security.
This role requires 3+ years in HPC or enterprise GPU infra, strong Linux, and experience with Slurm/Kubernetes, InfiniBand
We are seeking a high-caliber HPC AI Systems Administrator to serve as the foundational architect for our growing AI infrastructure. This role will be responsible for building a secure, scalable, and highly optimized environment to support our corporate data initiatives.
Operating at the critical intersection of infrastructure engineering and software application, you will design and maintain a robust compute platform. Your primary mission is to enable our Development Team to fine-tune and deploy production-level machine learning models smoothly, while ensuring the platform complies with enterprise-level security, governance, and data privacy policies.