An application made for this job — a tailored resume and cover letter that speak straight to the posting.
MAXISIQ, Inc. seeks a mid-career Systems Engineer to design, develop, and optimize GPU clusters for IC customers. This 100% on-site role is based at Bethesda's Intelligence Community Campus, requiring TS/SCI with polygraph consideration.
You will install HPC/GPU hardware, tune Linux performance, and implement job scheduling with Slurm, PBS, Kubernetes. Strong Linux, hardware, and scripting skills are essential for success.
Our partner, Vexterra Group,is looking for a mid-career Systems Engineer (Mid-Career) – HPC & GPU Infrastructure with a deep understanding of operating systems, hardware, Kubernetes, and NVIDIA GPU products. As a Systems Engineer (Mid-Career) – HPC & GPU Infrastructure, you will play a pivotal role in designing, developing, and optimizing GPU clusters for the IC community customers. This is a 100% on-site position. All work must be performed at the customer site in Bethesda at the Intelligence Community Campus.
Primary Responsibilities
1. HPC and GPU environment engineering: Contribute to the installation and maintenance of GPU
and HPC hardware on-prem and in the cloud, providing insights into hardware performance to
ensure efficient interaction with software components.
2. Performance Optimization: Analyze HPC/GPU cluster performance, identify bottlenecks, and
develop strategies to enhance performance across various applications in Linux, addressing both
hardware and software considerations. Regularly monitor and improve performance.
3. HPC/GPU tooling: Install and configure HPC/GPU job scheduling and workload management
platforms such as Slurm, PBS , Apache Airflow, Kubernetes
4. Power Efficiency: Work on power management techniques to optimize GPU power
consumption, ensuring efficient operation on both mobile and desktop Linux platforms.
Continuously assess and enhance power efficiency strategies.
5. Testing and Validation: Design and execute tests to validate GPU performance and functionality
on Linux, including stress testing, benchmarking, and debugging to ensure robust operation.
Maintain and expand the testing suite.
6. Documentation: Maintain comprehensive technical documentation, including architectural
specifications, code documentation, and Linux-specific best practices for GPU development.
Keep documentation up to date with changes and improvements.
7. Industry Insight: Stay updated on the latest trends, innovations, and competitive landscapes
within the GPU industry, contributing to research efforts and proposing Linux-specific
approaches to GPU design and optimization. Share regular updates and insights with the team.
Basic Qualifications
Preferred Qualifications
All your information will be kept confidential according to EEO guidelines.