Compensation: Competitive Base Salary + Performance Bonus
Overview
Our client is seeking an HPC & AI Solutions Architect to lead the technical design, integration, and delivery of high-performance computing and AI infrastructure solutions.
This is a highly technical, customer-facing role focused on designing scalable architectures across GPU/CPU compute, storage, networking, Kubernetes, orchestration, and security. The position spans the full solution lifecycle from technical discovery and workload analysis through proof-of-concept, deployment, and ongoing optimization.
The ideal candidate brings deep HPC and AI infrastructure expertise, strong hands-on system design and performance tuning experience, and the ability to translate complex customer requirements into scalable, production-ready solutions.
Key Responsibilities
Customer Engagement & Technical Discovery
- Work directly with customers to understand workload requirements, performance targets, and technical objectives.
- Lead technical discovery sessions focused on application behavior, bottlenecks, scalability, and infrastructure requirements.
- Serve as a trusted technical advisor throughout the solution lifecycle.
- Design end-to-end HPC and AI architectures across compute, storage, networking, orchestration, and security.
- Recommend hardware and software solutions aligned with performance, scalability, reliability, and efficiency goals.
- Develop architecture blueprints, integration plans, and technical documentation.
- Design solutions supporting GPU-intensive AI/ML, LLM, and advanced compute workloads.
Performance & Workload Optimization
- Support proof-of-concept, benchmarking, and performance-validation initiatives.
- Perform workload profiling, system tuning, and infrastructure optimization.
- Identify bottlenecks across compute, storage, networking, and orchestration layers.
- Recommend improvements that increase workload performance, scalability, and resilience.
Implementation & Delivery
- Provide technical leadership during deployment and integration.
- Partner with customers and internal Engineering, Product, and Operations teams throughout implementation.
- Support solutions from architecture through production deployment and optimization.
- Troubleshoot complex infrastructure and workload issues during delivery.
Technical Leadership
- Maintain expertise across emerging HPC, AI, GPU, storage, networking, and orchestration technologies.
- Build relationships with technology partners across GPU, networking, and storage ecosystems.
- Contribute to reference architectures, reusable design patterns, and technical best practices.
- Lead customer workshops, architecture reviews, and technical presentations.
Required Qualifications
- Strong technical expertise across:
- GPU and CPU architectures
- NVIDIA / CUDA ecosystem
- Slurm and Kubernetes
- InfiniBand, RDMA, and RoCE
- Lustre, GPFS / Spectrum Scale, Ceph, VAST, or similar storage platforms
- Kubernetes and container orchestration
- Identity, encryption, and infrastructure security
- Strong Linux systems knowledge, including tuning and performance analysis.
- Experience translating workload requirements into detailed technical architectures.
- Experience with proof-of-concept, benchmarking, or workload optimization.
- Strong customer-facing communication and presentation skills.
- Ability to work effectively with engineering, product, operations, and executive stakeholders.
Preferred Experience
- AI/ML, LLM, GPU, or HPC workloads.
- NVIDIA GPU infrastructure.
- Automation and Infrastructure-as-Code.
- Workload migration and performance engineering.
- Next-generation GPU and high-speed interconnect technologies.
- Bachelor's or Master's degree in Computer Science, Engineering, Physics, or related field.
- Relevant cloud, Linux, networking, Kubernetes, or security certifications.