- Explore, research, and prototype systems optimizations for advanced deep learning models at the intersection of high-level deep learning frameworks and low-level CUDA.
- Architect and optimize distributed computing systems from single-node to cluster-scale supercomputing environments.
- Design, implement, and optimize custom high-performance CUDA kernels for emerging neural network architectures and workloads.
- Analyze hardware-software interactions to identify and resolve performance bottlenecks in training and inference pipelines.
- Collaborate with AI researchers, hardware and software architects, kernel and compiler authors, and CUDA driver experts to co-design systems and algorithms.
- Develop exploratory tools and runtime systems to profile and accelerate new deep learning paradigms.
- Write clean, effective, and maintainable code and transition prototypes into open-source releases, framework integrations, internal tools, or commercial products.
Requirements
- BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience.
- 2+ years of relevant industry experience or equivalent academic experience after degree achievement.
- Strong proficiency in C++ and Python programming.
- Solid background in deep learning fundamentals, focused on transformers.
- Strong understanding of distributed computing, multi-node scaling, and cluster-scale performance challenges.
- Proven experience in systems programming, computer architecture, and low-level systems performance optimization.
- Familiarity with GPU deep learning accelerator architectures.
- Hands-on experience with CUDA programming, kernel optimization, and workload profiling.
- Experience profiling and optimizing generative AI models, including large language models.
- Research background in machine learning systems or adjacent fields.
- Experience profiling and optimizing innovative vision models, generative AI architectures, or diffusion models.
- Track record of initiative and willingness to deep-dive on problems across the stack.
- Preferred experience with PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, or Megatron internals and execution graphs.
- Preferred hands-on experience with NCCL, MPI, or UCX and distributed machine learning techniques.
- Preferred knowledge of numerical methods and low-precision arithmetic such as NVFP4, MXFP4, FP8, and INT8.
- Preferred background in deep learning compilers and ML systems, including Triton, XLA, and torch.compile.
- Preferred experience with highly parallel or reinforcement-learning-style simulation environments.
- Preferred experience designing and implementing agentic AI systems for complex systems and infrastructure problems.
Core Competencies
Demonstrates expertise in CUDA programming, high-performance computing, and deep learning optimization, with a strong foundation in distributed systems and machine learning architectures. Capable of collaborating with cross-functional teams to design and implement innovative solutions for advanced AI models.
Highest-signal resume keywords
- CUDA Programming
- C++ and Python Proficiency
- Deep Learning Optimization
- Distributed Computing Systems
- Performance Bottleneck Analysis
ATS Optimization Keywords
Hard Skills
- CUDA
- C++
- Python
- Deep Learning Fundamentals
- Systems Programming
- Computer Architecture
- Kernel Optimization
- Workload Profiling
- Generative AI Models
- Transformers
Soft Skills
- Collaboration
- Initiative
- Problem-Solving
Industry Keywords
- Machine Learning Systems
- High-Performance Computing
- AI Research
- Neural Network Architectures
- Cluster-Scale Performance
Tools & Technologies
- PyTorch
- JAX
- TensorRT
- NCCL
- MPI
- UCX
- Triton
- XLA
- Torch.compile
- SgLang