Senior AI Middleware Engineer for High-Performance Training
Cornelis Networks
Austin (TX)
On-site
USD 120,000 - 150,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Equity, cash, and incentives
Health coverage
401(k) with company match
Open Time Off
Paid holidays
Job summary
Cornelis Networks is seeking experienced engineers in Austin, Texas, to work on cutting-edge networking solutions for AI and HPC. The role focuses on performance-critical feature implementation, distributed training optimization, and profiling workloads. Candidates should possess over 8 years of experience in high-performance systems programming with strong knowledge of GPU communication stacks. The position is remote for U.S. residents and offers a competitive compensation package, including health and retirement benefits.
Qualifications
8+ years of experience in high-performance systems programming in C/C++ on Linux.
Strong experience with GPU communication stacks including CUDA/ROCm and NCCL/RCCL.
Ability to optimize distributed training performance using profiling and tracing.
Experience delivering production-quality code.
Responsibilities
Design and implement performance-critical features for CCL enablement.
Optimize distributed training performance across multi-node, multi-GPU configurations.
Improve GPU communication paths including GPU-direct transfers.
Profile distributed AI workloads and identify bottlenecks.
Skills
High-performance systems programming in C/C++
GPU communication stacks including CUDA/ROCm
Optimizing distributed training performance
Open-source contributions
Tools
AI frameworks such as PyTorch Distributed
CUDA
NCCL
Job description
Cornelis Networks is seeking experienced engineers in Austin, Texas, to work on cutting-edge networking solutions for AI and HPC. The role focuses on performance-critical feature implementation, distributed training optimization, and profiling workloads. Candidates should possess over 8 years of experience in high-performance systems programming with strong knowledge of GPU communication stacks. The position is remote for U.S. residents and offers a competitive compensation package, including health and retirement benefits.