Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
NVIDIA Corporation seeks a Senior Systems Software Engineer to build high-performance on-device AI software for RTX and DGX-class systems. You will focus on low-latency inference, memory efficiency, and robust deployment on resource-constrained platforms.
Collaborate with software, research, architecture, and product teams to optimize inference stacks using llama.cpp, vLLM, PyTorch, Windows ML, DXCGC, and TensorRT across diverse hardware. Strong C++ and ML background are essential.
NVIDIA has continuously reinvented itself for more than two decades. The invention of the GPU in 1999 fueled the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. GPU-powered deep learning helped ignite the era of modern AI, establishing GPUs as the foundation of intelligent applications across productivity, gaming, and creative workflows, and reinforcing NVIDIA’s position as a leading AI computing company. More recently, there is a growing focus on running AI models locally, closer to where data is generated. This approach reduces latency, enables real-time processing, and addresses privacy concerns by minimizing the need to send data to centralized servers. As technology continues to evolve, client-side AI will play an increasingly important role in shaping the digital landscape. The LocalAI team is seeking a Senior Systems Software Engineer to develop efficient on-device AI software for RTX and DGX-class systems. The role focuses on delivering high-performance local inference with low latency, optimized memory utilization, robust infrastructure, and practical deployment on resource-constrained platforms.
Partner with NVIDIA’s software, research, architecture, and product teams to align technical requirements and strategic priorities, fostering the AI ecosystem on RTX and DGX PCs. Build and optimize the local AI inference stack for RTX, RTX Pro, and DGX GPUs, with a focus on performance, stability, and scalability across diverse hardware architectures. Design and develop modern inference runtimes and execution stacks using frameworks such as llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX, supporting LLM, vision-language, TTS, ASR, and diffusion-based AI workloads. Perform end-to-end optimization of AI models, data pipelines, and inference runtimes to maximize performance on current and next-generation GPU architectures. Apply model optimization techniques, including quantization, pruning, sparsity, and distillation, to enable efficient deployment of large models on local and edge devices. Conduct system-level debugging, performance tuning, and performance-accuracy trade-off analysis; develop infrastructure for performance and accuracy sweeps; analyse results to identify gaps and drive fixes; and establish engineering guidelines to accelerate bring-up and ensure production readiness of new models and inference backends.
We're a top employer recognized for innovation, growth, and a commitment to diversity as an equal-opportunity workplace.
As our engineering teams continue to grow rapidly, we're looking for creative, self-driven engineers with a passion for technology to join us.
NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA.