An application made for this job — a tailored resume and cover letter that speak straight to the posting.
NVIDIA is hiring an architect to reshape AI inference with disaggregated systems and scalable reference architectures. You will guide partners, design end-to-end inference pipelines, and optimize performance across GPU clusters at scale.
You will mentor teams, drive strong customer engagements, and contribute to the deployment of cutting-edge NVIDIA Dynamo, Triton Inference Server, and TensorRT-LLM based solutions. Equity and comprehensive benefits are provided.
NVIDIA is hiring an architect to reshape AI inference with disaggregated systems and scalable reference architectures. You will guide partners, design end-to-end inference pipelines, and optimize performance across GPU clusters at scale.
You will mentor teams, drive strong customer engagements, and contribute to the deployment of cutting-edge NVIDIA Dynamo, Triton Inference Server, and TensorRT-LLM based solutions. Equity and comprehensive benefits are provided.