Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company

Sunnyvale (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Health insurance
Flexible work hours

Job summary

A leading AI technology firm in California is seeking an experienced Senior Software Engineer to develop and optimize AI infrastructure software using state-of-the-art GPU systems. Candidates should have a Bachelor's degree in a technical field and a minimum of 5 years of experience in software engineering, especially with AI frameworks. This role involves designing robust software for large-scale AI workloads and collaborating with cross-functional teams to innovate and drive product delivery.

Qualifications

  • 5+ years of experience in software engineering and infrastructure development.
  • Proven hands-on experience building AI systems.
  • Experience with CUDA, vLLM, or advanced inference engines.

Responsibilities

  • Design and develop robust systems software for AI workloads.
  • Deliver critical components for workload scheduling and orchestration.
  • Optimize AI workloads for extreme performance on large-scale GPU systems.

Skills

Python
C/C++
Machine Learning
Deep Learning
Kubernetes
AI/ML frameworks (PyTorch, TensorFlow)
Distributed Systems
AI frameworks

Education

Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field
Master's or PhD in Computer Science, AI/ML, or a related discipline

Tools

NVIDIA GB200
TensorRT
MLflow
Kubeflow

Job description

Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distributed AI-RAN Environments)

We challenge conventional limits by building transformative products that fully exploit our state-of-the-art (SOTA) infrastructure—including NVIDIA GB200 (e.g., liquid‑cooled NVL72 rack‑scale systems with 72 Blackwell GPUs and 36 Grace CPUs unified via massive NVLink domains), MGX modular architectures, and DGX Grace Hopper platforms—combined with cloud‑native software. These solutions power massive centralized AI data centers and pioneering distributed AI Radio Access Network (AI‑RAN) deployments, where AI is embedded directly into radio networks for superior efficiency, edge intelligence, and new capabilities.

We're seeking experienced senior engineers passionate about innovation, ready to shape foundational AI infrastructure software on world‑class GPU systems.

Minimum Qualifications
  • Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field.
  • 5+ years of experience in software engineering, hardware platforms, distributed systems, or infrastructure development.
  • 2+ years in technical lead roles, owning high‑impact projects and leading teams.
  • Proven hands‑on experience building systems software, AI frameworks, or applied AI systems.
Preferred Qualifications
  • Master's degree in a relevant technical discipline (e.g., CS, Systems Engineering).
  • Direct experience with Kubernetes and container orchestration at scale.
  • Hands‑on work with GPU‑accelerated systems and high‑performance computing (HPC) environments.
  • Familiarity with AI developer frameworks, MLOps tools, automation pipelines, and CI/CD systems.
Role Overview

Join our infrastructure team as a key contributor building foundational systems software atop next‑generation GPU platforms to support large‑scale AI workloads (training, fine‑tuning, and serving). Own significant portions of our new AI infrastructure stack, with a strong emphasis on Kubernetes orchestration and GPU resource management. Drive architectural innovation in systems software and automation to achieve maximum efficiency and utilization. As a Senior Software Engineer, collaborate closely with Engineering Leads, Product Management, and Program teams to execute toward commercialization and market impact.

Key Responsibilities
  • Design and develop robust systems software to enable AI workloads on large‑scale GPU clusters (e.g., GB200 NVL72 rack‑scale systems).
  • Deliver critical control plane components for workload scheduling, orchestration, and resource management; build management plane for underlying hardware platforms.
  • Create northbound APIs and interfaces for customer portals and self‑service access to the infrastructure.
  • Contribute to product requirements documents (PRDs), sprint planning, and agile program execution.
  • Help attract, mentor, and grow top engineering talent.
  • Exemplify and cultivate a culture of humility, bold innovation, and disciplined delivery to bring products to market.

We are pushing the boundaries of AI by challenging conventional approaches and building groundbreaking products that fully harness our state‑of‑the‑art (SOTA) infrastructure—including NVIDIA GB200 NVL72 (liquid‑cooled, 72‑GPU rack‑scale systems with massive NVLink domains for exascale performance), MGX modular architectures, and DGX Grace Hopper platforms. These products power both massive centralized AI data centers and innovative distributed AI Radio Access Network (AI‑RAN) deployments, where AI is natively integrated into radio networks for enhanced efficiency, edge intelligence, and new revenue opportunities.

We're seeking passionate, experienced practitioners eager to drive innovation and deliver transformative AI products on world‑class hardware.

Minimum Qualifications
  • Bachelor's degree in Computer Science, Electrical Engineering, Mathematics, Statistics, or a related technical field.
  • 3+ years of hands‑on experience in machine learning, deep learning, and software engineering.
  • Proficiency in Python; experience with C/C++.
  • Strong working knowledge of major AI/ML frameworks (PyTorch, TensorFlow, JAX, or similar).
  • Solid foundation in data structures, algorithms, and software design principles.
Preferred Qualifications
  • Master's or PhD in Computer Science, AI/ML, or a related discipline.
  • Experience with Large Language Models (LLMs), Generative AI, or Computer Vision.
  • Familiarity with distributed training frameworks and techniques (e.g., Ray, DeepSpeed, Megatron‑LM).
  • Proven expertise optimizing models for GPU inference (e.g., TensorRT, Triton Inference Server).
  • Knowledge of MLOps tools and practices (Kubeflow, MLflow, etc.).
Role Overview

As a core member of our AI engineering team, you will design, develop, and optimize cutting‑edge AI models and workloads that run natively on our high‑performance GPU clusters. Leverage our SOTA infrastructure to train, fine‑tune, and serve massive‑scale models at unprecedented efficiency. Collaborate across infrastructure, product, and research teams to align hardware capabilities with real‑world AI demands, driving breakthroughs in performance, scalability, and innovation.

Key Responsibilities
  • Design, implement, and train state‑of‑the‑art ML models for high‑impact applications (e.g., NLP, Computer Vision, Network Optimization).
  • Optimize AI workloads for extreme performance and scalability on large‑scale GPU systems like GB200 NVL72, using tools such as Dynamo, vLLM, and advanced inference engines.
  • Partner with cross‑functional teams to co‑design hardware‑software solutions that maximize AI processing efficiency.
  • Build robust tools, data pipelines, evaluation frameworks, and deployment systems.
  • Track and incorporate the latest AI research and technological advancements.
  • Contribute to product requirements (PRDs) and agile execution (sprint planning and delivery).
  • Champion a culture of humility, bold innovation, and high‑velocity product delivery.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Senior DGX Cloud AI Infrastructure Software Engineer
Senior DGX Cloud AI Infrastructure Software Engineer

NVIDIA • United States

On-site
USD 184,000 - 357,000
Equity
Benefits
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Senior Solutions Architect, Generative AI
Senior Solutions Architect, Generative AI

NVIDIA • California (MO)

Hybrid
USD 184,000 - 357,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 213,000 - 288,000
Early-stage equity
Direct access to leadership
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale - DGX Cloud
Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale - DGX Cloud

NVIDIA • Seattle (WA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Solutions Architect, AI Factory Infrastructure
Solutions Architect, AI Factory Infrastructure

NVIDIA • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits