Senior AI Infrastructure Engineer - ML Accelerators

Socket.dev

Sunnyvale (CA)

On-site

USD 207,000 - 300,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

High-value equity

Job summary

Socket.dev is seeking a software engineer with deep expertise in C++, GPU programming, and distributed systems to help scale enterprise AI infrastructure. The role emphasizes leadership, architecture, and collaboration with hardware teams to deliver reliable accelerator software across TPUs and GPUs.

The ideal candidate has multi-year experience in large-scale systems, machine learning infrastructure, and cross-functional programs, with a record of technical leadership and project direction.

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience programming in C++.
  • 5 years of experience testing and launching software products.
  • 5 years of building large-scale infrastructure, distributed systems or networks, or compute/storage/hardware architecture.
  • 3 years of experience with software design and architecture.
  • Experience with GPU programming, ML infrastructure, AI platform, ML optimization, and systems infrastructure.

Responsibilities

  • Define architecture and long-term technical roadmap for accelerator software stacks.
  • Collaborate with HW engineering to design, test, deploy, and debug low-level software across the stack.
  • Provide technical leadership and mentorship to a distributed team.
  • Manage priorities and end-to-end deliverables powering Google's AI hyper-computers and data centers.
  • Drive large-scale programs to maximize production health, reliability, and scalability.

Skills

C++ programming
Software testing
Large-scale infra
Software architecture
GPU programming
ML infrastructure
AI platforms
ML optimization

Education

Bachelor's degree
Master's degree

Job description

Socket.dev is seeking a software engineer with deep expertise in C++, GPU programming, and distributed systems to help scale enterprise AI infrastructure. The role emphasizes leadership, architecture, and collaboration with hardware teams to deliver reliable accelerator software across TPUs and GPUs.

The ideal candidate has multi-year experience in large-scale systems, machine learning infrastructure, and cross-functional programs, with a record of technical leadership and project direction.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Accelerator Systems Architect — Transformation Leader
AI Accelerator Systems Architect — Transformation Leader

Socket.dev • Sunnyvale (CA)

On-site
USD 262,000 - 364,000
Senior AI Systems Architect - GPU & Platform Internals
Senior AI Systems Architect - GPU & Platform Internals

Accellor • San Francisco (CA)

On-site
USD 210,000 - 320,000
Lead GPU Architect for AI Accelerators & Clusters
Lead GPU Architect for AI Accelerators & Clusters

EngineersOfAI • Milpitas (CA)

On-site
USD 140,000 - 190,000
Staff ML Systems Engineer for GPU Accelerators
Staff ML Systems Engineer for GPU Accelerators

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Senior AI Platform Architect — GPU & Inference
Senior AI Platform Architect — GPU & Inference

Accellor • Mountain View (CA)

On-site
USD 180,000 - 280,000
Senior AI Accelerator Performance Leader
Senior AI Accelerator Performance Leader

EPAM Systems • New York (NY)

On-site
USD 180,000 - 220,000
Medical, Dental and Vision Insurance
Health Savings Account
Flexible Spending Accounts
+10
Senior AI Accelerator Architect for ML Systems
Senior AI Accelerator Architect for ML Systems

Intel • Hillsboro (OR)

On-site
USD 171,000 - 315,000
Senior AI/ML HPC Systems Engineer - Accelerator Servers
Senior AI/ML HPC Systems Engineer - Accelerator Servers

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 173,000 - 236,000
Sign-on bonuses and RSUs
Senior AI Infrastructure Engineer - GPU Compute
Senior AI Infrastructure Engineer - GPU Compute

Unchain Data • United States

On-site
USD 120,000 - 160,000
Senior Distributed AI/ML Systems Engineer
Senior Distributed AI/ML Systems Engineer

Socket.dev • Cupertino (CA)

On-site
USD 193,300 - 261,500