Technical Success Engineer

Lambda

San Francisco (CA)

On-site

USD 251,000 - 335,000

Full time

16 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision coverage
Wellness and commuter stipends
401k Plan with 2% company match
Flexible paid time off

Job summary

Lambda, The Superintelligence Cloud, seeks a Superintelligence Technical Success Engineer to drive deployments from contract to live production environments. You will validate configurations, coordinate with Infrastructure and Engineering, and ensure customer requirements are met throughout deployment and early production.

Ideal candidates have 4+ years in GPU/HPC, cloud, Kubernetes, or large Linux systems, with strong troubleshooting and communication skills.

Qualifications

  • 4+ years of hands-on technical experience with GPU/HPC infrastructure, cloud platforms, Kubernetes, or large-scale Linux systems.
  • Comfortable validating builds, spotting gaps, and holding engineering teams accountable for resolutions.
  • Strong troubleshooting instincts and willingness to dive into networking, storage, or compute issues.

Responsibilities

  • Take signed deployments from contract to live, fully operational production environments, validating configuration, connectivity, storage, and compute against what was promised.
  • Bring technical depth to validation and troubleshooting conversations, coordinating with engineering and infrastructure teams to drive issues to resolution.
  • Coordinate with Infrastructure, Engineering, Product, and Data Center teams to close technical dependencies and resolve blockers.
  • Own a current, accurate technical picture of the deployment—what's built, what's open, what's at risk.
  • Keep stakeholders informed with clear, regular status updates: RAG status, top risks, and actions.
  • Guide the customer through onboarding to their first successful production workload.
  • Be the customer’s go-to technical contact through deployment and early production.
  • Feed recurring technical patterns back into reusable runbooks, checklists, or automation.
  • Transition out once the customer is stable and self-sufficient, keeping the broader account team informed.

Skills

GPU/HPC infra
Cloud platforms
Kubernetes
Large‑scale Linux
Troubleshooting
Technical communication
Ownership bias

Tools

Runbooks
Status reporting tools

Job description

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

  • Note: This position requires presence in our San Francisco or San Jose office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
About This Role

The Superintelligence Technical Success Engineer is part of Lambda's Superintelligence business unit, dedicated to our largest, most strategic customers operating in the most complex environments. This role owns taking signed deployment from contract to a live, fully operational production environment. The role requires strong technical acumen to validate the build against what was promised, spot gaps, and work effectively with engineering and infrastructure teams to get issues resolved and requirements clearly understood.

You'll work alongside the account team, to make sure the customer's requirements and expectations are clearly understood and addressed throughout deployment. This role calls for someone who's ready to dive in wherever the engagement needs them.

What You'll Do
  • Take signed deployments, from contract, to live, to working production environments, validating configuration, connectivity, storage, and compute against what was promised
  • Bring technical depth to validation and troubleshooting conversations — asking the right questions, spotting gaps, and working closely with engineering and infrastructure teams to drive issues to resolution
  • Coordinate with Infrastructure, Engineering, Product, and Data Center teams to close technical dependencies and resolve blockers
  • Own a current, accurate technical picture of the deployment — what's built, what's open, what's at risk
  • Keep stakeholders informed with clear, regular status updates: RAG status, top risks, and what's being done about them
  • Guide the customer through onboarding to their first successful production workload
  • Be the customer's go-to technical contact through deployment and early production
  • Feed recurring technical patterns back into reusable runbooks, checklists, or automation
  • Transition out once the customer is stable and self-sufficient, keeping the broader account team informed along the way
You
  • 4+ years of hands‑on technical experience with GPU/HPC infrastructure, cloud platforms, Kubernetes, or large‑scale Linux systems
  • Comfortable being the technical voice in the room able to validate builds, spot gaps, and hold engineering teams accountable for resolving them
  • Strong troubleshooting instincts and a willingness to get into the weeds of networking, storage, or compute issues
  • Track record of coordinating across engineering and infrastructure teams to close out technical dependencies
  • Clear written and verbal communication for status updates and technical documentation
  • A bias toward diving in and taking ownership, rather than waiting for a fully defined process
Nice to Have
  • Exposure to large-scale GPU cluster deployments
  • Familiarity with project/program tracking tools and structured status reporting
  • Experience building runbooks or checklists that outlived the engagement they were built for
Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda
  • Founded in 2012, with 500+ employees, and growing fast
  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In‑Q‑Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
  • Our values are publicly available: https://lambda.ai/careers
  • We offer generous cash & equity compensation
  • Health, dental, and vision coverage for you and your dependents
  • Wellness and commuter stipends for select roles
  • 401k Plan with 2% company match (USA employees)
  • Flexible paid time off plan that we all actually use
Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Compensation Range: $251K - $335K

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Success Engineer
Technical Success Engineer

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 190,000
Health, dental, and vision coverage
Wellness stipend
Commuter stipend
+1
Technical Success Engineer
Technical Success Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Health/Dental/Vision
401k with match
Wellness stipend
+2
Technical Success Engineer
Technical Success Engineer

Lambda • San Jose (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage
Wellness stipend
Commuter stipend
+2
Senior Software Engineer – Core Cloud Platform
Senior Software Engineer – Core Cloud Platform

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, and vision coverage
Wellness stipend
Commuter stipend
+1
Senior Site Reliability Engineer - Fleet
Senior Site Reliability Engineer - Fleet

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+4
Senior Site Reliability Engineer - Fleet
Senior Site Reliability Engineer - Fleet

Lambda • San Francisco (CA)

On-site
USD 180,000 - 230,000
Health, dental, and vision
401k with company match
Flexible paid time off
+2
Senior Site Reliability Engineer - Fleet
Senior Site Reliability Engineer - Fleet

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Health, dental, vision coverage
Wellness stipend
Commuter stipend
+2
Senior Site Reliability Engineer - Fleet
Senior Site Reliability Engineer - Fleet

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior Site Reliability Engineer – Core Cloud Platform
Senior Site Reliability Engineer – Core Cloud Platform

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Senior Software Engineer - Core Cloud Platform
Senior Software Engineer - Core Cloud Platform

Lambda • San Jose (CA)

On-site
USD 170,000 - 260,000
Health, dental, and vision coverage
Wellness and commuter stipends
401k Plan with company match
+1