Customer Solutions Engineer, Compute, Google Cloud

Jobs in JS

Kirkland (WA)

On-site

USD 102,000 - 144,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

Google Cloud is seeking Solutions Engineers for AI Infrastructure to own complex customer issues and provide expert support across teams. You will troubleshoot hardware and software, networking, and Linux administration to enable AI/ML workloads on Google Cloud AI Infrastructure products.

You will work with Engineering, Sales, and customer organizations to resolve deployment obstacles, develop diagnostic tools, and contribute to product improvements.

Qualifications

  • Bachelor's degree in Science, Technology, Engineering, Mathematics, or equivalent practical experience.
  • Experience reading/debugging code in a general purpose language and in virtualization/orchestration frameworks.
  • Experience troubleshooting and triaging customer issues across the stack.
  • System administrator level experience with Linux/Unix and debugging on enterprise-grade servers.

Responsibilities

  • Diagnose and resolve customer issues on AI/ML infrastructure through tools and new investigation methods.
  • Build understanding of AI/ML workloads and hardware architectures by debugging, reproducing, and root-cause analysis.
  • Consult with Engineering, Sales, and customer teams to resolve deployment and operational challenges.
  • Collaborate with Product and Engineering teams and SRE to improve products and production quality.
  • Be available for non-standard hours and weekend shifts as needed.

Skills

Code reading/debugging
Troubleshooting
Linux system administration
Hardware/software boundary debugging

Education

Bachelor's degree in STEM or equivalent

Tools

Kubernetes
Slurm

Job description

Minimum qualifications:
  • Bachelor's degree in Science, Technology, Engineering, Mathematics, or equivalent practical experience.
  • Experience in reading/debugging code written in a general purpose coding language (e.g., Java, C, C++, Python, Shell, Go or JavaScript, etc.) and in virtualization and orchestration frameworks.
  • Experience troubleshooting and advocating for customer needs, and triaging technical issues across the stack (e.g., hardware faults, low-level software, networking, virtualization, kernel drivers, firmware, performance).
  • System administrator level experience with Linux/Unix systems and experience in debugging issues across the hardware/software boundary on enterprise-grade server infrastructure.
Preferred qualifications:
  • Experience working with large-scale distributed systems, and familiarity with common solutions, design patterns, or best practices.
  • Experience working directly with AI/ML computing hardware, including GPUs or other accelerators.
  • Experience with ML frameworks (e.g., TensorFlow, PyTorch), and understanding of the AI/ML training and inference lifecycle.
  • Familiarity with containerization and orchestration technologies like Kubernetes or Slurm in an on-prem or cloud environment.
About the job:

Our Solutions Engineers for AI Infrastructure own complex customer issues and provide specialized support to other teams. In this role, you will be a part of a global team that provides 24x7 support to ensure customers can seamlessly deploy their AI and ML workloads on AI Infrastructure products. When customers encounter deep technical issues, your job is to ensure we have the expertise, tools, and processes to resolve the issue. You will troubleshoot technical problems with a mix of hardware and software debugging, networking, Linux system administration, coding/scripting, and updating documentation. You will help our customer’s success in the AI/ML space by making improvements to the product, internal tools, processes, and documentation. You'll help drive business growth by recognizing and advocating for our customers’ challenges related to AI deployments.

Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $102000 - $144000 (USD) + 15% bonus target + equity + benefits

Learn more about benefits at Google.

Responsibilities:
  • Manage customers' problems through effective diagnosis, resolution, or implementation of new investigation tools to increase productivity for customer issues on AI/ML infrastructure.
  • Develop an in-depth understanding of AI/ML workloads and underlying hardware architectures by troubleshooting, reproducing, determining the root cause for customer reported issues, and building tools for faster diagnosis.
  • Act as a consultant and subject matter expert for internal stakeholders in Engineering, Sales, and customer organizations to resolve complex deployment and operational obstacles in AI infrastructure environments.
  • Work closely with multiple Product and Engineering teams to find ways to improve the product, and interact with our Site Reliability Engineering (SRE) teams to drive high-quality production.
  • Be available for non-standard work hours or shifts which may include weekends as needed.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Field Application Engineer, Cloud AI Infrastructure
Field Application Engineer, Cloud AI Infrastructure

Google • Kirkland (WA)

On-site
USD 132,000 - 189,000
Health/dental/vision insurance
401(k) with company match
Paid time off 20 days/year
+3
Field Solutions Architect IV, AI Infrastructure, Google Cloud
Field Solutions Architect IV, AI Infrastructure, Google Cloud

Google • Atlanta (GA)

On-site
USD 233,000 - 324,000
Field Solutions Architect IV, AI Infrastructure, Google Cloud
Field Solutions Architect IV, AI Infrastructure, Google Cloud

Google • Sunnyvale (CA)

On-site
USD 233,000 - 324,000
Health Insurance
401(k) with company match
Paid time off 20 days
+4
Field Application Engineer Manager, Cloud AI Infrastructure
Field Application Engineer Manager, Cloud AI Infrastructure

Google • Austin (TX)

On-site
USD 236,000 - 329,000
Health insurance
Dental insurance
Vision insurance
+8
Field Application Engineer Manager, Cloud AI Infrastructure
Field Application Engineer Manager, Cloud AI Infrastructure

Google • Kirkland (WA)

On-site
USD 236,000 - 329,000
Health and welfare benefits
401(k) with company match
Paid time off 20 days
+4
Platform Application Engineer, Cloud AI Infrastructure
Platform Application Engineer, Cloud AI Infrastructure

Google • Austin (TX)

On-site
USD 159,000 - 230,000
Health insurance
Dental insurance
Vision insurance
+8
Platform Application Engineer, Cloud AI Infrastructure
Platform Application Engineer, Cloud AI Infrastructure

Google • Kirkland (WA)

Hybrid
USD 159,000 - 230,000
Health insurance
401(k) with company match
Paid time off
+4
Field Solutions Architect IV, AI Infrastructure, Google Cloud
Field Solutions Architect IV, AI Infrastructure, Google Cloud

Google • Chicago (IL)

On-site
USD 233,000 - 324,000
Health insurance
401(k) match
Paid time off
Customer Engineer, Cloud AI, Google Cloud
Customer Engineer, Cloud AI, Google Cloud

Google • Sunnyvale (CA)

On-site
USD 152,000 - 221,000
Customer Engineer, Cloud AI, Google Cloud
Customer Engineer, Cloud AI, Google Cloud

Google Inc. • Sunnyvale (CA), San Francisco (CA)

On-site
USD 152,000 - 221,000