Senior Software Engineer, AI/ML System Infrastructure

Google

Sunnyvale (CA)

On-site

USD 174,000 - 252,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
Life insurance
Disability insurance
401(k) with company match
Paid time off
Sick time
Holidays

Job summary

Google is seeking a Senior Software Engineer to drive software development for Tensor Processing Unit (TPU) system control planes in Sunnyvale, CA. You will design and implement health management systems that rely on hardware telemetry and build analytics for detecting hardware problems across TPU clusters and cloud infrastructure.

You will collaborate with teams to implement health rules for failure modes, generating repair actions and supporting TPU cluster operations.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 5 years of experience working with Go.
  • 3 years of experience with developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage or hardware architecture.
  • 3 years of experience in distributed computing.
  • 3 years of experience in infrastructure design.
  • 3 years of experience in system architecture.

Responsibilities

  • Build and develop systems to integrate in continuous running of TPU AI infrastructure.
  • Work cross-functionally to define the requirements for Diagnoser, and how will the customer benefit from it and define the Critical User Journeys (CUJs).
  • Influence and align the cross-functional teams on roadmap of PodCare and Diagnoser.
  • Deliver high quality code and timely project while working with cross-functional teams.
  • Contribute to CI/CD pipeline and integration and regression test for health monitoring system.

Skills

Go
Distributed systems
Infrastructure design
System architecture
C
C++
Data structures and algorithms

Education

Bachelor's degree
Master's degree or PhD

Job description

In accordance with Washington state law, we are highlighting our comprehensive benefits package, which is available to all eligible US based employees. Benefits for this role include:

  • Health, dental, vision, life, disability insurance
  • Retirement Benefits: 401(k) with company match
  • Paid Time Off: 20 days of vacation per year, accruing at a rate of 6.15 hours per pay period for the first five years of employment
  • Sick Time: 40 hours/year (increased to 69 hours/year for Seattle) including 5 discretionary sick days per instance
  • Maternity Leave (Short-Term Disability + Baby Bonding): 28-30 weeks
  • Baby Bonding Leave: 18 weeks
  • Holidays: 13 paid days per year

Note: By applying to this position you will have an opportunity to share your preferred working location from the following: Sunnyvale, CA, USA; Kirkland, WA, USA

Minimum qualifications
  • Bachelor’s degree or equivalent practical experience.
  • 5 years of experience working with Go.
  • 3 years of experience with developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage or hardware architecture.
  • 3 years of experience in distributed computing.
  • 3 years of experience in infrastructure design.
  • 3 years of experience in system architecture.
Preferred qualifications
  • Master's degree or PhD in Computer Science or related technical field.
  • 5 years of experience with data structures and algorithms.
  • 5 years of experience with C and C++.
  • 1 year of experience in a technical leadership role.
  • Experience developing accessible technologies.
About the job

Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day.

As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.

As the Senior Software Engineer, you will drive software development for Tensor Processing Unit (TPU) system control planes. You will design and implement health management systems that rely on hardware telemetry. You will build analytics for detecting hardware problems and develop different health rules specializing in detecting ICI, OCS, EMD, Tross, and out of band entities failure modes. You will build algorithms to generate correlated failures and suggest actions for repair workflows, integrating with both TPU cluster and cloud infrastructure. You will also build the Diagnoser for the specialized health rules to maintain and operate these TPU clusters.

Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits

Responsibilities
  • Build and develop systems to integrate in continuous running of TPU AI infrastructure.
  • Work cross-functionally to define the requirements for Diagnoser, and how will the customer benefit from it and define the Critical User Journeys (CUJs).
  • Influence and align the cross-functional teams on roadmap of PodCare and Diagnoser.
  • Deliver high quality code and timely project while working with cross-functional teams.
  • Contribute to CI/CD pipeline and integration and regression test for health monitoring system.

Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, AI/ML System Infrastructure
Senior Software Engineer, AI/ML System Infrastructure

Google • Kirkland (WA)

On-site
USD 174,000 - 252,000
Health insurance
Retirement benefits
Paid time off (PTO)
+4
Senior Software Engineer, AI/ML System Infrastructure
Senior Software Engineer, AI/ML System Infrastructure

Google Inc. • Kirkland (WA), Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Health insurance
Retirement plan
Paid time off
+3
Senior Software Engineer, AI/ML System Infrastructure
Senior Software Engineer, AI/ML System Infrastructure

Google • United States

On-site
USD 174,000 - 252,000
Health insurance
Dental insurance
Vision insurance
+8
Software Engineer III, Infrastructure, Platforms Infrastructure Engineering
Software Engineer III, Infrastructure, Platforms Infrastructure Engineering

Google • New York (NY)

On-site
USD 147,000 - 210,000
Health, dental, vision, life, and/or其它
401(k) with company match
Paid Time Off: 20 days/year
+4
Software Engineer III, Infrastructure, Platforms Infrastructure Engineering
Software Engineer III, Infrastructure, Platforms Infrastructure Engineering

Google • Sunnyvale (CA)

On-site
USD 147,000 - 210,000
Health insurance
Retirement Benefits: 401(k) with match
Paid Time Off: 20 days per year, 6.15h
+4
Senior Software Engineer, AI/TPU Infrastructure Firmware, On-prem
Senior Software Engineer, AI/TPU Infrastructure Firmware, On-prem

Google Inc. • Kirkland (WA), Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Health insurance
Dental insurance
Vision insurance
+8
Senior Technical Program Manager, AI and Infrastructure, TPU
Senior Technical Program Manager, AI and Infrastructure, TPU

Google Inc. • Kirkland (WA), Sunnyvale (CA)

On-site
USD 192,000 - 278,000
Health insurance
Dental insurance
Vision insurance
+8
Staff Software Engineer, TPU Performance
Staff Software Engineer, TPU Performance

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Senior Software Engineer, AI/ML, AI and Infrastructure
Senior Software Engineer, AI/ML, AI and Infrastructure

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Health insurance
401(k) with company match
Paid time off
+1
Software Engineer III, Infrastructure, Google Cloud AI
Software Engineer III, Infrastructure, Google Cloud AI

Google • Sunnyvale (CA)

On-site
USD 147,000 - 211,000
Health, dental, vision insurance
401(k) retirement plan with company match
20 days of vacation per year
+4