Senior Staff Software Engineer, Site Reliability Engineering, Workspace AI

engineeringjobs.net, Inc.

Sunnyvale (CA)

On-site

USD 200,000 - 280,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Bonus
Equity
Benefits

Job summary

engineeringjobs.net, Inc. seeks an experienced software architect to own the architecture and design of scalable AI inference and training infrastructure.

You will partner with development teams and tech leads to deliver cost-effective, high-performance systems for production use. You will lead incident response and blameless postmortems, promote AI-first development, and collaborate with SRE leaders across Google to implement scalable solutions.

Qualifications

  • Requires a bachelor's degree in computer science, a related field, or equivalent practical experience; at least eight years of software development experience; four years leading projects and providing technical leadership; and three years designing, analyzing, and troubleshooting distributed systems. Experience with machine learning or AI in a software development environment is required, while a master's degree in computer science or engineering is preferred.

Responsibilities

  • Own the architecture and design of reliable, scalable, cost-effective, and performant AI inference and training infrastructure, partnering with development teams and technical leads.
  • Lead production incident response and blameless postmortems, promote AI-first development, and collaborate with SRE leaders across Google to share scalable solutions.

Skills

Software Development
Technical Leadership
Distributed Systems
Machine Learning
Artificial Intelligence
Site Reliability Engineering
AI Infrastructure
Inference Infrastructure
Training Infrastructure
System Architecture
Large-Scale System Design
Incident Response
Blameless Postmortems
Automation
Scalability
Performance Optimization

Education

Bachelor's degree in computer science, related field, or equivalent practical experience
Master's degree in computer science or engineering preferred

Job description

Own the architecture and design of reliable, scalable, cost-effective, and performant AI inference and training infrastructure, partnering with development teams and technical leads. Lead production incident response and blameless postmortems, promote AI-first development, and collaborate with SRE leaders across Google to share scalable solutions.

Requirements

Requires a bachelor's degree in computer science, a related field, or equivalent practical experience; at least eight years of software development experience; four years leading projects and providing technical leadership; and three years designing, analyzing, and troubleshooting distributed systems. Experience with machine learning or AI in a software development environment is required, while a master's degree in computer science or engineering is preferred.

Key Skills

Software Development, Technical Leadership, Distributed Systems, Machine Learning, Artificial Intelligence, Site Reliability Engineering, AI Infrastructure, Inference Infrastructure, Training Infrastructure, System Architecture, Large-Scale System Design, Incident Response, Blameless Postmortems, Automation, Scalability, Performance Optimization

Benefits

Bonus, Equity, Benefits

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Staff Software Engineer, Site Reliability Engineering, Workspace AI
Senior Staff Software Engineer, Site Reliability Engineering, Workspace AI

Google LLC • Sunnyvale (CA)

On-site
USD 262,000 - 364,000
Software Engineering Manager II, Site Reliability Engineering, Data Intelligence
Software Engineering Manager II, Site Reliability Engineering, Data Intelligence

Google Inc. • San Jose (CA), Northern (KY)

On-site
USD 207,000 - 300,000
Equity
Benefits
Bonus target
Senior AI Infra Engineer - SRE for Inference (Equity)
Senior AI Infra Engineer - SRE for Inference (Equity)

engineeringjobs.net, Inc. • Sunnyvale (CA)

On-site
USD 200,000 - 280,000
Bonus
Equity
Benefits
Software Engineering Manager II, Site Reliability Engineering, AI Foundry SRE
Software Engineering Manager II, Site Reliability Engineering, AI Foundry SRE

Google • San Jose (CA)

On-site
USD 207,000 - 300,000
Equity
Bonus target
Benefits
Staff Software Engineer, Infrastructure, Google Cloud AI
Staff Software Engineer, Infrastructure, Google Cloud AI

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Software Engineering Manager II, Site Reliability Engineering, Data Intelligence
Software Engineering Manager II, Site Reliability Engineering, Data Intelligence

Google • San Jose (CA)

On-site
USD 207,000 - 300,000
Staff Software Engineer, Applied AI Engineering
Staff Software Engineer, Applied AI Engineering

Google • Town of Montana (WI)

On-site
USD 207,000 - 300,000
Staff Software Engineer - Lead Large-Scale AI Infra
Staff Software Engineer - Lead Large-Scale AI Infra

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Health insurance
Dental insurance
Vision insurance
+8
Senior Staff Software Engineer, AI Developer Tools, Cloud Gemini Enterprise
Senior Staff Software Engineer, AI Developer Tools, Cloud Gemini Enterprise

Google • Kirkland (WA)

On-site
USD 262,000 - 364,000
Health insurance
401(k) with company match
PTO 20 days/year
+4
Senior Staff Software Engineer, AI Developer Tools, Cloud Gemini Enterprise
Senior Staff Software Engineer, AI Developer Tools, Cloud Gemini Enterprise

Google • Seattle (WA)

On-site
USD 262,000 - 364,000
Health plan
401(k) with company match
Paid time off
+4