Senior Platform Engineer, Metal Dev — AI Infra Reliability

Weights & Biases

United States

Hybrid

USD 153,000 - 242,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance - 100% paid for
Company-paid Life Insurance
Voluntary supplemental life insurance
Short and long-term disability insurance
Flexible Spending Account
Health Savings Account
Tuition Reimbursement
Employee Stock Purchase Program (ESPP)
Mental Wellness Benefits
Paid Parental Leave
Flexible, full-service childcare support
401(k) with a generous employer match
Flexible PTO
Catered lunch each day
Casual work environment
Innovative disruption culture

Job summary

A leading technology company is seeking a Senior Platform Engineer to join its Hardware Engineering Dev team. The role involves managing incident responses, improving system observability, and enhancing operational reliability. Candidates should have 7+ years of experience in cloud operations, proficiency in Go, and familiarity with tools like Prometheus and Grafana. This position offers a competitive salary and benefits package, including health insurance, a 401(k) plan, and flexible PTO.

Qualifications

  • 7+ years of experience in cloud operations, site reliability engineering (SRE), or related technical roles.
  • Understanding of cloud platforms (e.g., Kubernetes, AWS, GCP) and basic knowledge of cloud infrastructure.
  • Familiarity with incident management practices and frameworks (e.g., ITIL, SRE best practices).
  • Proficiency with Go.
  • Prior experience with Prometheus / Grafana.
  • Previous experience deploying containerized applications using Kubernetes.
  • Excellent documentation skills and attention to detail.
  • Strong analytical and problem-solving abilities.
  • Served on an on-call rotation supporting production services.

Responsibilities

  • Lead incident response efforts by identifying and resolving service disruptions quickly.
  • Lead the documentation of incidents, conduct in-depth root cause analysis (RCA).
  • Own the development and continuous improvement of incident response playbooks.
  • Clearly communicate efforts during incidents to the management and stakeholders.
  • Master clear understanding of various services on how they work in production.
  • Build a strategy around making core services perform at scale.
  • Own system observability and health leveraging tools like Prometheus and Grafana.
  • Lead automation efforts to streamline incident detection and recovery.
  • Define and drive KPIs and SLAs for incident management.
  • Collaborate with engineers across teams to improve platform reliability.
  • Design and implement solutions to build operational efficiency and stability.
  • Document hardware automation workflows and processes.
  • Create CI/CD pipelines.
  • Ensure smooth operation of server hardware lifecycle.
  • Partner with the Fleet Operations Team to design scalable tooling.
  • Build out dashboards and alerts for operational troubleshooting.
  • Participate in on-call rotation.

Job description

A leading technology company is seeking a Senior Platform Engineer to join its Hardware Engineering Dev team. The role involves managing incident responses, improving system observability, and enhancing operational reliability. Candidates should have 7+ years of experience in cloud operations, proficiency in Go, and familiarity with tools like Prometheus and Grafana. This position offers a competitive salary and benefits package, including health insurance, a 401(k) plan, and flexible PTO.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Engineer, Metal Dev — AI Infra Reliability
Senior Platform Engineer, Metal Dev — AI Infra Reliability

CoreWeave • New York (NY)

Hybrid
USD 153,000 - 242,000
Senior Platform Infrastructure Engineer – Scale & Reliability
Senior Platform Infrastructure Engineer – Scale & Reliability

Verkada • San Mateo (CA)

On-site
USD 130,000 - 280,000
Healthcare programs
Mental health support
Paid parental leave
+2
Senior Platform Engineer — Remote Backend & Infra Lead
Senior Platform Engineer — Remote Backend & Infra Lead

Saga.xyz • Los Altos (CA)

Remote
USD 180,000 - 240,000
Remote work
Flexible working hours
Flexible vacation policy
+3
Senior Platform Engineer: AI-Driven Infra & Security
Senior Platform Engineer: AI-Driven Infra & Security

7AI, Inc. • Boston (MA)

On-site
USD 120,000 - 150,000
Senior Platform Reliability Engineer (Kubernetes)
Senior Platform Reliability Engineer (Kubernetes)

Saviynt • Atlanta (GA)

On-site
USD 120,000 - 150,000
Competitive compensation
Benefits package
Career growth opportunities
Senior DevOps Engineer—AI Platform Reliability
Senior DevOps Engineer—AI Platform Reliability

Flux Enterprise • San Francisco (CA)

On-site
USD 120,000 - 150,000
Senior Platform Engineer — Scale & Observability
Senior Platform Engineer — Scale & Observability

CloudDevs • United States

Remote
USD 140,000 - 210,000
Competitive compensation package
WeWork membership
Regular team offsites
+2
Staff Platform Engineer — AI-Driven, Scalable Infra Lead
Staff Platform Engineer — AI-Driven, Scalable Infra Lead

Beam Benefits • United States

Remote
USD 130,000 - 180,000
Senior Platform Security Engineer — AI Infra, Equity
Senior Platform Security Engineer — AI Infra, Equity

Pallet • New York (NY), San Francisco (CA)

On-site
USD 195,000 - 265,000
Health, Vision, and Dental benefits
Life Insurance and Accidental Insurance
Daily catered lunches
+1
Senior Platform Engineer
Senior Platform Engineer

7AI • Boston (MA)

On-site
USD 120,000 - 160,000