ML Infrastructure Engineer: GPU Training & Serving

Character.AI

San Francisco (CA)

On-site

USD 130,000 - 207,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Flexible working hours
Professional development opportunities

Job summary

Join a dynamic team as a Machine Learning Infrastructure Engineer at Character.AI, where you'll enhance infrastructure for machine learning endeavors. This role requires substantial experience and expertise in cloud platforms, GPU management, and diagnostic tooling. Contribute to pioneering AI technology in an award-winning company recognized for its innovative approach to interactive entertainment.

Qualifications

  • 4+ years in ML infrastructure roles.
  • Experience with diagnosing problems in ML environments.
  • Familiar with GPU usage and high-performance computing.

Responsibilities

  • Support infrastructure for ML research and product.
  • Build tools for diagnosing cluster issues.
  • Monitor deployments and manage experiments.

Skills

Cloud platforms
GPU management
Diagnostic tools

Education

Bachelor’s degree in Computer Science or related field

Tools

Kubernetes
Compute Engine
Cloud Storage

Job description

Join a dynamic team as a Machine Learning Infrastructure Engineer at Character.AI, where you'll enhance infrastructure for machine learning endeavors. This role requires substantial experience and expertise in cloud platforms, GPU management, and diagnostic tooling. Contribute to pioneering AI technology in an award-winning company recognized for its innovative approach to interactive entertainment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Engineer: Scale GPU Training & Inference
ML Infra Engineer: Scale GPU Training & Inference

Reducto • San Francisco (CA)

On-site
USD 120,000 - 160,000
Unlimited PTO
Free lunch
Reimbursed transportation
+3
AI Infrastructure Engineer: Scalable GPU Inference, On-Site
AI Infrastructure Engineer: Scalable GPU Inference, On-Site

Spellbrush • San Francisco (CA)

On-site
USD 90,000 - 150,000
Equity
Health Insurance
Dental Insurance
+1
Principal ML Infra Engineer - GPU Inference & C++ Systems
Principal ML Infra Engineer - GPU Inference & C++ Systems

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
Senior GPU ML Infra Engineer — Mid-Training & Inference
Senior GPU ML Infra Engineer — Mid-Training & Inference

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Infrastructure Engineer — GPU, Kubernetes & Automation
AI Infrastructure Engineer — GPU, Kubernetes & Automation

MARS-TECHNOMINDS, INC • Town of Florida (NY)

On-site
USD 100,000 - 130,000
Senior ML Infrastructure Engineer - GPU & Scale
Senior ML Infrastructure Engineer - GPU & Scale

TensorWave • Las Vegas (NV)

On-site
USD 120,000 - 150,000
Competitive Salary
Stock Options
100% paid Medical, Dental, and Vision insurance
+8
ML Training Platform Architect for Large-Scale GPU Clusters
ML Training Platform Architect for Large-Scale GPU Clusters

Scale AI • Seattle (WA), New York (NY), San Francisco (CA)

On-site
USD 216,000 - 270,000
Comprehensive health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
Architect AI & ML Infrastructure for Scalable Systems
Architect AI & ML Infrastructure for Scalable Systems

Vultr • United States

Remote
USD 145,000 - 160,000
100% company-paid insurance premiums
401(k) plan with matching
Professional Development Reimbursement
+4
AI/ML Engineer: Next‑Gen Platforms & GPU Workloads
AI/ML Engineer: Next‑Gen Platforms & GPU Workloads

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Staff ML Infra Engineer: Distributed Training & Inference
Staff ML Infra Engineer: Distributed Training & Inference

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000