Distributed Systems Engineer - AI Infra & GPU Clusters
krea.ai
San Francisco (CA)
On-site
USD 120,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A leading AI technology firm in San Francisco is seeking systems-oriented candidates to enhance their infrastructure used for advanced AI research. Ideal candidates will have a strong background in Python and experience managing distributed systems, particularly within GPU environments. The position offers the chance to work on cutting-edge projects involving large-scale data processing and custom infrastructure development. A passion for building and optimizing systems is a must.
Qualifications
Strong aptitude for building and optimizing distributed systems.
Good mental model of system interactions under varying conditions.
Experience with large-scale data processing and ETL systems.
Responsibilities
Work with distributed training and inference systems on GPU clusters.
Optimize data pipelines for research and operational tasks.
Collaborate with researchers to enhance ML infrastructure.
Skills
Python
Kubernetes
PyTorch
SQL
DuckDB
NumPy
Pandas
Tools
K8s
InfiniBand
Job description
A leading AI technology firm in San Francisco is seeking systems-oriented candidates to enhance their infrastructure used for advanced AI research. Ideal candidates will have a strong background in Python and experience managing distributed systems, particularly within GPU environments. The position offers the chance to work on cutting-edge projects involving large-scale data processing and custom infrastructure development. A passion for building and optimizing systems is a must.