An application made for this job — a tailored resume and cover letter that speak straight to the posting.
OpenTalent in San Francisco is seeking a systems-minded engineer to build and operate inference services for large language models at scale. You will own distributed pipelines and GPU memory optimization, collaborating with researchers and ML engineers to deliver reliable, cost-efficient systems.
The role emphasizes production reliability, incident ownership, and Bay Area in-person collaboration, with opportunities to shape the stack from data ingest to model serving.
OpenTalent in San Francisco is seeking a systems-minded engineer to build and operate inference services for large language models at scale. You will own distributed pipelines and GPU memory optimization, collaborating with researchers and ML engineers to deliver reliable, cost-efficient systems.
The role emphasizes production reliability, incident ownership, and Bay Area in-person collaboration, with opportunities to shape the stack from data ingest to model serving.