Get more replies from employers
Send a job-specific resume in minutes.
United States Digital Space LLC is seeking a Senior Software Engineer to design and scale backend systems for AI development. You’ll work across distributed training infrastructure, workload orchestration, and platform APIs, partnering with multiple teams to deliver robust tooling.
This role involves building scalable services, ensuring reliability, and shaping the developer experience for researchers and enterprises. Office hubs include London, SF, and NYC with a hybrid work model.
the company is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.
Through our merger with Voltage Park, a neocloud and AI Factory, the company combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.
We serve solo researchers, startups, and large enterprises. the company operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.
The people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:
The Training \& Experimentation team at the company builds the platform that enables developers to experiment at scale and train their own intelligence. Whether customers are training foundation models, fine-tuning open-source models, launching distributed training jobs, or iterating on experiments, this team creates the infrastructure and developer experience that powers every stage of the AI lifecycle.
We're looking for a Senior Software Engineer to help design and scale the backend systems that make AI development faster, more reliable, and easier to use. You'll work across distributed training infrastructure, workload orchestration, experiment management, developer tooling, and platform APIs while partnering closely with our Managed Infrastructure, Core Platform, and Optimized Compute teams.
This is an opportunity to build systems that support real-world AI workloads at scale while shaping the developer experience used by researchers, startups, and enterprise AI teams around the world.
*This role is based in one of our San Francisco, NYC, or London office hubs, with a minimum of 2 in-office days per week and occasional team and company offsites. We are not able to provide visa sponsorship for this position at this time.*
We offer a comprehensive and competitive benefits package designed to support our employees' health, well-being, and long-term success:
Benefits may vary by location, team, and role.
*At the company, we are committed to fostering an inclusive and diverse workplace. We believe that diverse teams drive innovation and create better products. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential.*