Platform Systems Engineer: Scalable AI Infra & Observability
OpenAI
Greater London
On-site
GBP 60,000 - 80,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A global AI research company in London is seeking a Software Engineer for Platform Systems to enhance large-scale AI training infrastructure. Key responsibilities include designing failure detection systems, improving observability, and collaborating with various teams to ensure the platform's reliability. Ideal candidates should have experience in distributed systems, performance optimization, and debugging complex issues. Join us in shaping the future of AI technology with state-of-the-art engineering solutions.
Qualifications
Experience with large-scale distributed systems and their performance.
Experience writing low-level software.
Understanding of hardware, operating systems, and networking.
Responsibilities
Design and build failure detection and tracing systems for AI training jobs.
Develop tools to identify issues in large-scale systems.
Improve observability and reliability of training infrastructure.
Skills
Performance analysis
Debugging
Systems engineering
Distributed systems
Observability
Job description
A global AI research company in London is seeking a Software Engineer for Platform Systems to enhance large-scale AI training infrastructure. Key responsibilities include designing failure detection systems, improving observability, and collaborating with various teams to ensure the platform's reliability. Ideal candidates should have experience in distributed systems, performance optimization, and debugging complex issues. Join us in shaping the future of AI technology with state-of-the-art engineering solutions.