Get more replies from employers
Send a job-specific resume in minutes.
ByteDance in Singapore is seeking a Site Reliability Engineer for their Machine Learning Systems team. You'll ensure operational efficiency, stability, and disaster recovery of ML systems while collaborating with a global team.
Ideal candidates possess a Bachelor's degree in Computer Science and proficiency in programming languages such as Go, Python, or Shell, alongside hands-on experience with Kubernetes. This role offers the opportunity to enhance your coding and performance analysis skills within a cutting-edge AI research environment.
Employment Type: Regular
Job Code: A247694
The ByteDance Large Model Team is committed to developing the most advanced AI large model technology in the industry, becoming a world‑class research team, and contributing to technological and social development. The Large Model Team has a long‑term vision and determination in the field of AI, with research directions covering NLP, CV, speech, and other areas. Relying on the abundant data and computing resources of the platform, the team has continued to invest in relevant fields and has launched its own general large model, providing multi‑modal capabilities. The Machine Learning (ML) System sub‑team combines system engineering and the art of machine learning to develop and maintain massively distributed ML training and inference system/services around the world, providing high‑performance, highly reliable, scalable systems for LLM/AIGC/AGI. In our team, you'll have the opportunity to build the large‑scale heterogeneous system integrating with GPU/NPU/RDMA/Storage and keep it running steadily and reliably, enrich your expertise in coding, performance analysis and distributed system, and be involved in the decision‑making process. You'll also be part of a global team with members from the United States, China and Singapore working collaboratively towards unified project direction.