Enable job alerts via email!

Site Reliability Engineer - ARK Large Model Platform (Singapore)

Byteplus

Singapore

On-site

SGD 80,000 - 120,000

Full time

Today

Be an early applicant

Generate a tailored resume in minutes

Land an interview and earn more. Learn more

Job summary

A technology company is seeking a qualified candidate for the Applied Machine Learning Enterprise team to develop and enhance the Ark Large Model Platform on Volcano Engine. This role requires extensive experience in cloud computing and large-scale model systems, as well as proficiency in programming languages such as Golang, Python, or Java. The successful candidate will manage large model systems, ensuring reliability and performance while striving to reduce IT costs. Opportunities for career growth are available.

Benefits

Positive team atmosphere

Career growth opportunity

Flat organization

Qualifications

Minimum of 5 years of R&D experience in cloud computing or large-scale model systems.
Ability to use programming languages proficiently in a professional setting.
Understanding of relevant technology stack.

Responsibilities

Develop and oversee the Ark Large Model Platform on Volcano Engine.
Manage stability of large-scale model systems through DevOps practices.
Ensure efficient operation and maintenance of large model systems.

Skills

Proficiency in cloud-native technologies

Expertise in Golang

Expertise in Python

Expertise in Java

Familiarity with log collection and monitoring

Education

B. Sc or higher degree in Computer Science or related fields

Tools

Terraform

About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

Diversity & Inclusion

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Job highlights

Positive team atmosphere, Career growth opportunity, Flat organization

Responsibilities

The Applied Machine Learning (AML) - Enterprise team provides machine learning platform products on VolcanoEngine with cloud native resource scheduling system which intelligently orchestrates different tasks and jobs with minimised costs of every experiment and maximised resource utilisation, rich modelling tools including customised machine learning tasks and web IDE, and multi-framework high performance model inference services.

In 2021, through VolcanoEngine, we released this machine learning infrastructure to the public, to provide more enterprises with reduced costs of computation power, lower barriers to machine learning engineering and deeper developments in AI capabilities.

Responsibilities

Responsible for Ark Large Model Platform development on Volcano Engine, researching systematic solutions on large model solution implementations and applications in various industries, striving to reduce the IT cost of large model applications, meeting the users' ever-growing demand for intelligent interaction and improving the lifestyle and communications of users in the future world.

Manage and oversee the stability of both control and data aspects of large-scale model systems through effective DevOps practices.
Develop and enhance observability systems for monitoring the stability of large model systems, ensuring high reliability and performance.
Handle super large-scale cluster management and ensure efficient operation and maintenance of large model systems.

Qualifications

Minimum Qualifications

B. Sc or higher degree in Computer Science or related fields from accredited and reputable institutions.
Minimum of 5 years of R&D experience in the fields of cloud computing or large-scale model systems.
Proficiency in cloud-native technologies and understanding of the relevant technology stack.
Expertise in one of the following programming languages: Golang, Python, or Java, with the ability to use it proficiently in a professional setting.
Familiarity with cloud-native technologies for log collection, monitoring, and alerting.

Preferred Qualifications:

Prior experience in the construction and maintenance of stability systems for large-scale infrastructures.
Experience in operating and maintaining large-scale systems.
Experience with infrastructure as code, particularly Terraform, is highly desirable.

Get your free, confidential resume review.

or drag and drop a PDF, DOC, DOCX, ODT, or PAGES file up to 5MB.

Top companies

Popular jobs