Join us as we work together to inspire creativity and enrich life around the globe.
Location:
Team:
Employment Type:
Regular
Job Code:
A106244
Share this listing:
Responsibilities
- Responsible for the overall architecture design and implementation of model inference services, building a high-performance, highly available, and scalable enterprise-level inference system for large-parameter, high-complexity AI models, overcoming various architectural challenges in the implementation of complex model inference, and supporting the efficient launch of models across all business scenarios.
- Responsible for the R&D and optimization of the core modules of the inference framework, covering core capabilities such as inference engine scheduling, monitoring and alerting, canary release, etc., continuously iterating on the framework performance, and resolving performance bottlenecks, resource bottlenecks, and stability issues in high-concurrency and large-model inference scenarios.
- Keep track of the latest inference technologies in the industry, conduct technology selection and innovation in combination with business scenarios, accumulate distributed high-concurrency service architecture solutions, and promote the upgrade and standardization of the team's technical system.
Qualifications
Minimum Qualifications
- Individuals who are completing or have recently completed a Bachelor's/ Master's degree in computing or a related discipline.
- Familiar with basic Linux commands, with solid C/C++ programming skills and knowledge of data structures and algorithms.
- Familiar with the basic principles of multi-threaded concurrency, proficient in basic usages such as thread usage, synchronization locks, and thread pools, able to identify common concurrency issues, and possess the ability to perform basic performance tuning in multi-threaded scenarios.
- Have experience in R&D projects of high-concurrency distributed services, and be familiar with service latency and resource optimization.
- Possess good learning and execution abilities, be willing to proactively understand the model inference service architecture, and have problem analysis and abstraction capabilities.
- Possess good cross-team collaboration skills, communication and presentation skills, and document writing skills, have strong sense of responsibility and stress tolerance, and be able to drive the resolution of complex technical issues and the implementation of projects.
Preferred Qualifications
- Have practical project experience and understanding of source code in high-concurrency services/frameworks such as Redis, RocksDB, BRPC, GRPC, etc.
- Understand the operating mechanism of GPUs, and have relevant project experience and optimization capabilities in GPU service resource management and control.
Job Information
About Us
Why Join ByteDance
Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.
As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.
Diversity & Inclusion
ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives.
Our platform connects people from across the globe and so does our workplace.
At ByteDance, our mission is to inspire creativity and enrich life.
To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach.
We are passionate about this and hope you are too.