Join us as we work together to inspire creativity and enrich life around the globe.
Location:
Team:
Employment Type:
Regular
Job Code:
A145010
Share this listing:
Responsibilities
- Design and build a unified platform/middleware system that can support diverse business requirements across different scenarios, including low cost, high availability, high throughput, high performance, and large-scale storage capacity.
- Design and optimize complex multi-tier storage architectures beyond GPU memory, CPU memory, and external storage, with a focus on efficient data placement and resource utilization.
- Keep up with the latest advances in software and hardware architectures, and proactively evaluate and experiment with emerging technologies.
- As an internal platform serving multiple teams, plan and optimize the utilization of large volumes of heterogeneous resources across multiple hardware generations, data centers, service tiers, and resource pools. Develop automated and dynamic optimization strategies based on changes in model size, service traffic, and workload characteristics.
Qualifications
Minimum Qualification(s)
- Bachelor's degree or above in Computer Science, Software Engineering, or a related field
- Proficient in C++ and Python programming in Linux environments.
- Strong understanding of distributed systems principles, with hands-on experience in the design, development, maintenance, and continuous optimization of large-scale distributed systems. Able to identify potential issues and bottlenecks in complex distributed systems.
- Experience working on distributed systems in areas such as recommendation, search, or machine learning, with exposure to resource scheduling, task orchestration, model training, model inference, feature extraction, ML Systems (MLSys), or AIOps.
- Strong logical and analytical thinking skills, with the ability to abstract and decompose complex business and technical requirements effectively. Strong teamwork and collaboration skills.
Preferred Qualification(s)
- Experience optimizing systems similar to Parameter Server, or optimizing indexing structures in large-scale search systems.
- Experience with KV Cache systems, such as Mooncake, including system optimization and performance tuning and with mainstream machine learning frameworks such as TensorFlow, PyTorch, or MXNet.
- Familiarity with open-source storage systems such as Redis, LevelDB/RocksDB, and MongoDB, or hands-on experience using or optimizing large-scale distributed storage systems such as HDFS or Ceph.
Job Information
About Us
Why Join ByteDance
Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect - and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.
As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.
Diversity & Inclusion
ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.