A pioneering AI tech company in Poland is seeking a Machine Learning Engineer to enhance model inference performance at scale. You will optimize latency and cost for large-scale ML models, collaborate with research engineers to deploy architectures, and profile inference pipelines to ensure reliability. The ideal candidate has strong experience in ML optimization, a solid grasp of deep learning concepts, and is comfortable in fast-paced startup settings. Competitive compensation and equity options are offered.
Qualifications
Strong experience in ML inference optimization or high-performance ML systems.
Solid understanding of deep learning internals like attention and memory layout.
Hands-on experience with PyTorch (or similar) and model deployment.
Familiarity with GPU performance tuning techniques.
Responsibilities
Optimize inference latency and cost for large-scale ML models in production.
Profile GPU/CPU inference pipelines for memory and process bottlenecks.
Collaborate with research engineers to productionize new model architectures.
Benchmark performance across hardware like NVIDIA and AMD GPUs.
Skills
ML inference optimization
Deep learning internals
PyTorch or similar
GPU performance tuning
Scaling inference for real users
Startup environments
Tools
CUDA
Triton
TensorRT
Job description
A pioneering AI tech company in Poland is seeking a Machine Learning Engineer to enhance model inference performance at scale. You will optimize latency and cost for large-scale ML models, collaborate with research engineers to deploy architectures, and profile inference pipelines to ensure reliability. The ideal candidate has strong experience in ML optimization, a solid grasp of deep learning concepts, and is comfortable in fast-paced startup settings. Competitive compensation and equity options are offered.