About The Role
We’re looking for a mid-to-senior MLOps engineer to help build the intelligence layer of our AI platform. This is not a traditional MLOps role focused only on maintaining pipelines and deployments. You will help create the underlying infrastructure that makes sophisticated AI workflows accessible through a simple, self-service product experience: model deployment, training, notebooks, functions, UI-driven agent creation, and RAG pipelines.
You’ll work at the intersection of platform engineering, machine learning infrastructure, distributed systems, and developer experience.
What You’ll Do
- Design and build the infrastructure powering our Intelligence layer.
- Enable reliable, automated workflows for model training, deployment, lifecycle management, and inference.
- Build scalable foundations for users to create, configure, and operate AI agents and RAG pipelines through the platform UI.
- Develop the platform capabilities behind managed notebooks, functions, experiments, training jobs, model registries, and serving endpoints.
- Improve model serving, observability, versioning, evaluation, promotion, and rollback capabilities.
- Optimize GPU inference and training deployments for performance, reliability, and cost efficiency.
- Explore efficient approaches for deploying models across centralized GPU infrastructure and edge devices.
- Automate workflows that otherwise require manual MLOps effort, with a focus on safe, self-service capabilities for platform users.
- Partner with backend, product, and AI teams to turn complex infrastructure into intuitive platform features.
- Help define standards for security, multi-tenancy, resource isolation, model governance, and operational reliability.
Our current stack
You’ll Work With Technologies Including
- Kubernetes and Knative
- KServe and vLLM
- MLflow
- LangFuse
- GPU infrastructure and model-serving workloads
- RAG architectures, vector retrieval, agent workflows, and LLM applications
Our intelligence backend is built in Rust, so Rust experience is a strong advantage.
What We’re Looking For
- Strong experience in MLOps, AI platform engineering, machine learning infrastructure, or distributed systems.
- Hands-on Kubernetes experience, including deploying and operating stateful, training, notebook, serverless, or GPU-intensive workloads.
- Experience building ML platforms that support some combination of training, experimentation, notebooks, feature/data workflows, model registries, and production serving.
- Experience with model serving frameworks such as KServe, vLLM, Triton, Ray Serve, or similar.
- Familiarity with ML lifecycle tooling such as MLflow, experiment tracking, model registries, and CI/CD or GitOps for ML systems.
- Practical experience optimizing training and inference workloads for latency, throughput, availability, and cost.
- Understanding of GPU scheduling, resource allocation, model quantization, batching, autoscaling, and multi-model serving.
- Experience with RAG systems, LLM applications, AI agents, or their supporting infrastructure.
- Strong software engineering fundamentals and a product-minded approach to platform design.
- Ability to make complex operational workflows reliable and approachable for users.
Nice to have
- Rust experience or interest in working with Rust-based backend systems.
- Experience running notebook environments such as JupyterHub, Kubeflow Notebooks, VS Code Server, or equivalent.
- Experience creating secure, scalable execution environments for user-defined functions or jobs.
- Experience deploying AI workloads on edge devices.
- Familiarity with model optimization techniques, including quantization, compilation, pruning, and hardware-aware serving.
- Experience with multi-tenant AI platforms, security controls, and model governance.
- Familiarity with observability and evaluation tooling for ML and LLM systems.
What Success Looks Like
You will help us evolve from manually operated AI infrastructure to a platform where users can train, evaluate, deploy, and operate models; work in managed notebooks; run functions; create agents; and configure RAG workflows - with the platform handling as much of the operational complexity as possible.