As a Senior MTS, you will own systems, not just individual components. You will be responsible for architecture, technical trade-offs, operational health, and driving workstreams from ambiguous requirements to production with minimal scaffolding. You will work closely with production systems, customers, and the underlying infrastructure. Depending on business needs, you may contribute across engineering, infrastructure, AI, and customer-facing problems.
Technical Tracks
Track 1: AI Platforms
Best suited for backend and platform engineers with hands-on experience building production applications and services around LLMs and enterprise AI.
Responsibilities:
- Build RAG and multi-agent platforms
- Develop inference services
- Build evaluation and benchmarking systems
- Develop enterprise integrations
- Implement metering and billing systems
- Build governance and access-control capabilities
- Implement audit, privacy, PII handling, and guardrails
- Develop AI observability, tracing, and production monitoring
- Support customer VPC and on-premises deployments
Relevant Experience:
- Strong Python backend engineering
- REST/gRPC APIs and production services
- PostgreSQL, Redis, and schema design
- Kafka, NATS, or similar messaging systems
- Event-driven and distributed architectures
- LLM applications and production model APIs
- RAG pipelines
- Agentic applications and tool calling
- Vector databases
- LangChain, LangGraph, LlamaIndex, CrewAI, AutoGen, or equivalent frameworks
- LLM evaluation, observability, tracing, cost, and latency monitoring
- Enterprise AI integrations and platform engineering
Track 2: AI Infra
Best suited for systems, infrastructure, and platform engineers interested in building large-scale GPU infrastructure for AI workloads.
Responsibilities:
- Build GPU fleet management and control planes
- Develop schedulers and multi-tenant infrastructure
- Implement GPU sharing and isolation
- Build GPU health monitoring and recovery systems
- Develop Kubernetes controllers, exporters, daemons, and platform services
- Support on-premises and cloud GPU infrastructure
- Improve infrastructure reliability and observability at scale
Relevant Experience:
- Strong Go or Rust
- Linux systems and infrastructure engineering
- Kubernetes and container platforms
- Docker/containerd
- GPU infrastructure and fleet management
- CUDA, NVML, DCGM, MPS, or MIG
- Kubernetes controllers and operators
- REST and gRPC platform APIs
- Kafka, NATS, or RabbitMQ
- PostgreSQL or MySQL
- Prometheus, Grafana, or OpenTelemetry
- Distributed systems and fault tolerance
- Networking, resource isolation, and reliability engineering
- Experience with GPU infrastructure, bare-metal systems, Kubernetes platform engineering, or HPC environments is particularly valuable.
Track 3: Applied Research
Best suited for ML engineers who combine strong software engineering fundamentals with hands-on experimentation and a focus on research that reaches production.
Responsibilities:
- Build offline and online evaluation systems
- Develop LLM-as-a-judge and regression frameworks
- Improve agent reasoning, planning, and tool use
- Work on fine-tuning and post-training
- Improve retrieval quality
- Develop simulation environments for agent testing
- Build data curation and synthetic data pipelines
- Turn AI research into production-ready capabilities
Relevant Experience:
- Strong Python and ML engineering fundamentals
- PyTorch or JAX
- LLM evaluation and benchmarking
- Fine-tuning and post-training
- Agent evaluation
- RL/RLHF
- Embedding and reranker experiments
- Prompt optimisation and tool-use training
- Simulation environments
- Synthetic data generation
- Large-scale data curation
- Model training or serving infrastructure
- Statistical experimentation and rigorous evaluation
Requirements:
- Approximately 5 years of professional software engineering experience.
- Strong track record of building and operating production systems.
- Ability to own systems or significant workstreams end-to-end, from architecture and design through deployment, operations, debugging, and continuous improvement.
- Strong engineering fundamentals and practical problem-solving skills.
- Ability to make sound technical decisions in ambiguous environments.
Core Engineering Requirements:
- Strong programming skills in Python, Go, or Rust.
- Experience designing, building, and operating production services at scale.
- Strong understanding of APIs, databases, distributed systems, and software architecture.
- Hands-on experience with Docker and Kubernetes.
- Experience with at least one major cloud platform: AWS, Azure, or GCP.
- Strong understanding of Git, CI/CD, automated testing, and engineering best practices.
- Understanding of reliability concepts including fault tolerance, retries, idempotency, failure handling, logging, monitoring and observability.
- Ability to debug production issues, analyse system behaviour, and drive incidents to resolution.
- Ability to evaluate technical trade-offs and make maintainable architecture decisions.