Our client is looking for an AI Engineer who can move fluidly between research and production: someone who can prototype a new model or prompting approach quickly, then harden it into a reliable, observable, cost-efficient service that runs at scale. You'll work closely with product and full-stack engineering to turn ambiguous problems into shipped AI features.
Responsibilities
- Design, build, and deploy machine learning and LLM-based systems, including retrieval-augmented generation (RAG) pipelines, fine-tuned models, and agentic workflows.
- Own the full lifecycle of AI features: data collection and evaluation, model/prompt selection, experimentation, deployment, and monitoring in production.
- Build and maintain data pipelines, embeddings stores, and vector databases that support search, retrieval, and personalization.
- Evaluate and integrate third-party model APIs (OpenAI, Anthropic, etc.) alongside open-source and self-hosted models, balancing quality, latency, and cost.
- Fine-tune, train, and deploy open-source models on our own infrastructure - including data prep, training/fine-tuning runs, quantization, and self-hosted inference serving at scale.
- Stay current with the fast-moving AI/ML landscape and bring back practical recommendations on tools, models, and techniques.
Qualifications & Experience
- 4+ years of experience building and shipping machine learning or AI-powered systems in production.
- Strong Python skills, with hands-on experience using frameworks such as PyTorch, TensorFlow, or similar.
- Practical experience with large language models - prompting, fine-tuning, RAG, or agent frameworks (e.g., LangChain, LlamaIndex, or custom implementations).
- Hands-on experience training and fine-tuning open-source models (e.g., Llama, Mistral, Qwen) - full fine-tuning or parameter-efficient methods (LoRA/QLoRA) - and deploying them on your own infrastructure rather than relying solely on hosted APIs.
- Experience with vector databases and embedding-based retrieval (e.g., Pinecone, Weaviate, pgvector, FAISS).
- Comfort working with cloud infrastructure (AWS, GCP, or Azure) and containerized deployments (Docker, Kubernetes), including GPU-backed compute.
- Proficiency working with AI coding assistants (e.g., Claude Code, GitHub Copilot, Cursor) to accelerate development, while critically reviewing and validating AI-generated code.
- Experience with AI coding assistants such as Cursor, Claude, or similar Experience with AI coding assistants such as Cursor, Claude, or similar
The Reference Number for this position is NG61533 which is a Permanent, Remote role offering a salary of up to R800k per annum salary negotiable based on experience.