Principal Machine Learning Engineer – Production Systems

SoftInWay Inc

Bristol

On-site

GBP 70,000 - 100,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology company specializing in ML solutions is seeking a Senior/Principal ML Systems Architect to design and implement a scalable architecture for machine learning. The successful candidate will bridge research and commercial deployment, focusing on performance, reliability, and maintainability. Key qualifications include expertise in TensorFlow and Python, experience in API development, and familiarity with MLOps practices. Joining means working on cutting-edge ML solutions and collaborating with experts across various fields.

Qualifications

  • Expert in ML frameworks like TensorFlow and ONNX.
  • Advanced Python programming for machine learning.
  • Proficient in performance optimization techniques.

Responsibilities

  • Architect the ML Solver Platform and define modular architecture.
  • Convert research code into robust services.
  • Optimize performance and scalability across deployments.

Skills

Expert in TensorFlow (TF2/Keras)
Advanced Python for ML
Proficiency in gRPC/Protobuf and REST
GPU acceleration (CUDA/cuDNN)
Metrics, tracing, structured logging
SBOM, image signing, role-based access

Tools

TensorFlow
ONNX Runtime
FastAPI

Job description

Senior/Principal ML Systems Architect (TensorFlow + Python)
Overview

We are seeking a highly experienced ML Systems Architect to design and implement a scalable, production-grade architecture for our machine learning solver. This role bridges research prototypes and commercial deployment, ensuring reliability, maintainability, and performance in a mixed technology stack.

Responsibilities
  • Architect the ML Solver Platform:
    • Define modular architecture for data preprocessing, model execution, and post-processing.
    • Establish clear API contracts between Python/TensorFlow and C# services.
  • Convert research code into robust, testable, and observable services.
  • Implement CI/CD pipelines, automated testing, and reproducibility standards.
  • Design REST/gRPC endpoints for cross-language communication.
  • Ensure compatibility with C#/.NET services.
  • Performance & Scalability:
    • Optimize GPU/CPU utilization, batching strategies, and memory management.
    • Plan for multi-model and multi-tenant scenarios.
  • MLOps & Lifecycle Management:
    • Implement model versioning, artifact registries, and deployment workflows.
    • Set up monitoring, logging, and alerting for solver performance.
  • Security & Compliance:
    • Apply best practices for secrets management, dependency scanning, and secure artifact storage.
Required Skills & Experience
  • ML Frameworks: Expert in TensorFlow (TF2/Keras), experience with ONNX Runtime for inference.
  • Programming: Advanced Python for ML; strong understanding of packaging, type checking, and performance profiling.
  • APIs: Proficiency in gRPC/Protobuf and REST for cross-language integration.
  • Performance Optimization: GPU acceleration (CUDA/cuDNN), mixed precision, XLA, profiling.
  • Observability: Metrics, tracing, structured logging, dashboards.
  • Security: SBOM, image signing, role-based access, vulnerability scanning.
Preferred Qualifications
  • Experience with ONNX Runtime Training, PyTorch, or hybrid ML architectures.
  • Familiarity with distributed training strategies and multi-GPU setups.
  • Knowledge of feature stores and data validation frameworks.
  • Exposure to regulated environments and compliance frameworks.
Tools & Technologies
  • ML: TensorFlow, ONNX Runtime, tf2onnx.
  • APIs: FastAPI, gRPC.
Why Join Us?
  • Work on cutting-edge ML solutions integrated into commercial engineering software.
  • Define architecture that scales across global deployments.
  • Collaborate with a team of experts in ML, software engineering, and UI development.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer
Senior Machine Learning Engineer

BJAK • Greater London

On-site
GBP 80,000 - 120,000
Principal Machine Learning Engineer
Principal Machine Learning Engineer

NLP PEOPLE • Greater London

On-site
GBP 80,000 - 120,000
Senior ML Systems Architect: Scalable AI Platform
Senior ML Systems Architect: Scalable AI Platform

SoftInWay Inc • Bristol

On-site
GBP 70,000 - 100,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Loop Recruitment • Greater London

On-site
Lead ML Engineer - AI Solutions Architect, Hybrid London
Lead ML Engineer - AI Solutions Architect, Hybrid London

Loop Recruitment • Greater London

On-site
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Dimensionalconsultingltd • Manchester

Hybrid
GBP 70,000 - 90,000
Competitive salary with annual performance bonus
£2,000 annual conference and learning budget
Private healthcare and dental cover
+4
Machine Learning Specialist
Machine Learning Specialist

Stanford Black Limited • Greater London

On-site
GBP 100,000 - 150,000
Competitive compensation
Bonus structure
Autonomy from day one
+1
Lead Machine Learning Engineer
Lead Machine Learning Engineer

Xcede • Greater London

Hybrid
GBP 120,000 - 150,000
Senior ML Platform Engineer - Artificial Intelligence
Senior ML Platform Engineer - Artificial Intelligence

Bloomberg LP • Greater London

On-site
GBP 80,000 - 100,000
Machine Learning Engineer
Machine Learning Engineer

Generative Engineering • Greater London

On-site
GBP 70,000 - 100,000