Production Support Engineer — LLM Platform

Qube Research & Technologies

Hong Kong

On-site

HKD 500,000 - 750,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Qube Research & Technologies is hiring for a production support role focused on the LLM gateway platform. You will deliver first- and second-line support, triage incidents, and coordinate with engineering and cloud teams to maintain availability and performance across a fast-growing environment.

Responsibilities include incident resolution, tooling and automation, dashboards and runbooks, with an emphasis on observability and cross-team collaboration to minimize disruption and maintain service

Qualifications

  • Experience in a production support, SRE or platform operations role within a fast-paced environment, with ownership of issue resolution end to end.
  • Strong Linux and Windows system administration skills.
  • Proficiency scripting in Python, Bash and/or PowerShell.
  • Solid experience with relational databases and writing queries for investigation.
  • Practical understanding of monitoring and observability (metrics, logs, traces, dashboards).
  • Debugging distributed, API-driven services (HTTP, timeouts, retries, latency).
  • Familiarity with LLM concepts, hosting environments, and API gateways.
  • Exposure to public cloud (AWS Bedrock, networking, IAM) and containerised workloads.

Responsibilities

  • Provide first- and second-line support for the LLM gateway platform, investigating and resolving issues raised by engineering and business users across the firm.
  • Monitor and maintain the platform’s underlying infrastructure to ensure availability, stability and predictable performance under rapidly growing load.
  • Support and troubleshoot model-serving backends and provider integrations, covering latency, throughput, error-rate and capacity issues.
  • Triage incidents affecting model availability and drive to resolution with vendors where required.
  • Coordinate with platform engineering, cloud infrastructure and end-user teams to resolve incidents and minimise disruption.
  • Support release management and change processes to keep production stable, including staged rollouts and rollbacks.
  • Build tooling and automation for monitoring, diagnostics and operational visibility.
  • Contribute to the design and implementation of dashboards, alerting and status communication.
  • Own and improve operational documentation and runbooks.

Skills

Ownership
Linux
Windows
Python
Bash/PowerShell
SQL
Observability
Cloud/AWS
Kubernetes
Communication
Incident response

Tools

Grafana
Prometheus
Terraform
Ansible
CI/CD tooling

Job description

Qube Research & Technologies (QRT) is a global quantitative and systematic investment manager, operating in all liquid asset classes across the world. We are a technology and data driven group implementing a scientific approach to investing. Combining data, research, technology, and trading expertise has shaped our collaborative mindset, which enables us to solve the most complex challenges. QRT’s culture of innovation continuously drives our ambition to deliver high quality returns for our investors.

Your future role within QRT
  • Provide first- and second-line support for LLM gateway platform, investigating and resolving issues raised by engineering and business users across the firm
  • Monitor and maintain the platform’s underlying infrastructure to ensure availability, stability and predictable performance under rapidly growing load
  • Support and troubleshoot model-serving backends and provider integrations, model providers, covering latency, throughput, error-rate and capacity issues
  • Triage incidents affecting model availability — provider instability, connection resets, timeouts, regional slowness — determine whether the cause is platform-side or upstream, and drive to resolution with vendors where required
  • Support the tooling layer built on top of LLM gateway: integrations, developer workspaces (e.g. Coder), coding assistants and API clients, including diagnosing issues introduced by upstream vendor releases running against a gateway-fronted API
  • Coordinate with platform engineering, cloud infrastructure and end-user teams to resolve incidents and minimise disruption
  • Support release management and change processes to keep production stable, including staged rollouts, non-prod validation and rollback
  • Build tooling and automation to improve monitoring, diagnostics and operational visibility, and to reduce repetitive manual work
  • Contribute to the design and implementation of monitoring, dashboards and alerting — for example extending Grafana dashboards covering TTFT, TPOT, percentile latency and failure-rate reporting
  • Own and improve operational documentation, runbooks and user-facing status communication
Your present skillset
  • Experience in a production support, SRE or platform operations role within a fast-paced environment, with strong ownership of issue resolution end to end
  • Strong Linux and Windows system administration skills
  • Proficiency scripting and automating in Python, Bash and/or PowerShell
  • Solid experience with relational databases such as PostgreSQL or SQL Server, including writing queries for investigation and supporting routine operational processes
  • Practical understanding of monitoring and observability: metrics, logs, traces, dashboards and alerting, and the ability to analyse system data to distinguish a platform-wide problem from a localised one
  • Comfortable debugging distributed, API-driven services: HTTP status and error semantics, timeouts, retries, connection resets, rate limiting, caching and latency percentiles
  • Familiarity with large language model concepts and hosting environments — inference APIs, model gateways/proxies, prompt and context handling, token accounting, streaming responses, prompt caching
  • Exposure to public cloud, ideally AWS (Bedrock, networking, IAM, logging/metrics), and to containerised or Kubernetes-based workloads
  • Ability to communicate clearly with both engineers and non-technical users, and to manage expectations of senior stakeholders during live incidents
  • Awareness of data-sensitivity and access-control considerations when routing workloads to third-party model providers
Beneficial
  • Experience supporting developer tooling and AI coding assistants (e.g. Claude Code, OpenCode) or IDE/workspace platforms
  • Experience with Grafana, Prometheus or equivalent observability stacks, including building dashboards and alert rules
  • Experience with CI/CD and infrastructure-as-code (Terraform, Ansible, or similar)
  • Experience operating multi-region services and troubleshooting region-specific performance issues (e.g. APAC latency)
  • Experience acting as the operational interface to third-party vendors and cloud providers during degradations

QRT is an equal opportunity employer. We welcome diversity as essential to our success. QRT empowers employees to work openly and respectfully to achieve collective success. In addition to professional achievement, we are offering initiatives and programs to enable employees achieve a healthy work-life balance.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production Support Engineer - LLM Platform
Production Support Engineer - LLM Platform

Qube Research & Technologies • Hong Kong

On-site
HKD 600,000 - 900,000
Network Operations Engineer
Network Operations Engineer

Quberesearchandtechnologies • Hong Kong

On-site
HKD 70,000 - 90,000
Healthy work-life balance initiatives
Professional development programs
Diversity and inclusion initiatives
Cloud DevOps Engineer
Cloud DevOps Engineer

Quberesearchandtechnologies • Hong Kong

On-site
HKD 500,000 - 700,000
Data Support Engineer
Data Support Engineer

Quberesearchandtechnologies • Hong Kong

On-site
HKD 300,000 - 450,000
Quantitative Strategist
Quantitative Strategist

Quberesearchandtechnologies • Hong Kong

On-site
HKD 300,000 - 500,000
Mentorship from industry professionals
Healthy work-life balance initiatives
Senior Software Engineer - Quant Research Framework
Senior Software Engineer - Quant Research Framework

Qube Research & Technologies • Hong Kong

On-site
HKD 650,000 - 1,100,000
Production Support Engineer
Production Support Engineer

Tribus • Hong Kong

On-site
HKD 420,000 - 660,000
Software Engineer - Data & Lifecycle (Equities/CA)
Software Engineer - Data & Lifecycle (Equities/CA)

Quberesearchandtechnologies • Hong Kong

On-site
HKD 600,000 - 800,000
Market Microstructure Researcher
Market Microstructure Researcher

Quberesearchandtechnologies • Hong Kong

On-site
HKD 500,000 - 800,000
Cloud Engineer
Cloud Engineer

Quberesearchandtechnologies • Hong Kong

On-site
HKD 480,000 - 720,000