Staff Engineer (Core & MLOps)

Jobgether SRL

United States

Remote

USD 140,000 - 190,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Fully remote
Flexible working hours
Global collaboration
Conference access

Job summary

Jobgether SRL seeks a Staff Engineer (Core & MLOps) to shape foundational infrastructure powering large-scale web data products and distributed teams. You will own architecture of core control and context planes enabling AI-driven workflows to operate reliably across multi-cloud environments.

The role demands hands-on architecture, mentoring across squads, and strong collaboration in a remote-first, globally distributed setting.

Qualifications

  • 10+ years of experience building scalable distributed backend systems.
  • Advanced Java with reactive frameworks and strong Python proficiency.
  • Deep experience with gRPC and Protocol Buffers and schema evolution.
  • Hands-on production experience with Kubernetes, Terraform, and Kafka.
  • Experience designing telemetry pipelines and reliable service contracts.
  • Strong reliability engineering, SLO/SLI and incident post-mortems.
  • Excellent technical writing and cross-team alignment skills.
  • Experience with multi-cloud infrastructure and dev tooling is a plus.

Responsibilities

  • Architect and evolve core control and context planes for multi-language clients.
  • Own service chassis, pipelines, and deployment tooling (Helm, CI/CD).
  • Define inter-service contracts, API gateway, and schema evolution standards.
  • Operate core platform infra across Kubernetes, Terraform, and Kafka stacks.
  • Lead architectural strategy and write RFDs for platform initiatives.
  • Establish SRE practices including SLOs, error budgets, and canary deployments.
  • Mentor engineers across squads and drive reliable software practices.
  • Lead incident post-mortems and translate insights into platform improvements.

Skills

Java
Python
gRPC
Protocol Buffers
Distributed systems design
Technical writing
Mentoring

Tools

Kubernetes
Terraform
Kafka
Confluent Kafka
Envoy
SPIRE
Helm

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Engineer (Core & MLOps) based in United Arab Emirates.

This role offers the opportunity to shape foundational infrastructure powering large-scale web data products and distributed engineering teams. You'll own the architecture of core control and context planes that enable services and AI-driven workflows to operate reliably and efficiently. Working across Kubernetes, Kafka, Java, Python, gRPC, and multi-cloud infrastructure, you'll tackle complex distributed systems challenges at production scale. You'll establish engineering standards, reliability practices, and service contracts that influence multiple product squads. The role combines hands-on architecture with technical leadership, mentoring, and cross-functional alignment. In a globally distributed, remote-first environment, you'll have significant autonomy to solve challenging infrastructure problems and influence long-term platform strategy.

Accountabilities
  • Architect and evolve the control and context planes, advancing service and schema registries, SLO enforcement, health-aware routing, automated canary releases, and operational feedback loops.
  • Own the service chassis and golden path, maintaining and improving multi-language Java and Python client libraries, standardized workload specifications, Helm charts, and deployment pipelines.
  • Define and govern inter-service contracts, including gRPC and Protocol Buffer definitions, API gateway transcoding, versioning policies, and schema evolution standards.
  • Operate and improve the core platform infrastructure across Kubernetes, Terraform, HAProxy/Nginx, Confluent Kafka, real-time billing pipelines, Valkey, and database modernization initiatives.
  • Lead architectural strategy through Requests for Discussion (RFDs) covering workflow orchestration, gateway orchestration, multi-cluster routing, automated failover, and other critical platform initiatives.
  • Establish reliability engineering practices, including SLOs, SLIs, error budgets, fault isolation, and automated weighted canary deployments.
  • Participate in shared infrastructure on-call rotations, lead incident post-mortems, and convert operational insights into platform improvements.
  • Mentor engineers across multiple squads, review architectural proposals, and establish engineering practices that make reliable software development more consistent and efficient.
Requirements
  • 10+ years of experience building scalable distributed backend systems, with a strong track record of creating internal platforms or core libraries adopted across engineering organizations.
  • Advanced Java expertise, including reactive frameworks such as Vert.x or Netty, combined with strong Python proficiency.
  • Deep experience with gRPC and Protocol Buffers, including schema evolution and backward compatibility in mission-critical systems.
  • Hands-on production experience with Kubernetes at scale, Terraform, and event-streaming platforms such as Kafka.
  • Experience designing automated telemetry pipelines, materialized views, feature stores, or other feedback systems that use production data to dynamically improve system behavior.
  • Strong reliability engineering background, including SLO/SLI definition, blast-radius analysis, fault tolerance, and rigorous service contracts.
  • Exceptional technical writing skills and the ability to communicate complex architectural concepts clearly while driving alignment across teams.
  • Strong written and interpersonal communication skills suited to a globally distributed, remote-first environment.
  • A curious, continuous-learning mindset with an interest in evaluating new technologies, architectures, and engineering approaches.
  • Experience with Temporal, DBOS, or similar durable execution platforms is a plus.
  • MLOps experience, including model serving, performance monitoring, or production drift detection, is advantageous.
  • Familiarity with zero-trust networking and service meshes such as SPIRE, mTLS, Cilium, Istio, or Envoy is beneficial.
  • Experience building developer tooling such as CLIs, SDKs, or project generators is a plus.
  • Experience with large-scale web scraping or crawling, or contributions to distributed-systems and data-extraction open-source projects, is advantageous.
Benefits
  • Fully remote, remote-first working environment with flexible working hours.
  • Freedom and flexibility to work from the location where you are most productive.
  • Opportunity to work on core infrastructure supporting large-scale web data pipelines and distributed systems.
  • Exposure to cutting-edge open-source technologies, tools, and evolving AI and web data infrastructure.
  • Opportunities to attend conferences and connect with colleagues across the globe.
  • Collaboration with a diverse, multicultural, and globally distributed engineering community.
  • High level of autonomy and organizational trust.
  • Opportunities to influence platform architecture, engineering standards, and technical strategy across multiple teams.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Engineer, Core & MLOps — Remote (UAE)
Staff Engineer, Core & MLOps — Remote (UAE)

Jobgether SRL • United States

Remote
USD 140,000 - 190,000
Fully remote
Flexible working hours
Global collaboration
+1
Director of Technology & Chief Architect
Director of Technology & Chief Architect

Jobgether SRL • United States

Remote
USD 180,000 - 280,000
Fully remote in United States
Executive tech leadership
Paid time off
Senior Backend Engineer: Machine Learning Infrastructure
Senior Backend Engineer: Machine Learning Infrastructure

Lever, Inc. • Town of Italy (NY)

On-site
EUR 71,000 - 106,000
Fully remote work environment
Unlimited vacation
Home-office stipend
+3
Engineering Manager - Foundations & Enablement
Engineering Manager - Foundations & Enablement

Lever, Inc. • Germany (OH)

Remote
USD 125,000 - 159,000
Remote stipend
Home office package
Equipment upgrades
+8
Senior/Staff Platform Engineer
Senior/Staff Platform Engineer

Jobgether SRL • United States

Remote
USD 150,000 - 210,000
Fully remote Canada
Technical ownership
Remote‑first collaboration
+1
Cloud Engineer - REMOTE - 80/HR
Cloud Engineer - REMOTE - 80/HR

ContractStaffingRecruiters.com • Branford (CT)

Remote
USD 130,000 - 160,000
Platform Engineer
Platform Engineer

The HT Group • Round Rock (TX)

Hybrid
USD 130,000 - 190,000
Platform Cloud Engineer - 1 year contract - 80/hr - REMOTE
Platform Cloud Engineer - 1 year contract - 80/hr - REMOTE

ContractStaffingRecruiters.com • Stamford (CT)

Remote
USD 120,000 - 150,000
Platform Cloud Engineer - 1 year contract - 80/hr - REMOTE
Platform Cloud Engineer - 1 year contract - 80/hr - REMOTE

ContractStaffingRecruiters.com • Branford (CT)

Remote
USD 120,000 - 160,000
Platform Engineer
Platform Engineer

Jobgether • Kentucky

On-site
USD 110,000 - 140,000
Remote-First Flexibility
Access to premium hardware and software
Generous time off
+3