Site Reliability Engineer

CodeRound AI

Hinoba-an

On-site

PHP 1,000,000 - 1,800,000

Full time

7 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

CodeRound AI in Hinoba-an, Philippines, seeks an experienced SRE to own the reliability, scalability, and operational excellence of a GenAI platform powering the ML lifecycle. You’ll work across Kubernetes, cloud infrastructure, production systems, observability, automation, and model-serving workloads, while helping build strong SRE and incident-management practices.

Join a fast-growing GenAI startup as an SRE and takeownership of the reliability, scalability, and operational excellence of a

Qualifications

  • Experience building scalable, reliable production platforms.
  • Strong knowledge of Kubernetes, cloud infrastructure, and incident-management practices.
  • Familiarity with observability, monitoring, and release engineering.

Responsibilities

  • Own platform uptime, reliability, scalability, and performance.
  • Manage Kubernetes clusters, cloud infrastructure, and production environments.
  • Establish and improve incident response, on-call, RCA, and postmortem processes.
  • Drive deployment, rollback, and change-management practices.
  • Handle capacity planning and disaster recovery.
  • Build and enhance monitoring, alerting, and operational dashboards.
  • Automate deployments, scaling, and repetitive operational workflows.
  • Troubleshoot complex production infrastructure and application issues.
  • Support GPU workloads and model-serving infrastructure.

Skills

SRE Principles
DevOps Practices
Incident Management
Observability

Tools

Kubernetes
AWS
GCP
Azure
Terraform
Helm
CI/CD
Python
Bash
Go
Linux
Networking
Model Serving
GPU Workloads

Job description

Client: Provider of platform for machine learning model training and deployment. It offers tools for fine-tuning, deploying, and observing machine learning models.

Requirements:

SRE Site Reliability Engineering DevOps Platform Engineering Kubernetes AWS GCP Azure Linux Networking Terraform Helm CI/CD Python Bash Go MLOps AI Infrastructure GPU Model Serving Observability Monitoring Incident Management Production Infrastructure Cloud Infrastructure Release Engineering Reliability Engineering GenAI

Join a fast-growing GenAI startup as an SRE and take ownership of the reliability, scalability, and operational excellence of a platform powering the end-to-end ML lifecycle. You’ll work across Kubernetes, cloud infrastructure, production systems, observability, automation, and model-serving workloads, while helping build strong SRE and incident-management practices.

Responsibilities:
  • Own platform uptime, reliability, scalability, and performance
  • Manage Kubernetes clusters, cloud infrastructure, and production environments
  • Establish and improve incident response, on-call, RCA, and postmortem processes
  • Drive deployment, rollback, and change-management practices
  • Handle capacity planning and disaster recovery
  • Build and enhance monitoring, alerting, and operational dashboards
  • Automate deployments, scaling, and repetitive operational workflows
  • Troubleshoot complex production infrastructure and application issues
  • Support GPU workloads and model-serving infrastructure
Nice to have:

SRE Site Reliability Engineering DevOps Platform Engineering Kubernetes AWS GCP Azure Linux Networking Terraform Helm CI/CD Python Bash Go MLOps AI Infrastructure GPU Model Serving Observability Monitoring Incident Management Production Infrastructure Cloud Infrastructure Release Engineering Reliability Engineering GenAI

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Platform SRE: Reliability, Scale & Automation
GenAI Platform SRE: Reliability, Scale & Automation

CodeRound AI • Hinoba-an

On-site
PHP 1,000,000 - 1,800,000
Site Reliability / Cloud Platform Engineer
Site Reliability / Cloud Platform Engineer

Global Recruitment and Consultancy OPC • Cebu City

On-site
PHP 1,200,000 - 2,400,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Acquire Intelligence • Taguig

On-site
PHP 900,000 - 1,500,000
Site Reliability Engineer
Site Reliability Engineer

Pyramid Consulting, Inc • Mexico

On-site
PHP 5,618,000 - 8,115,000
Senior Site Reliability Engineer SRE Kubernetes
Senior Site Reliability Engineer SRE Kubernetes

Accenture in the Philippines • Cebu City

On-site
PHP 900,000 - 1,500,000
Site Reliability Engineer
Site Reliability Engineer

V2 Solutions • Hinoba-an

On-site
PHP 900,000 - 1,500,000
Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000
Technical Lead - Site Reliability Engineering
Technical Lead - Site Reliability Engineering

LSEG • Taguig

On-site
PHP 4,914,000 - 7,372,000
Healthcare
Retirement planning
Paid volunteering days
+1
Site Reliability Manager
Site Reliability Manager

Google Inc. • Hinoba-an

On-site
PHP 3,000,000 - 6,000,000
Software Engineer- AI-Driven SRE & Cloud SRE
Software Engineer- AI-Driven SRE & Cloud SRE

Keka Technologies Private Limited • Mexico

On-site
PHP 2,215,000 - 3,322,000