DevOps (GitLab-based platform CICD 30*3 pipelines)

Bitdeer Technologies Group

San Jose (CA)

On-site

USD 160,000 - 220,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bitdeer Technologies Group is building an AI-driven GPU cloud and seeks a Cloud Senior DevOps Engineer to shape CI/CD, IaC, and MLOps pipelines. You will connect research with production, operating the platform that supports model, data-science, and product teams.

You will own high-availability design, automation, and observability across multi-cloud environments, driving scalable governance and efficient deployment workflows in a fast-paced AI infrastructure group.

Qualifications

  • Bachelor's degree or higher in a technical field with 5+ years in DevOps, SRE, or Cloud Infra.
  • Expert Linux and core networking concepts (TCP/IP, DNS, HTTP, Load Balancing, VPCs) are required.
  • Deep Docker and Kubernetes knowledge with production-grade cluster management.
  • Experience across major cloud platforms (AWS, GCP, Azure, etc.) and multi/hybrid cloud setups.
  • Strong coding/scripting skills (Go, Python, Shell) focused on automation and tooling.
  • Solid understanding of CI/CD, IaC, observability, and SRE principles.
  • Excellent problem-solving and cross-team collaboration abilities.

Responsibilities

  • Design, implement, and maintain end-to-end CI/CD pipelines for software and ML models.
  • Automate build, test, deployment, and rollback processes to production.
  • Build and scale cloud-native infrastructure using Kubernetes and Docker.
  • Provision GPU clusters and high-performance resources for AI workloads.
  • Own high-availability design, DR strategies, capacity planning, and tuning.
  • Champion IaC with Terraform, Ansible, and Helm for automated provisioning.
  • Develop comprehensive monitoring and logging with Prometheus, Grafana, and ELK/EFK.
  • Create and maintain an Internal Developer Platform to accelerate deployments.
  • Lead cross-functional collaboration with R&D, Data Science, Security, and Business.

Skills

Problem solving
Cross-team communication
Leadership
Automation mindset
Strategic thinking

Education

Bachelor's degree in Computer Science, Engineering, or related field

Tools

Docker
Kubernetes
Terraform
Ansible
Helm
Prometheus
Grafana
ELK Stack
Linux

Job description

About Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia. To learn more, visit https://ir.bitdeer.com/.

Job Description

You are the paved road for every AI product we ship - the CI/CD, IDP, and MLOps substrate that lets model teams deploy without opening a ticket.

Position Summary

Bitdeer is building an AI-operated GPU cloud, and the AI Cloud team ships the products on top of it. As Cloud Senior DevOps Engineer you build the paved road those products travel down: CI/CD, infrastructure-as-code, MLOps pipelines, and the Internal Developer Platform that lets model, data-science, and product teams deploy at speed without stepping around governance. You are the bridge between research and production, and the operator of the platform substrate the AIOps team plugs into.

Key Responsibilities
  • CI/CD & MLOps Pipeline Management: Design, implement, and maintain end-to-end CI/CD pipelines for both software applications and machine learning models. Automate build, test, deployment, and rollback processes to ensure seamless transitions from innovation to production.
  • Cloud-Native & AI Infrastructure: Build, optimize, and scale cloud-native infrastructure using Kubernetes (K8s) and Docker. Manage and provision specialized computing resources (e.g., GPU clusters) to support high-performance AI workloads and model inferencing.
  • High Availability Architecture: Take ownership of high-availability design in production environments. Implement disaster recovery (DR) strategies, self-healing mechanisms, capacity planning, and performance tuning to meet stringent business SLAs.
  • Infrastructure as Code (IaC): Champion IaC practices utilizing tools such as Terraform, Ansible, and Helm to achieve fully automated, reproducible, and auditable infrastructure provisioning across multiple cloud environments.
  • Observability & Monitoring: Architect and refine comprehensive monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, ELK/EFK stack) to provide deep visibility into system health, application performance, and AI model metrics. Feed the AIOps substrate with clean, well-labeled telemetry.
  • Internal Developer Platform (IDP): Build the paved road that lets product, model, and data-science teams deploy without opening a ticket. Make golden paths so obvious that shortcuts feel harder than doing it right.
  • Cross-functional Collaboration: Work closely with R&D, Data Science, Security, and Business teams to streamline workflows, eliminate bottlenecks, and continuously elevate engineering efficiency.
  • Governance, Security & Compliance: Establish and enforce system stability and security standards. Manage release workflows, implement Zero Trust access controls, oversee secrets management, and ensure compliance (e.g., SOC2, ISO27001).
  • Incident Management & Resolution: Act as the technical lead during complex system anomalies and major incidents. Spearhead rapid troubleshooting, thorough RCA, and preventative remediation — and turn each incident into an automation that stops the next one before it pages a human.
Basic Qualifications
  • Experience & Education: Bachelor's degree or above in Computer Science, Engineering, or a related technical field, with 5+ years of hands-on experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles.
  • Networking & OS: Expert-level knowledge of Linux operating systems and core networking principles (TCP/IP, DNS, HTTP, Load Balancing, VPCs).
  • Containerization & Orchestration: Deep mastery of Docker and Kubernetes orchestration, including a thorough understanding of underlying principles, cluster management, and production-level best practices.
  • Cloud Platforms: Proven proficiency in designing and managing infrastructure on major Public or Hybrid Cloud platforms (e.g., AWS, GCP, Azure, Alibaba Cloud), including multi-cloud and hybrid-cloud strategies.
  • Programming Skills: Strong coding and scripting capabilities in at least one major language (Go, Python, Shell, etc.) with a solid engineering-oriented mindset focused on automation and tooling development.
  • Domain Knowledge: Systematic and practical understanding of CI/CD methodologies, Infrastructure as Code (IaC), Observability paradigms, and Site Reliability Engineering (SRE) principles.
  • Soft Skills: Exceptional problem-solving abilities, sharp technical judgment, and excellent cross-team communication skills to effectively collaborate in a fast-paced, dynamic environment.
Preferred Qualifications (Plus)
  • AI/ML Infrastructure Experience: Familiarity with MLOps practices, model serving/inferencing frameworks (e.g., vLLM, TGI, Triton Inference Server), and experience managing GPU clusters for AI/ML workloads.
  • Large-Scale Systems: Proven track record working with large-scale distributed systems or high-concurrency environments (e.g., Fintech, Trading, Real-time processing, or AI platforms).
  • Platform Engineering: Hands-on experience in designing and building Internal Developer Platforms (IDP) to enhance developer autonomy and productivity.
  • Advanced Security: Deep familiarity with Zero Trust architecture, automated security testing (DevSecOps), and implementing strict compliance frameworks (e.g., SOC2, ISO27001).
  • Leadership: Prior experience acting as a Technical Lead, mentoring junior engineers, or managing DevOps teams.
  • AIOps Builder Mindset: You see incidents, tickets, and manual runbooks as source material for the next automation, not as steady-state work. You've either wired an LLM-driven code/config helper into a pipeline or have strong opinions on how to.
  • Paved-Road Philosophy: You think in golden paths and defaults, not policies; you'd rather make the right thing easy than write a doc scolding people for doing the wrong thing.

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps (GitLab-based platform CICD 30*3 pipelines)
DevOps (GitLab-based platform CICD 30*3 pipelines)

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 190,000
Cloud Service Security Platform DevOps & Maintenance
Cloud Service Security Platform DevOps & Maintenance

Bitdeer • San Jose (CA)

On-site
USD 140,000 - 210,000
Cloud Service Security Platform DevOps & Maintenance
Cloud Service Security Platform DevOps & Maintenance

Bitdeer Technologies Group • Austin (TX)

On-site
USD 130,000 - 190,000
Cloud Service Security Platform DevOps & Maintenance
Cloud Service Security Platform DevOps & Maintenance

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 190,000
Cloud Service Security Platform DevOps Maintenance
Cloud Service Security Platform DevOps Maintenance

Bitdeer Technologies Group • United States

On-site
USD 140,000 - 230,000
AI Cloud Senior DevOps Engineer
AI Cloud Senior DevOps Engineer

Bitdeer • San Jose (CA)

On-site
USD 150,000 - 210,000
AI Cloud Senior DevOps Engineer
AI Cloud Senior DevOps Engineer

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 180,000 - 280,000
AI Cloud Senior DevOps Engineer
AI Cloud Senior DevOps Engineer

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 190,000
AI Cloud Senior DevOps Engineer
AI Cloud Senior DevOps Engineer

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 180,000 - 240,000
SRE Platform Software Engineer (Early Career / Temporary)
SRE Platform Software Engineer (Early Career / Temporary)

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 90,000 - 130,000