Software Engineer (Platform Data Reliability)

PlayStation

San Mateo (CA)

On-site

USD 140,000 - 180,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

PlayStation is seeking a Software Engineer II focused on Platform Data Reliability & Automation. You will build, automate, and operate scalable data platforms using Infrastructure as Code and cloud technologies across AWS and GCP.

You’ll improve reliability of NoSQL, streaming, and caching services and contribute to observability tooling and operational playbooks. Working with platform and product teams, you’ll write Go code, participate in on-call rotations, and help define SLIs/SLAs while

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or a related field, or equivalent practical experience.
  • Observability tools and practices: metrics, logging, tracing, alerting, dashboards.
  • Hands-on experience with Terraform or Ansible for IaC/configuration management.
  • Experience writing production Go code, with testing, concurrency, and maintainability in mind.
  • 3+ years in software engineering, site reliability or platform engineering.
  • Familiarity with NoSQL, caching, or streaming tech (Cassandra, Aerospike, Kafka, Redis).
  • Experience with AWS or GCP and managed services like MSK, DynamoDB, ElastiCache, Memorystore.

Responsibilities

  • Develop, automate, and operate scalable data platforms using IaC and cloud technologies.
  • Improve reliability and automation of NoSQL, streaming, and caching services across AWS/GCP.
  • Build tooling for deployment, observability, and operational workflows.
  • Contribute to incident response, root-cause analysis, and permanent fixes.
  • Collaborate with engineering, platform, security, and operations teams.

Skills

Go programming
Infrastructure as Code
Terraform
Ansible
Observability
Linux
Kubernetes
Cloud services (AWS/GCP)
Incident response
Communication

Education

Bachelor’s or Master’s degree in Computer Science or related field

Tools

Kafka
Cassandra
Aerospike
Redis
AWS MSK
DynamoDB
ElastiCache
Memorystore

Job description

  • We are seeking a Software Engineer II (Platform Data Reliability & Automation) to help build, automate, and operate scalable data platforms using Infrastructure as Code (IaC) and cloud technologies
  • This role focuses on improving the reliability and automation of NoSQL, streaming, and caching services across AWS and GCP environments. You’ll develop automation, observability, and operational tooling supporting technologies such as Cassandra, Aerospike, Kafka, and Redis
  • Working alongside senior engineers, platform teams, and product teams, you’ll contribute to highly available infrastructure supporting billions of transactions and millions of players globally. By applying software engineering and database reliability engineering principles, you’ll help reduce manual work, improve system uptime, and make data services easier and safer for engineering teams to use
  • Develop, maintain, and improve Infrastructure as Code and configuration-management automation using tools such as Terraform and Ansible to provision, configure, monitor, scale, and manage NoSQL, streaming, and caching platforms
  • Build automation that enables repeatable and reliable deployment of data services across cloud and hybrid environments
  • Contribute to the reliability, availability, scalability, performance, and resiliency of platform data services
  • Contribute to defining, measuring, and improving service-level indicators, service-level objectives, and error budgets
  • Develop automation for operational activities such as scaling, failover, backup, recovery, upgrades, and routine maintenance
  • Build and enhance observability solutions using metrics, logging, tracing, dashboards, and alerts
  • Troubleshoot issues affecting Cassandra, Aerospike, Kafka/MSK, Redis, and related platform services
  • Participate in on-call rotations and incident response, contributing to root-cause analysis and the implementation of permanent fixes
  • Write reliable, maintainable, and well-tested Go code for infrastructure automation, platform services, and operational tooling
  • Collaborate with engineering, platform, security, and operations teams to integrate and deliver reliable data services
  • Create and maintain operational documentation, procedures, runbooks, and automation playbooks
  • Participate in code reviews, technical design discussions, and continuous improvement initiatives
  • Explore practical applications of AI-assisted automation, anomaly detection, automated remediation, and developer-productivity tooling where appropriate
  • Bachelor’s or Master’s degree in Computer Science or a related field, or equivalent practical experience
  • Familiarity with observability tools and practices, including metrics, logging, tracing, alerting, and dashboard creation
  • Hands-on experience developing or maintaining Infrastructure as Code or configuration-management tools such as Terraform or Ansible
  • Experience developing production software in Go, with an understanding of idiomatic code, testing, concurrency, error handling, and maintainability
  • 3+ years of experience in software engineering, database reliability engineering, site reliability engineering, platform engineering, or a related field
  • Working knowledge of one or more NoSQL, caching, or streaming technologies, such as Cassandra, Aerospike, Kafka, AWS MSK, or Redis
  • Experience with AWS or GCP and familiarity with managed services such as MSK, DynamoDB, ElastiCache, Memorystore, or equivalent technologies
  • Working knowledge of Linux, networking, storage, and common system-troubleshooting techniques
  • Strong written and verbal communication skills, with the ability to collaborate effectively across teams
  • Understanding of distributed systems concepts, including availability, consistency, replication, fault tolerance, and horizontal scaling
  • Experience deploying or operating workloads on Kubernetes
  • Ability to independently diagnose and resolve technical problems within a defined scope, learn unfamiliar systems, and seek guidance when addressing complex or ambiguous challenges
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform Data Reliability Engineer II | Go & IaC Automation
Platform Data Reliability Engineer II | Go & IaC Automation

PlayStation • San Mateo (CA)

Hybrid
USD 150,000 - 225,000
Medical insurance
Dental insurance
Vision
+4
Senior Platform Engineer
Senior Platform Engineer

Selby Jennings • New York (NY)

On-site
USD 140,000 - 180,000
Platform Engineer
Platform Engineer

Synergy • Chicago (IL)

On-site
USD 100,000 - 150,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Senior Engineer, Platform & Site Reliability
Senior Engineer, Platform & Site Reliability

ICE Clear Europe Limited • Jacksonville (FL)

On-site
USD 120,000 - 170,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Level AI • Auckland (CA)

On-site
USD 140,000 - 210,000
Senior Engineer, Platform & Site Reliability
Senior Engineer, Platform & Site Reliability

ICE • Jacksonville (FL)

On-site
USD 140,000 - 190,000
Platform Engineer
Platform Engineer

Compunnel, Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Level AI • Bellevue (WA)

On-site
USD 150,000 - 190,000
Senior Engineer, Platform & Site Reliability
Senior Engineer, Platform & Site Reliability

Intercontinental Exchange Holdings, Inc. • Jacksonville (FL)

On-site
USD 140,000 - 180,000