Site Reliability Engineer (OpenSearch) - Remote

Information Consulting Services

Herndon (VA)

On-site

USD 120,000 - 190,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Information Consulting Services is seeking a Site Reliability Engineer to ensure availability, performance, scalability, and security for cloud-hosted search services built on OpenSearch. You will work on reliability engineering, operations, automation, and continuous improvement for distributed platforms within a globally distributed team.

The role requires deep OpenSearch knowledge, expert Kubernetes operations, and strong Linux administration.

Qualifications

  • 8+ years of experience in SRE/DevOps/cloud operations with distributed systems.
  • Hands-on experience designing, building, deploying and operating OpenSearch clusters in production.
  • Expert-level Kubernetes operations, troubleshooting, and service management.
  • Strong Linux administration (SUSE and Ubuntu) and scripting.
  • Experience with Git, Concourse pipelines, and cloud monitoring.
  • Experience with Kafka and ZooKeeper; automation for testing and deployment.

Responsibilities

  • Provision, build, deploy, monitor, operate, and support cloud services in a global team environment.
  • Architect and maintain high-performance OpenSearch clusters and platforms from ground up.
  • Optimize OpenSearch for high availability, resiliency, security, and performance.
  • Monitor cluster health, indexing throughput, latency, shard allocation, and storage utilization.
  • Lead incident response, RCA, remediation, and on-call rotation.

Skills

OpenSearch
Kubernetes
Linux
Git
Concourse
Automation
Monitoring

Tools

OpenSearch
AWS
Prometheus
Grafana
Terraform
Jenkins
Kafka
Zookeeper

Job description

  • Type: Contract with potential to convert to full-time after ~12 months (not guaranteed)
About the Opportunity

Seeking a Site Reliability Engineer to ensure availability, performance, scalability, and security for mission-critical, cloud-hosted search and analytics services built on OpenSearch. You will focus on reliability engineering, operations, automation, and continuous improvement for distributed platforms, working within a diverse, globally distributed team.

Key Responsibilities
  • Provision, build, deploy, monitor, operate, and support cloud services in a global team environment.
  • Architect, build, deploy, and maintain high-performance OpenSearch clusters and platforms from the ground up.
  • Optimize OpenSearch for high availability, resiliency, scalability, security, and performance.
  • Monitor and troubleshoot cluster health, node performance, indexing throughput, search latency, shard allocation/replication, and storage utilization.
  • Analyze and resolve operational issues across infrastructure, platform, and application layers; lead incident response, RCA, and remediation.
  • Maintain integrity and security of servers, systems, and OpenSearch platform infrastructure.
  • Support lifecycle activities: installation, configuration, upgrades/patching, backup/restore, and disaster recovery.
  • Develop and maintain monitoring policies, alerting standards, runbooks, and support procedures.
  • Automate testing, deployment, scaling, recovery, and operational workflows for OpenSearch and related cloud services.
  • Plan capacity for compute, memory, storage, and network; partner with engineering to enhance reliability and operational readiness.
  • Support log ingestion, index management, lifecycle/retention, and search performance tuning.
  • Participate in an on-call rotation; support occasional weekend/after-hours needs.
Required Qualifications
  • US citizenship required; dual citizenship not permitted.
  • 8+ years of experience in SRE/DevOps/cloud operations with distributed systems.
  • Proven, hands-on experience designing, building, deploying, operating, and optimizing OpenSearch clusters from scratch in production.
  • Expert-level Kubernetes experience (operations, troubleshooting, management, configuration of complex services).
  • Deep OpenSearch administration: cluster architecture, performance tuning, scaling, upgrades, and troubleshooting; index/shard/replica strategy; sizing; snapshot/restore; backup/DR.
  • Strong Linux expertise (SUSE and Ubuntu).
  • Expertise with Git and Concourse (pipeline setup, management, troubleshooting).
  • Experience with Kafka and Zookeeper; strong automation for testing, deployment, scalability, and cloud service management.
  • Experience building/implementing/supporting cloud monitoring and observability; solid knowledge of cloud computing, infrastructure operations, databases, web services, networking, virtualization, and internet protocols.
  • Security fundamentals for SaaS multi-tenant application systems; excellent communication and prioritization skills; ability to multitask.
Preferred Qualifications
  • AWS experience (e.g., Route 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, VPC); experience deploying/operating OpenSearch in AWS.
  • Experience with Cloud Foundry environments.
  • Experience with Jenkins, Chef, and/or Terraform.
  • Experience with Prometheus and Grafana.
  • Background with log ingestion pipelines, index lifecycle management, retention strategies, and search platform security controls.
  • Familiarity with capacity forecasting, performance benchmarking, and resilience testing for distributed search platforms.
  • Collaborative, globally distributed team with cross-training opportunities.
  • Participation in an on-call rotation and occasional after-hours/weekend support.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

OpenSearch SRE: Cloud Reliability & Ops (Contract)
OpenSearch SRE: Cloud Reliability & Ops (Contract)

Information Consulting Services • Herndon (VA)

On-site
USD 120,000 - 190,000
Remote Site Reliability Engineer: OpenSearch on Kubernetes
Remote Site Reliability Engineer: OpenSearch on Kubernetes

KPG99 INC • Virginia (IL)

On-site
USD 140,000 - 180,000
Remote Site Reliability Engineer - OpenSearch on Kubernetes
Remote Site Reliability Engineer - OpenSearch on Kubernetes

Johnson Technology Systems Inc • United States

On-site
USD 155,000 - 160,000
Software Development Engineer , AWS OpenSearch Intelligent Search Team
Software Development Engineer , AWS OpenSearch Intelligent Search Team

Amazon • Bellevue (WA)

On-site
USD 150,000 - 230,000
Sr. Software Dev Engineer, AWS OpenSearch Service
Sr. Software Dev Engineer, AWS OpenSearch Service

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 168,000 - 228,000
Health insurance
401(k) matching
Paid time off
+1
Software Engineer (SRE + Java/Kotlin /Golang)
Software Engineer (SRE + Java/Kotlin /Golang)

MPower Plus • Austin (TX)

On-site
USD 90,000 - 130,000
Sr. SDE, Amazon OpenSearch Service
Sr. SDE, Amazon OpenSearch Service

Amazon Web Services (AWS) • Santa Clara (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
RSUs
+2
Sr. SDE, Amazon OpenSearch Service
Sr. SDE, Amazon OpenSearch Service

Amazon • Santa Clara (CA)

On-site
USD 193,000 - 262,000
Health insurance (medical, dental, and
401(k) matching
Paid time off
+1
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000
Site Reliability Engineer
Site Reliability Engineer

asobbi • California (MO)

On-site
USD 170,000 - 220,000
Fully remote (US timezone)