Site Reliability Engineer

Knack Solutions

Richmond (VA)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology solutions provider seeks a Site Reliability Engineer in Richmond, Virginia. The role involves developing and operating secure cloud infrastructure, ensuring the availability, performance, and compliance of services based on AWS and Kubernetes. The ideal candidate has expertise in container management, automation, operational monitoring, and is familiar with the full software development lifecycle. Strong collaboration with engineering and product teams is essential to enhance the efficiency of cloud services.

Qualifications

  • Experience with container management technologies including Docker and Kubernetes.
  • Experience with AWS, covering EKS, ECS, IAM, S3, and more.
  • Familiarity with automation/configuration management using Terraform or similar solutions.

Responsibilities

  • Develop, deploy, and operate secure infrastructure using cloud services.
  • Ensure high availability and compliance of cloud services.
  • Define SLA standards for SaaS solutions.

Skills

Container management technologies (Docker, Kubernetes)
AWS services (EKS, ECS, IAM, S3, RDS)
Automation/configuration management (Terraform)
CI tools (Jenkins)
Operational monitoring tools (Datadog, NewRelic, Splunk)
Linux tools and scripting
Large-scale distributed systems
Software development lifecycle

Job description

Position: Site Reliability Engineer (SRE)

Work Authorization: All Work Authorizations

Contract: 24 months

As one of the Site Reliability Engineers, youll be able to work closely with customers, product management, and other subject matter experts in the technology industry to drive forward solutions that have immediate impact on the day-to-day ability for other data scientists and machine learning engineers to productionize their models by iteratively improving how we operate and scale our cloud based containerized service.

What You'll Do
  • Develop, deploy, and operate our secure infrastructure built on cloud services (AWS, Kubernetes, etc)
  • Ensure the high availability, resiliency, performance, business continuity and compliance capabilities of our cloud services.
  • Define SLA standards for SAAS solutions that are used by several groups within the company.
  • Work with our engineering teams to deploy and operate cloud services, scale our development, QA and production environments.
  • Build solutions for developer productivity. Develop and operate our build automation and continuous delivery systems.
  • Participate in an on-call rotation, drive incident resolution and improve platform resiliency
Basic Qualifications
  • Experience with container management technologies including Docker and Kubernetes.
  • Experience with AWS including EKS, ECS, IAM, S3, RDS, Security Groups, Route53, VPC Flow Logs, etc.
  • Experience with automation/configuration management using Terraform or similar solutions.
  • Experience with CI tools such as Jenkins.
  • Experience with operational monitoring tools, such as Datadog, NewRelic and Splunk.
  • Proficient in Linux tools and shell scripting or other Linux automation
  • An interest in designing, analyzing and troubleshooting large-scale distributed systems.
  • Well-versed with the entire software development lifecycle, devops, and SRE practices.
Preferred Qualifications
  • Experience with automated unit and integration testing of infrastructure code
  • Experience with container security and vulnerability management
  • Experience in one or more languages such as Python or GoLang
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Madison-Davis, LLC • United States

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Latent • San Francisco (CA)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Hampshire

On-site
USD 110,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Veriipro • Washington

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000