Senior Reliability & Optimization Engineer

Kforce Inc

Englewood (CO)

On-site

USD 140,000 - 200,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
HSA
FSA
401(k) plan
Life insurance
Disability insurance
Paid time off
Paid sick leave

Job summary

Kforce is seeking a Senior Reliability & Optimization Engineer to support an enterprise client in Greenwood Village, CO. You will own platform engineering, manage Terraform modules, and oversee AWS/GitLab/CD pipelines from development through production.

The role emphasizes cost-aware scaling, monitoring with DataDog/Splunk, and incident response, with collaboration across dev teams to improve reliability and performance.

Qualifications

  • Bachelor's degree or equivalent professional experience in a related field.
  • Strong hands-on AWS operations across multiple services (EKS, S3, DocumentDB, Aurora/RDS, Redis, Lambda, Route53, WAFv2, MQ).
  • Experience maintaining GitLab CI/CD pipelines for multi-stage environments.
  • Production experience with Docker/Kubernetes and containerized microservices.
  • Proven incident response and root-cause analysis under SLA pressure.
  • Ability to benchmark, performance-test, and optimize resource usage.

Responsibilities

  • Maintain Terraform modules and audit AWS against state, reconciling drift.
  • Own the health of GitLab CI/CD pipelines and enforce multi-environment deployment.
  • Operate the AWS footprint (EKS, Helm, Istio; Aurora/DocumentDB/Redis; Route53, WAFv2, CloudFront; S3).
  • Keep resources current and tagged; remediate vulnerabilities via IaC.
  • Build, deploy, and validate releases across environments; document release notes/runbooks.
  • Right-size resources to meet SLAs and optimize costs; collaborate with developers to improve performance.
  • Own monitoring and alerting end-to-end (DataDog/Splunk); perform incident response and post-incident reviews.

Skills

AWS operations
CI/CD pipelines
Containerized microservices
Incident response
Performance optimization
Observability
GitLab workflows
Terraform / IaC

Education

Bachelor's degree in Computer Science, Engineering, or related field

Tools

Terraform
GitLab
Docker
Kubernetes
DataDog
Splunk

Job description

Responsibilities

Kforce is immediately seeking an experienced Senior Reliability & Optimization Engineer in support of our enterprise telecommunications and mass media client based in Greenwood Village, CO. Responsibilities: Platform Engineering (Primary):

  • Maintain and extend the Terraform modules that define our infrastructure; Audit AWS against state, detect drift, and reconcile it rather than letting the console diverge
  • Own the health of the GitLab CI/CD pipelines and enforce pipeline-only deployment with progressive promotion (dev, qa, uat, stage, prod); Never skip environments or patch production in the console
  • Operate the AWS footprint: EKS (blue/green upgrades, node groups, IRSA), Helm, and Istio; Data and messaging (Aurora, DocumentDB, Redis, Amazon MQ); Networking and edge (Route53, WAFv2, CloudFront); And storage (S3)
  • Keep resources current and tagged to standard and remediate scan-flagged vulnerabilities through IaC
  • Build, deploy, and validate releases across environments, including off-hours windows; Document release notes and runbooks
Responsibilities

Kforce is immediately seeking an experienced Senior Reliability & Optimization Engineer in support of our enterprise telecommunications and mass media client based in Greenwood Village, CO. Responsibilities: Platform Engineering (Primary):

  • Maintain and extend the Terraform modules that define our infrastructure; Audit AWS against state, detect drift, and reconcile it rather than letting the console diverge
  • Own the health of the GitLab CI/CD pipelines and enforce pipeline-only deployment with progressive promotion (dev, qa, uat, stage, prod); Never skip environments or patch production in the console
  • Operate the AWS footprint: EKS (blue/green upgrades, node groups, IRSA), Helm, and Istio; Data and messaging (Aurora, DocumentDB, Redis, Amazon MQ); Networking and edge (Route53, WAFv2, CloudFront); And storage (S3)
  • Keep resources current and tagged to standard and remediate scan-flagged vulnerabilities through IaC
  • Build, deploy, and validate releases across environments, including off-hours windows; Document release notes and runbooks
Optimization
  • Right-size and scale resources to meet SLAs at the lowest sustainable cost, distinguishing real demand growth from regressions and leaks, and recommending configuration changes that keep latency and spend on target
  • Partner with developers and test engineers to improve application performance and strengthen automated test coverage for reliability-sensitive changes
Reliability & Incident Response:
  • Own monitoring and alerting end to end: build DataDog/Splunk dashboards, tune thresholds and routing against SLAs, and pre-empt degradation
  • Provide first response, mitigation, recovery, and root-cause analysis for incidents and outages; Participate in on-call rotation and drive follow-ups to closure
Requirements
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent professional experience
  • AWS operations: Strong hands-on experience operating AWS through both the console and Infrastructure as Code, across services such as EKS, S3, DocumentDB, Aurora/RDS, ElastiCache Redis, Lambda, Route53, WAFv2, and Amazon MQ
  • CI/CD: Experience maintaining GitLab (or comparable) CI/CD pipelines for infrastructure and application deployments, including multi-stage environment promotion
  • Containerized microservices: Production experience operating containerized microservice and web-based applications (Docker/Kubernetes)
  • Incident response: Demonstrated experience owning production incident triage, mitigation, and root-cause analysis under SLA pressure
  • Performance work: Experience benchmarking, performance testing, and optimizing resource usage and system behavior
  • Observability: Expertise with monitoring and alerting tooling such as DataDog and/or Splunk - building dashboards, tuning alert thresholds, and using telemetry to investigate latency, errors, and saturation
  • Source control: Familiar with Git-based source control and branch/merge-request workflows (e.g., GitLab)
  • Infrastructure as Code: Proficient with Terraform - reading, maintaining, and authoring modules; Managing state; and detecting and reconciling drift; Comfortable making changes through code and pipelines rather than the console

The pay range is the lowest to highest compensation we reasonably in good faith believe we would pay at posting for this role. We may ultimately pay more or less than this range. Employee pay is based on factors like relevant education, qualifications, certifications, experience, skills, seniority, location, performance, union contract and business needs. This range may be modified in the future.

  • We offer comprehensive benefits including medical/dental/vision insurance, HSA, FSA, 401(k), and life, disability & ADD insurance to eligible employees.
  • Salaried personnel receive paid time off.
  • Hourly employees are not eligible for paid time off unless required by law.
  • Hourly employees on a Service Contract Act project are eligible for paid sick leave.
  • This job is not eligible for bonuses, incentives or commissions.

Kforce is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, pregnancy, sexual orientation, gender identity, national origin, age, protected veteran status, or disability status.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Platform Engineer
Senior Platform Engineer

Flexential Corp. • Denver (CO)

On-site
USD 150,000 - 165,000
Medical+Dental+Vision
401(k)
HSA/FSA
+7
Senior DevOps Engineer
Senior DevOps Engineer

Kforce Inc • New York (NY)

On-site
USD 140,000 - 200,000
DevOps Engineer
DevOps Engineer

Kforce Inc • Stamford (CT)

On-site
USD 140,000 - 170,000
Cloud Platform Engineer
Cloud Platform Engineer

Kforce Inc • Fort Collins (CO)

On-site
USD 110,000 - 160,000
Medical insurance
Dental insurance
Vision insurance
+4
Senior Platform Engineer
Senior Platform Engineer

Far Coder • Northern (KY)

Hybrid
USD 150,000 - 165,000
Medical, Telehealth, Dental and Vision
401(k)
HSA and FSA
+5
Cloud Infrastructure Engineer
Cloud Infrastructure Engineer

Eliassen Group • Greenwood Village (CO)

On-site
USD 76,000 - 80,000
Resiliency Junior IT SME (Phoenix Initiative)
Resiliency Junior IT SME (Phoenix Initiative)

Kforce Inc • New York (NY)

On-site
USD 85,000 - 140,000
Medical, dental, vision
401(k) retirement plan
Paid time off
Cloud Engineer
Cloud Engineer

Kforce Inc • San Mateo (CA)

On-site
USD 150,000 - 190,000
Cloud Platform SRE Engineer #11145
Cloud Platform SRE Engineer #11145

ECCO Select • Dallas (TX)

Hybrid
USD 110,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Skill • Southlake (TX)

On-site
USD 66,000 - 73,000
Health insurance
Vision insurance
Dental insurance
+2