AWS DevOps / Agentic SRE Engineer (Active TS/SCI Clearance)

Strategic Business Systems, Inc (SBS)

Chantilly, Northern (VA, KY)

Hybrid

USD 180,000 - 270,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Telework where mission permits
Professional development support for A

Job summary

Strategic Business Systems, Inc. (SBS) in Chantilly, VA is seeking an AWS DevOps / Agentic SRE Engineer with an active TS/SCI clearance to own deployment, reliability, and observability of secure AWS environments supporting an agentic AI platform.

This hands-on role combines DevOps, SRE, and release automation, covering the path from local development to production in AWS EKS, CDK, Helm, Flux GitOps, and modern tracing and observability tools.

Qualifications

  • Active TS/SCI clearance required.
  • Current Security+ or equivalent certification.
  • 3+ years of SRE/DevOps/platform/infrastructure engineering experience.
  • Hands-on experience operating Kubernetes in production environments.
  • Experience authoring and maintaining Helm charts.
  • Experience implementing GitOps-based continuous delivery using Flux, Argo CD, or equivalent technologies.
  • Hands-on AWS experience with EKS, RDS, S3, IAM/IRSA, and ECR.
  • Experience developing and maintaining Infrastructure as Code.
  • Experience with TypeScript / AWS CDK in IaC environment.
  • Experience implementing production observability and tracing (OpenTelemetry, Grafana, Tempo).
  • Demonstrated incident investigation and RCA skills.
  • Proficiency with Python or other scripting for automation.
  • Strong CI/CD, containers, networking, security, and modern cloud architecture knowledge.

Responsibilities

  • Design, build, maintain, and operate highly available AWS infrastructure for an Agentic AI platform.
  • Own the software delivery lifecycle from local development through build, test, packaging, promotion, and production deployment into AWS environments.
  • Deploy and operate Kubernetes workloads in production AWS EKS environments.
  • Author, maintain, and troubleshoot Helm charts for platform applications and services.
  • Build and maintain GitOps-based CD using Flux, Helm, and registries.
  • Develop Infrastructure as Code using AWS CDK and TypeScript.
  • Manage AWS infrastructure including EKS, RDS, S3, IAM/IRSA, and ECR.
  • Maintain container-image and Helm-chart promotion processes.
  • Configure Envoy Gateway routing, certificates, and TLS.
  • Implement observability and distributed tracing via OpenTelemetry, Grafana, Tempo.
  • Establish end-to-end tracing across services.
  • Monitor health and proactively address reliability issues.
  • Diagnose Helm/Flux failures, drift, and deployment issues.
  • Lead production incident investigations and root-cause analyses.
  • Develop preflight validation, diagnostic, QA, and tooling.
  • Automate infrastructure and ops with Python and scripting.
  • Support credential, certificate, and secrets rotation.
  • Collaborate with software and AI/ML teams to deploy agentic apps and model-serving infra.
  • Support platform security, compliance, IATT, and ATO.
  • Use AI coding assistants with validation before deployment.
  • Drive engineering via telemetry, testing, and verification.

Skills

TS/SCI Clearance
Security+ Certification
3+ years SRE/DevOps
Kubernetes in production
Helm charts
GitOps (Flux/Argo CD)
AWS (EKS/RDS/S3/IAM/IRSA/ECR)
Infrastructure as Code
TypeScript / AWS CDK
Observability (OpenTelemetry/Grafana/"
Incident RCA
Python scripting
CI/CD / Containers / Networking / Sec

Tools

Kubernetes
Flux
Argo CD
Helm
AWS CDK
Python

Job description

AWS DevOps / Agentic SRE Engineer (Active TS/SCI Clearance)
  • Chantilly, Virginia
AWS DevOps / Agentic SRE Engineer – TS/SCI

Clearance: Active TS/SCI
Certification: Active Security+ or equivalent required
Experience: 3+ years of SRE, DevOps, Platform Engineering, or Infrastructure Engineering experience

Position Overview

We are seeking an SRE / DevOps & Release Engineer to own the deployment, reliability, observability, and operation of secure AWS environments supporting an advanced Agentic AI platform.

This is a hands‑on engineering position combining DevOps, Site Reliability Engineering, platform engineering, and release automation. The engineer will own the path from local development through production deployment into AWS Kubernetes environments, as well as the operational health and reliability of those environments.

The platform is currently operating within IATT environments and progressing toward scale and ATO. At the current team size, build/release and production operations are intentionally combined — the engineer responsible for deploying the platform also has responsibility for ensuring that it operates reliably.

The environment includes AWS EKS, CDK, Kubernetes, Helm, Flux GitOps, container registries, Envoy Gateway, distributed tracing, and modern observability technologies.

Key Responsibilities
  • Design, build, maintain, and operate highly available AWS infrastructure supporting an Agentic AI platform.
  • Own the software delivery lifecycle from local development through build, test, packaging, promotion, and deployment into AWS environments.
  • Deploy and operate Kubernetes workloads in production AWS EKS environments.
  • Author, maintain, and troubleshoot Helm charts supporting platform applications and services.
  • Build and maintain GitOps-based continuous delivery utilizing Flux, Helm, and container registries.
  • Develop and maintain Infrastructure as Code using AWS CDK and TypeScript.
  • Build and manage AWS infrastructure utilizing EKS, RDS, S3, IAM/IRSA, ECR, and related AWS services.
  • Build and maintain container-image and Helm-chart promotion processes.
  • Configure and support Envoy Gateway routing, certificates, and TLS.
  • Implement and operate observability and distributed-tracing capabilities utilizing OpenTelemetry, Grafana, Tempo, or equivalent technologies.
  • Establish cross-service tracing to provide end‑to‑end visibility into distributed application and agent workflows.
  • Monitor cluster, application, and service health and proactively identify reliability issues.
  • Diagnose failed or stuck Helm and Flux reconciliations, configuration drift, pod failures, and deployment issues.
  • Lead production incident investigation and root‑cause analysis using Kubernetes state, logs, metrics, and distributed traces.
  • Develop preflight validation, diagnostic, QA, and operational tooling.
  • Automate infrastructure and operational processes using Python and other scripting technologies.
  • Support credential, certificate, and secrets rotation.
  • Partner with software engineers and AI/ML teams to deploy and operate agentic applications, model‑serving infrastructure, and supporting services.
  • Support platform security, compliance, IATT, and ATO activities.
  • Use AI‑assisted development tools to accelerate engineering while independently validating generated code and configuration before deployment.
  • Drive an engineering approach based on measurable evidence, automated testing, telemetry, and verification.
Required Qualifications
  • Current and active TS/SCI security clearance
  • Current Security+ certification or equivalent certification supporting privileged‑user access.
  • 3+ years of professional experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, or Infrastructure Engineering.
  • Hands‑on experience operating Kubernetes in production environments.
  • Experience authoring and maintaining Helm charts.
  • Experience implementing GitOps‑based continuous delivery using Flux, Argo CD, or equivalent technologies.
  • Hands‑on AWS experience with services such as EKS, RDS, S3, IAM/IRSA, and ECR.
  • Experience developing and maintaining Infrastructure as Code.
  • Experience working with TypeScript/AWS CDK or demonstrated ability to work within a TypeScript‑based IaC environment.
  • Experience implementing and utilizing production observability and distributed‑tracing technologies such as OpenTelemetry, Grafana, and Tempo.
  • Demonstrated experience diagnosing infrastructure and application failures using telemetry, logs, metrics, and traces.
  • Experience leading incident investigation, root‑cause analysis, remediation, and validation.
  • Proficiency with Python or another scripting language for infrastructure automation and operational tooling.
  • Strong understanding of CI/CD, containers, networking, security, and modern cloud architecture.
Preferred Qualifications
  • TS/SCI with Poly
  • Experience implementing registry‑based GitOps architectures utilizing ECR or similar container registries.
  • Experience troubleshooting Flux and Helm reconciliation issues, including HelmRelease failures and configuration drift.
  • Experience deploying and operating ML/LLM‑serving infrastructure such as KServe, MLflow, or model‑inference endpoints.
  • Experience supporting AI or Agentic AI platforms.
  • Familiarity with Amazon Bedrock, LLM services, MCP tools, and agent‑based architectures.
  • Ability to analyze model and agent performance characteristics such as latency, token utilization, and cost using distributed traces.
  • Experience with Envoy, API gateways, routing, certificate management, and TLS.
  • Experience supporting secure Federal or Intelligence Community AWS environments.
  • Experience supporting systems through IATT and ATO processes.
  • Experience using AI coding assistants while independently validating generated code and infrastructure changes before deployment.
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline, or equivalent practical experience.
Why SBS

SBS is a trusted AWS partner supporting both federal and commercial customers. Our teams focus on delivery—building and operating secure, scalable solutions that matter.

You’ll be part of a team that:

  • Works directly with AWS on meaningful, mission‑driven programs
  • Operates in high‑impact national security environments
  • Builds and operates modern AWS and Kubernetes platforms
  • Works across cloud infrastructure, DevOps, SRE, GitOps, and AI
  • Values engineers who can build, automate, troubleshoot, and deliver

COMPENSATION & BENEFITS

SBS offers a comprehensive total‑rewards package, including a market competitive salary along with:

Comprehensive medical, dental, and vision coverage; HSA‑eligible plan options available

401(k) retirement plan with company match (vesting schedule per Plan Document)

Paid Time Off, federal holidays, and floating holiday for personal observance

Annual professional development support for AWS certifications, training, and conferences

Employee referral program where applicable and documented by program policy

Life, AD&D, and short- / long‑term disability insurance

Telework and flexible‑schedule support where mission and contract permit

Mission‑focused federal contractor supporting national‑security customers

Target Salary Range: $180,000-270,000. This range reflects the anticipated compensation for this role. Actual salary will be based on a combination of factors, including the position’s scope and level of responsibility, the candidate’s relevant experience, education, technical expertise, skills and qualifications, geographic location, and applicable business or contractual requirements.
About SBS
Strategic Business Systems, Inc. (SBS) is a national Information Technology services company headquartered in the Washington, D.C. metropolitan area. SBS provides IT infrastructure design, integration, and operational services. Our expertise spans the full spectrum of infrastructure technologies, including networking, servers, data storage, disaster recovery, cybersecurity, and internet technologies.
Equal Employment Opportunity
SBS is an equal opportunity employer; all qualified applicants will receive consideration for employment without regard to age, gender, gender identity, sex, sexual orientation, color, race, creed, national origin, religion, marital status, parental status, citizenship status, ancestry, physical or mental disability, genetic information, veteran status, military status, or any other classification protected by federal, state, or local laws.
Accommodations
If you need an accommodation while seeking employment with SBS, please emailhr@sbsplanet.com. Accommodations are made on a case‑by‑case basis.
No Unsolicited Agency Referrals
SBS does not accept unsolicited resumes from staffing agencies. Any resumes submitted without a prior agreement will be considered the property of SBS.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AWS DevOps / Agentic SRE Engineer (Active TS/SCI Clearance)
AWS DevOps / Agentic SRE Engineer (Active TS/SCI Clearance)

Strategic Business Systems • Chantilly (VA)

On-site
USD 180,000 - 270,000
Comprehensive medical, dental, and V/V
401(k) with company match
Paid Time Off and holidays
+1
Full-Stack Engineer – Agentic AI / AWS (Active TS/SCI Clearance)
Full-Stack Engineer – Agentic AI / AWS (Active TS/SCI Clearance)

Strategic Business Systems • Chantilly (VA)

Hybrid
USD 180,000 - 270,000
Telework and flexible-schedule
Comprehensive medical, dental, and vis
401(k) retirement plan with company
+2
AWS DevOps Engineer (Intermediate)
AWS DevOps Engineer (Intermediate)

Strategic Business Systems, Inc (SBS) • Chantilly (VA)

On-site
USD 90,000 - 125,000
Medical, dental, vision
401(k) with company match
Paid time off
+2
AWS Senior Software Developer (Active TS/SCI Preferred)
AWS Senior Software Developer (Active TS/SCI Preferred)

Strategic Business Systems (SBS) • Chantilly (VA)

Hybrid
USD 100,000 - 130,000
Comprehensive medical, dental, and vision coverage
401(k) retirement plan with company match
Annual professional development support for AWS certifications
Senior AWS Cloud Architect (Active TS/SCI )
Senior AWS Cloud Architect (Active TS/SCI )

Strategic Business Systems, Inc (SBS) • Chantilly (VA)

On-site
USD 150,000 - 210,000
Medical, dental, and vision coverage
401(k) with company match
Paid time off and holidays
+2
AWS Senior Software Developer (Active TS/SCI Preferred)
AWS Senior Software Developer (Active TS/SCI Preferred)

Strategic Business Systems • Fort Meade (MD)

On-site
USD 100,000 - 130,000
Comprehensive medical, dental, and vision coverage
401(k) retirement plan with company match
Paid Time Off and floating holiday
+2
Cloud Developer (Intermediate)
Cloud Developer (Intermediate)

Strategic Business Systems (SBS) • Chantilly (NC)

On-site
USD 90,000 - 125,000
Medical coverage
401(k) with match
Paid time off
+5
AWS Software Developer (Cloud / DevOps) w/ CI Poly
AWS Software Developer (Cloud / DevOps) w/ CI Poly

Strategic Business Systems • Fort Meade (MD)

On-site
USD 100,000 - 130,000
Comprehensive medical, dental, and vision coverage
401(k) retirement plan with company match
Paid Time Off and federal holidays
+1
AWS Software Developer (Cloud / DevOps) w/ CI Poly
AWS Software Developer (Cloud / DevOps) w/ CI Poly

Strategic Business Systems (SBS) • Fort Meade (MD)

Hybrid
USD 120,000 - 150,000
Medical, dental, vision coverage
401(k) with company match
Paid time off and holidays
+2
Senior Security Analyst – ATO / RMF (Active TS/SCI Clearance)
Senior Security Analyst – ATO / RMF (Active TS/SCI Clearance)

Strategic Business Systems • Chantilly (VA)

On-site
USD 140,000 - 195,000
Telework and flexible schedule where
401(k) retirement plan with company 4%
Paid Time Off and holidays
+1