Observability Platform Architect

Booz Allen Hamilton

McLean (VA)

On-site

USD 87,000 - 198,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Booz Allen Hamilton is seeking an Observability Platform Architect to set the technical direction for enterprise monitoring across portfolios and accredited environments. You will own instrumentation strategy, automate scale, and define agentic detection and remediation while expanding into Dynatrace, OpenTelemetry, and agentic operations.

The role combines hands-on engineering with oversight, requiring ownership of the platform, its standards, and rigorous validation of automation across

Qualifications

  • 8+ years of engineering experience.
  • 3+ years maintaining responsibility for an enterprise observability or APM platform, including its architecture, deployment, upgrades, standards, and cost.
  • Experience with enterprise observability platforms (Dynatrace, Datadog, New Relic, Grafana, Prometheus, Elastic) including writing/optimizing complex queries.
  • Experience replacing recurring manual platform work with production automation and agentic operations, ML in production, fraud or risk scoring, or security detection.
  • Experience with OpenTelemetry collectors, instrumentation, semantic conventions, and vendor-neutral instrumentation.
  • Experience with IaC and CI/CD (Terraform, GitHub Actions, Azure DevOps, GitLab) and deploying agents/collectors at scale across Windows/Linux/Kubernetes/cloud infra

Responsibilities

  • Architect and build the observability platform across all accredited environments.
  • Drive OpenTelemetry adoption hands-on, including collector configuration and instrumentation rollout.
  • Automate the platform itself: service onboarding, agent deployment, tagging, dashboards, alert provisioning, and access management.
  • Design and deploy agentic detection, triage, and remediation, defining what an agent may do unattended.
  • Own the platform's AI capability end to end, including what is enabled and how decisions are validated.
  • Build the entity model, dashboards, and SLOs used by leadership.

Skills

8+ years of experience
Observability platform experience
OpenTelemetry
Infrastructure as code/CI/CD
Code writing/shipping capabilities

Education

Bachelor's degree in CS, Information Systems, or Engineering

Tools

Dynatrace
Datadog
New Relic
Grafana
Prometheus
Elastic

Job description

Join a culture of empowerment and connectivity.

Develop your craft

Learn the skills you need to accelerate your career.

Discover benefits that your life and work.

Innovate with intention

Build mission-ready tech that protects the nation.

Are you looking for an opportunity to combine hands-on engineering with big picture thinking at the point where AI stops being a demonstration and starts running production operations? You understand that autonomous operations is not a tooling problem but a trust problem — what a system is permitted to do on its own, and how you prove that decision was sound.

As an observability platform architect on our team, you'll set the technical direction for how the enterprise monitors its business services across multiple portfolios and separately accredited environments. Your customers will trust you not only to design the platform, but to build it alongside the team — owning instrumentation strategy and the entity model, automating an estate this size so it stays manageable by a team this size, and defining what agentic detection, triage, and remediation are permitted to do unattended. On our team, you'll broaden into Dynatrace, OpenTelemetry, and agentic operations in environments most engineers never get the opportunity to work in. This role sets the standards and then builds against them alongside the team, so you'll stay in the code. The ideal candidate comes from a Platform Engineering, Observability, Monitoring Infrastructure, or Site Reliability Engineering background with genuine ownership of an observability platform: its architecture, standards, and operation across an organization.

Join us. The world can't wait.

What You'll Work On:

Architect and build the observability platform across every accredited environment, and keep them from drifting into separately maintained estates.

Drive OpenTelemetry adoption hands-on, including collector configuration, instrumentation, and rollout, and set the standards, naming conventions, and onboarding patterns the team builds against.

Automate the platform itself, including service onboarding, agent deployment, tagging, dashboard and alert provisioning, and access management, delivered as code through CI/CD.

Design and deploy agentic detection, triage, and remediation, defining what an agent may do unattended, how we verify the service actually recovered, how it rolls back, and what evidence it leaves behind.

Own the platform's AI capability end to end, including what is enabled, what it is trusted to decide, how its accuracy holds up over time, and how incorrect actions get caught.

Build the entity model, dashboards, and SLOs that leadership depends on, and drive down cost per monitored service.

You Have:

8+ years of experience with engineering experience

3+ years of experience maintaining responsibility for an enterprise observability or APM platform itself, including its architecture, deployment, upgrades, standards, and cost across an organization

Experience with an enterprise observability platform, such as Dynatrace, Datadog, New Relic, Grafana, Prometheus, or Elastic, including writing and optimizing complex queries against its data

Experience replacing recurring manual platform work with production automation, and shipping a system that makes decisions automatically, including agentic operations, ML in production, fraud or risk scoring, or security detection, including owning whether those decisions stayed correct over time

Experience with OpenTelemetry, including collectors, instrumentation, semantic conventions, and the practical limits of vendor-neutral instrumentation

Experience with infrastructure as code and CI/CD, such as Terraform, GitHub Actions, Azure DevOps, or GitLab, and deploying agents or collectors at scale across Windows, Linux, Kubernetes, cloud platforms, and network infrastructure

Ability to write, review, and ship code

Bachelor's degree in CS, Information Systems, or Engineering

Nice If You Have:

Experience with Dynatrace in depth, including Grail, DQL, Workflows, Site Reliability Guardian, Davis, and Dashboards, with configuration managed through the Terraform provider

Experience building automated remediation in production, including an automation that acted wrongly and what changed afterward

Experience with multiple observability platforms, such as Grafana, Prometheus, Datadog, or Elastic

Experience operating observability in DoD cloud environments such as Azure or AWS IL5/IL6, including working around feature gaps between commercial and authorized offerings

Experience operating one platform across more than one environment or security boundary, and managing the drift and duplication that creates

Knowledge of what observability platforms operating models costs to run

Secret clearance

Bachelor's degree in CS, Information Systems, or Engineering

Clearance:

Applicants selected will be subject to a security investigation and may need to meet eligibility requirements for access to classified information; Secret clearance is required.

Compensation

At Booz Allen, we celebrate your contributions, provide you with opportunities and choices, and support your total well-being. Our offerings include health, life, disability, financial, and retirement benefits, as well as paid leave, professional development, tuition assistance, work-life programs, and dependent care. Our recognition awards program acknowledges employees for exceptional performance and superior demonstration of our values. Full-time and part-time employees working at least 20 hours a week on a regular basis are eligible to participate in Booz Allen’s benefit programs. Individuals that do not meet the threshold are only eligible for select offerings, not inclusive of health benefits. We encourage you to learn more about our total benefits by visiting the Resource page on our Careers site and reviewing Our Employee Benefits page.

Salary at Booz Allen is determined by various factors, including but not limited to location, the individual’s particular combination of education, knowledge, skills, competencies, and experience, as well as contract‑specific affordability and organizational requirements. The projected compensation range for this position is $86,800.00 to $198,000.00 (annualized USD). The estimate displayed represents the typical salary range for this position and is just one component of Booz Allen’s total compensation package for employees. This posting will close within 90 days from the Posting Date.

Observability Platform Architect

The Opportunity:

Are you looking for an opportunity to combine hands‑on engineering with big picture thinking at the point where AI stops being a demonstration and starts running production operations? You understand that autonomous operations is not a tooling problem but a trust problem — what a system is permitted to do on its own, and how you prove that decision was sound.

As an observability platform architect on our team, you'll set the technical direction for how the enterprise monitors its business services across multiple portfolios and separately accredited environments. Your customers will trust you not only to design the platform, but to build it alongside the team — owning instrumentation strategy and the entity model, automating an estate this size so it stays manageable by a team this size, and defining what agentic detection, triage, and remediation are permitted to do unattended. On our team, you'll broaden into Dynatrace, OpenTelemetry, and agentic operations in environments most engineers never get the opportunity to work in. This role sets the standards and then builds against them alongside the team, so you'll stay in the code. The ideal candidate comes from a Platform Engineering, Observability, Monitoring Infrastructure, or Site Reliability Engineering background with genuine ownership of an observability platform: its architecture, standards, and operation across an organization.

Join us. The world can't wait.

What You'll Work On:

  • Architect and build the observability platform across every accredited environment, and keep them from drifting into separately maintained estates.

  • Drive OpenTelemetry adoption hands‑on, including collector configuration, instrumentation, and rollout, and set the standards, naming conventions, and onboarding patterns the team builds against.

  • Automate the platform itself, including service onboarding, agent deployment, tagging, dashboard and alert provisioning, and access management, delivered as code through CI/CD.

  • Design and deploy agentic detection, triage, and remediation, defining what an agent may do unattended, how we verify the service actually recovered, how it rolls back, and what evidence it leaves behind.

  • Own the platform's AI capability end to end, including what is enabled, what it is trusted to decide, how its accuracy holds up over time, and how incorrect actions get caught.

  • Build the entity model, dashboards, and SLOs that leadership depends on, and drive down cost per monitored service.

You Have:

  • 8+ years of experience with engineering experience

  • 3+ years of experience maintaining responsibility for an enterprise observability or APM platform itself, including its architecture, deployment, upgrades, standards, and cost across an organization

  • Experience with an enterprise observability platform, such as Dynatrace, Datadog, New Relic, Grafana, Prometheus, or Elastic, including writing and optimizing complex queries against its data

  • Experience replacing recurring manual platform work with production automation, and shipping a system that makes decisions automatically, including agentic operations, ML in production, fraud or risk scoring, or security detection, including owning whether those decisions stayed correct over time

  • Experience with OpenTelemetry, including collectors, instrumentation, semantic conventions, and the practical limits of vendor‑neutral instrumentation

  • Experience with infrastructure as code and CI/CD, such as Terraform, GitHub Actions, Azure DevOps, or GitLab, and deploying agents or collectors at scale across Windows, Linux, Kubernetes, cloud platforms, and network infrastructure

  • Ability to write, review, and ship code

  • Secret clearance

  • Bachelor's degree in CS, Information Systems, or Engineering

Nice If You Have:

  • Experience with Dynatrace in depth, including Grail, DQL, Workflows, Site Reliability Guardian, Davis, and Dashboards, with configuration managed through the Terraform provider

  • Experience building automated remediation in production, including an automation that acted wrongly and what changed afterward

  • Experience with multiple observability platforms, such as Grafana, Prometheus, Datadog, or Elastic

  • Experience operating observability in DoD cloud environments such as Azure or AWS IL5/IL6, including working around feature gaps between commercial and authorized offerings

  • Experience operating one platform across more than one environment or security boundary, and managing the drift and duplication that creates

  • Knowledge of what observability platforms operating models costs to run

  • ITIL 4 Foundation certification

Clearance:

Applicants selected will be subject to a security investigation and may need to meet eligibility requirements for access to classified information; Secret clearance is required.

Compensation

At Booz Allen, we celebrate your contributions, provide you with opportunities and choices, and support your total well-being. Our offerings include health, life, disability, financial, and retirement benefits, as well as paid leave, professional development, tuition assistance, work‑life programs, and dependent care. Our recognition awards program acknowledges employees for exceptional performance and superior demonstration of our values. Full‑time and part‑time employees working at least 20 hours a week on a regular basis are eligible to participate in Booz Allen’s benefit programs. Individuals that do not meet the threshold are only eligible for select offerings, not inclusive of health benefits. We encourage you to learn more about our total benefits by visiting the Resource page on our Careers site and reviewing Our Employee Benefits page.

Salary at Booz Allen is determined by various factors, including but not limited to location, the individual’s particular combination of education, knowledge, skills, competencies, and experience, as well as contract‑specific affordability and organizational requirements. The projected compensation range for this position is $86,800.00 to $198,000.00 (annualized USD). The estimate displayed represents the typical salary range for this position and is just one component of Booz Allen’s total compensation package for employees. This posting will close within 90 days from the Posting Date.

Identity Statement

As part of the hiring process, we will ask you to complete an identity verification process that leverages advanced biometrics and artificial intelligence to ensure authenticity and protect against identity fraud. You are expected to be on camera during interviews and assessments. We reserve the right to take your picture to verify your identity and prevent fraud.

Candidate AI Usage Policy

AI is a part of our daily work at Booz Allen, and we are committed to the responsible and ethical use of AI tools. However, we want to ensure a fair candidate process based on your own skills and knowledge. As part of this commitment, the use of artificial intelligence (AI) or other tools to assist with responses during interviews (whether in‑person or virtual) is prohibited unless permission is explicitly provided.

Work Model
Our people‑first culture prioritizes the benefits of collaboration, whether it occurs in person or virtually. To support engagement and effective communication, employees working virtually are generally expected to have their cameras on during meetings.

  • Remote: If this position is listed as remote, there may still be occasions when you are required to work in person at a Booz Allen or customer facility.

  • Hybrid: If this position is listed as hybrid, you will be expected to work from a Booz Allen facility frequently, in alignment with leadership expectations and the needs of the role. You may also be required to work from or visit a customer facility.

  • Onsite: If this position is listed as onsite, work will primarily be performed at a Booz Allen office or customer facility, where employees will collaborate directly with colleagues and customers as required by the role.

Commitment to Non-Discrimination

All qualified applicants will receive consideration for employment without regard to disability, status as a protected veteran or any other status protected by applicable federal, state, local, or international law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform and Data Engineer
Platform and Data Engineer

Booz Allen Hamilton • Honolulu (HI)

On-site
USD 62,000 - 141,000
Cyber Analytics Platform Architect
Cyber Analytics Platform Architect

Booz Allen Hamilton • Fort Meade (MD)

On-site
USD 87,000 - 198,000
Observability Pipeline Engineer
Observability Pipeline Engineer

Lehigh Heavy Forge Corporation • McLean (VA)

Hybrid
USD 87,000 - 198,000
Health benefits
Professional development
Site Reliability Engineer, Lead
Site Reliability Engineer, Lead

Booz Allen Hamilton • Chantilly (VA)

On-site
USD 99,000 - 225,000
Technical Data Engineer, Senior
Technical Data Engineer, Senior

Booz Allen Hamilton • Herndon (VA)

On-site
USD 99,000 - 225,000
DevOps Infrastructure Engineer
DevOps Infrastructure Engineer

Booz Allen Hamilton • Nellis Air Force Base Census-Designated Place (NV)

On-site
USD 78,000 - 176,000
DevSecOps Platform Engineer
DevSecOps Platform Engineer

Booz Allen Hamilton • Fayetteville (NC)

On-site
USD 77,500 - 176,000
Data Engineer, Senior
Data Engineer, Senior

Booz Allen Hamilton • Maryland

On-site
USD 78,000 - 176,000
DevOps Infrastructure Engineer
DevOps Infrastructure Engineer

Booz Allen Hamilton • Nevada (IA)

Hybrid
USD 78,000 - 176,000
Operations Analyst
Operations Analyst

Booz Allen Hamilton • Wharton (NJ)

On-site
USD 53,000 - 108,000