A complete application in a minute — tailored resume and cover letter, ready to send.
Flexential Corp. is seeking a Senior Platform Engineer to design, build and operate automated, highly available IT platforms in a remote US role. You will own hands-on tasks across observability, devops, incident and release management, and platform integrations on a Kubernetes-driven stack.
You will mentor engineers, drive engineering standards, and influence the technology roadmap with secure, scalable solutions.
Flexential Corp. is hiring a remote candidate for Senior Platform Engineer. This is a full time position. Work location: USA. The role typically involves technologies such as Kubernetes, Terraform, CI/CD, Python, AWS, GCP. The Senior Platform Engineer is a hands‑on engineering role on a platform development team responsible for building and operating Flexential's IT platforms including observability, devops, ITSM incident and release mgmt, and Integrations technologies. This role develops and manages critical platform subsystems for high availability, operational resiliency, security and scalability utilizing native-AI enablement for all outcomes. This is an individual contributor role with significant technical ownership and direct impact on critical Flexential technology roadmap. You will work across infrastructure, automation, and application layers — deploying Kubernetes workloads, authoring Terraform modules, building Ansible playbooks, and building GitLab pipelines that other engineers depend on daily.
Design, develop and operationally manage automated, resilient, high availability, self‑healing, secure platforms with native‑AI capabilities for IT needs, serving both internal as well as customer business capabilities.
Develop, and manage the Observability OpenTelemetry Central Backend Stack: Grafana Enterprise, Mimir, Loki, Tempo, and Alertmanager on Kubernetes/RKE2 via Helm and GitLab CI-CD.
Build and manage iaC and CI-CD for automated provisiong and deployment, including Terraform modules for Infra/VM/storage provisioning, Ansible AWX playbooks for OS/App bootstrap, ArgoCD and Helm for Kubernetes configuration.
Develop and manage OpenTelemetry Prometheus scrape profile library including SNMP exporters, REST API exporters, and cloud provider exporters (CloudWatch, Azure Monitor, GCP) for multiple device classes.
Develop AIOps capabilities on platforms for e.g Observability use‑cases: anomaly detection integrations, event correlation rules in Alertmanager, and synthetic monitoring patterns to reduce alert noise.
Configure and maintain Zabbix auto‑discovery: network range scanning, device classification, and Prometheus service discovery integration.
Build and harden Edge Stack deployments (Prometheus + OTel collector) per data center site using GitOps templates.
Integrate Alertmanager with ServiceNow: webhook routing, ticket enrichment, auto‑close logic, and escalation policy configuration.
Maintain platform security: Conjur/CyberArk secret injection at runtime, mTLS between stack components, RBAC in Grafana Enterprise.
Author and maintain Grafana dashboards in JSON/GitLab — facility overview, network health, RED metrics, application telemetry.
Mentor mid‑level engineers, lead code reviews, and establish engineering standards for the team.
Represent platform engineering in cross‑functional architecture reviews and executive‑level program updates.
Perform other duties as required and assigned.
Experience and/or knowledge of ITSM processes and workflow automation e.g Incident & Response Mgmt (IRM), Release mgmt, ServiceNow ITSM integration, alert routing, escalation policy design, SLA-driven on‑call workflows. Hand‑on's experience or working knowledge of Boomi integrations PaaS(iPaaS) technologies. Experience working with BAS / BMS systems in a Datacenter / OT environment. Hands‑on experience working with AWS products in a Well‑architected Framework and multi‑account model to develop various compute, storage, network iaaaS and PaaS services for IT applications.
Base Pay Range Annualized/Hourly salary range offered for this position is estimated to be $150,000 - $165,000 . However, the actual pay range depends on each candidate’s experience, location, and qualifications.
Benefits are subject to change at the Company's discretion.
At this time, we are only able to employ individuals who reside and perform work in states where the company is authorized to do business and maintain employment registrations. We are currently unable to consider applicants who reside or will perform work in the following states: Arkansas, Alaska, Alabama, Delaware, Hawaii, Iowa, Kansas, Louisiana, Maine, Maryland, North Dakota, Rhode Island, South Dakota, Wisconsin, Wyoming, Washington DC. Applicants must be legally authorized to work in the United States and must reside in a state where the company maintains employment eligibility at the time of hire and throughout their employment. State eligibility may be subject to change based on business needs and regulatory requirements. Flexential participates in the E‑Verify program.
Flexential is an equal opportunity employer, and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity or expression, pregnancy, age, national origin, disability status, genetic information, protected veteran status, or any other characteristic protected by law.
Skills required for this role include Kubernetes, Terraform, CI/CD, and related tools for day-to-day development. Seniority level: Senior.
At Flexential, we value diversity and believe that different perspectives and unique skills make us stronger as a team.
While we've outlined the qualities and skills we are looking for in this job posting, we understand that there may be attributes and abilities that aren't explicitly listed but could be a perfect fit for our team.
We're enthusiastic about assembling a diverse team of individuals who bring various talents to the table.
Your unique perspective and abilities may be just what we need to drive innovation and success.