Platform Reliability Lead

Compunnel, Inc.

Berkeley Heights (NJ)

On-site

USD 110,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Compunnel, Inc. is seeking an OMS Platform Reliability Lead to oversee the health and stability of the Fluent Commerce Order Management ecosystem. This role will lead the technical RUN support team and focus on transitioning operations to a 'Self-Healing' model using automated solutions.

The ideal candidate will have at least 5 years of experience in OMS Technical Operations, advanced knowledge of Fluent Commerce, and be proficient in Java and GraphQL. Strong leadership skills are essential for mentoring the support team and collaborating with development teams.

Qualifications

  • 5+ years in OMS Technical Operations or Platform Engineering.
  • Advanced technical knowledge of Fluent Commerce, Webhooks, and Fluent GraphQL API.
  • Experience with ITIL and SRE principles.

Responsibilities

  • Manage the technical RUN support team and implement automation for 'Self‑Healing' operations.
  • Design automated Order Replay mechanisms and enhance observability dashboards.
  • Oversee the incident management lifecycle ensuring documentation of technical fixes.

Skills

Java
GraphQL
RESTful architectures
Monitoring tools (Datadog, Splunk)
GIT

Education

Bachelor’s degree in Computer Science, Software Engineering, or related field

Tools

ELK Stack
Fluent Commerce SDK

Job description

JOB SUMMARY

The OMS Platform Reliability Lead is a highly technical role responsible for the health, stability, and automated evolution of the Fluent Commerce Order Management ecosystem. This position leans heavily into Systems Engineering, requiring the ability to read and debug Java extensions, design complex GraphQL mutations, and build automated remediation tools for the "RUN" team. You will manage the technical RUN support team and serve as the bridge between software engineering and IT operations. Your primary focus is to transition from manual support to "Self‑Healing" operations by implementing automation for order replays, data deduplication, and predictive alerting.

Key Responsibilities
Technical Automation & Self-Healing Operations
  • Order Remediation Automation: Design and implement automated "Order Replay" mechanisms within Fluent Commerce to resolve synchronization failures between event-driven integrations without manual intervention.
  • Enhanced Observability: Build advanced telemetry dashboards (using tools like Splunk, Datadog, or New Relic) to monitor GraphQL query performance, API latency, and webhook success rates.
  • Smart Alerting: Design and tune threshold‑based alerting for the RUN team to identify "Stuck Orders" or inventory mismatches before they impact the customer experience.
  • Tooling Development: Script custom utilities using the Fluent Commerce SDK or REST APIs to facilitate bulk updates and system cleanups.
Technical Incident Management & Platform Monitoring
  • Deep-Dive Troubleshooting: Act as the ultimate technical escalation point for incidents requiring code‑level analysis of Java custom extensions or complex GraphQL mutations.
  • Root Cause Engineering: Lead technical Root Cause Analysis (RCA) by performing deep dives into application logs and event‑driven architecture to identify architectural bottlenecks.
  • Performance Tuning: Analyze API response times and database interaction patterns to propose platform optimizations to the development team.
  • ITSM Compliance: Oversee the incident management lifecycle, ensuring documentation includes code‑level workarounds and technical "bug‑fixes" for future reference.
Stakeholder & Vendor Engineering Collaboration
  • Technical Liaison: Serve as the primary technical point of contact for e‑commerce and architecture teams to ensure operational requirements are included in the dev roadmap.
  • Vendor Management: Collaborate with Fluent Commerce product engineers to align on platform upgrades and API versioning impacts.
  • Team Leadership: Mentor the RUN support team in technical skills including GraphQL query optimization and Java debugging.
Change Management & Release Integrity
  • Technical Oversight: Validate technical configurations and platform extensions during the release cycle to ensure deployment integrity and performance stability.
  • CI/CD Awareness: Manage version control using GIT, ensuring proper branching strategies for operational hotfixes and configuration changes.
Required Qualifications
  • Education: Bachelor’s degree in Computer Science, Software Engineering, or a related technical field.
  • Experience: 5+ years in OMS Technical Operations or Platform Engineering, with specific experience in high‑volume, event‑driven SaaS environments.
  • Fluent Commerce Expertise preferred: Advanced technical knowledge of Fluent Commerce (specifically Webhooks, Essential Rules, and the Fluent GraphQL API).
  • Core Technical Stack:
    • Java: Proficiency in reading, debugging, and identifying performance issues in custom Java extensions.
    • GraphQL: Expert proficiency in query/mutation design, including the use of aliases, fragments, and variables for complex data manipulation.
    • Integration: Comprehensive understanding of RESTful architectures, JSON schemas, and event‑driven patterns (Pub/Sub, Kafka, or Event Grid).
    • Observability: Experience with monitoring tools such as Datadog, Splunk, ELK Stack, or New Relic.
    • GIT: Deep experience with repository management and deployment pipelines.
    • Process Knowledge: Strong mastery of ITIL with an SRE (Site Reliability Engineering) mindset—focusing on automation over manual "toil."
    • Analytical Skills: Ability to parse complex system logs and use data to drive proactive stability improvements.
    • Communication: Ability to explain a "race condition" or "API timeout" to a business stakeholder in terms of revenue and customer impact.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

OMS Platform Reliability Lead
OMS Platform Reliability Lead

Veriipro • Berkeley Heights (NJ)

On-site
USD 100,000 - 130,000
Platform Reliability Engineer: Self-Healing & Observability
Platform Reliability Engineer: Self-Healing & Observability

Compunnel, Inc. • Berkeley Heights (NJ)

On-site
USD 110,000 - 150,000
Sr. Product Mgr
Sr. Product Mgr

TechDigital Group • Johns Creek (GA)

On-site
USD 120,000 - 150,000
Staff Software Engineer, Orders Platform
Staff Software Engineer, Orders Platform

United States Digital Space LLC • United States

Remote
USD 150,000 - 230,000
SRE/Java Engineer
SRE/Java Engineer

Compunnel, Inc. • Quincy (MA)

On-site
USD 100,000 - 130,000
Java Developer
Java Developer

SQLI • Georgia

Hybrid
USD 90,000 - 120,000
Senior Product Manager, Order Management Systems
Senior Product Manager, Order Management Systems

BetterCloud • Buffalo (NY)

On-site
USD 130,000 - 170,000
AWS Engineer – OMS
AWS Engineer – OMS

Lumovy Technology Solutions • Seattle (WA)

On-site
USD 110,000 - 130,000
Sr. Product Mgr
Sr. Product Mgr

TechDigital Group • United States

On-site
USD 110,000 - 140,000
Product Operations Support Engineer
Product Operations Support Engineer

TechDigital Group • Nashville (TN)

On-site
USD 80,000 - 100,000