Staff PlatformOps Engineer (Emerging AI)

Socket.dev

Seattle (WA)

Hybrid

USD 198,000 - 246,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Medical, Dental, Vision
Parental Leave
Flexible Vacation
401k matching
Transit benefits

Job summary

DAT Freight & Analytics is seeking a Staff AI Platform Engineer to own the platform powering all AI workloads. You will shape architecture from gateway to storage, and lead the design and governance of scalable, observable AI infrastructure on AWS.

You will mentor engineers and data scientists, drive best practices in CI/CD, testing, and incident response, and partner with Product and Engineering across the company to ship reliable AI services.

Qualifications

  • 10+ years of software engineering experience with production infrastructure ownership.
  • 4+ years building and operating ML/AI platforms in production.
  • Proficiency in TypeScript/Node.js, Java, or Go, with library design mindset.
  • Strong AWS production experience incl. EKS/Kubernetes, Lambda, S3, Kafka/MSK.
  • Experience with containerized inference, GPU scheduling, autoscaling, batching.
  • Hands-on with LLM systems in production: retrieval-augmented generation, vector DBs, tool-calling.
  • Experience building evaluation/monitoring: offline/online metrics, drift checks.
  • Background in event-driven distributed systems using Redpanda or Kafka.

Responsibilities

  • Own the platform design and operations for AI workloads across teams.
  • Lead architecture discussions and align Product, Eng, and Business stakeholders.
  • Architect scalable AI infrastructure on AWS and define IaC with Pulumi.
  • Build offline/online evaluation systems and A/B/shadow testing pipelines.
  • Own cost and performance metrics, optimize GPU/serverless capacity and latency.
  • Implement guardrails, logging, data governance, and security across models.
  • Promote modern software practices: CI/CD, automated testing, code reviews.
  • Mentor engineers in production ML and distributed systems.
  • Lead incident response and root-cause analysis for AI services.

Skills

Platform ownership
AWS expertise
LLM systems
Observability
Mentoring
Communication

Education

Bachelor's degree in CS/Engineering

Tools

Pulumi
Terraform
Kubernetes
S3
Redpanda
Kafka

Job description

About DAT

DAT Freight & Analytics is an award-winning employer of choice and a next-generation SaaS technology company that has been at the leading edge of freight and logistics innovation for nearly five decades. Founded in 1978, DAT operates the largest freight marketplace in North America — processing 250 million+ load posts annually and maintaining one of the largest repositories of freight market transaction data in the world. On a defined path to $1 billion in revenue, DAT deploys a suite of software solutions, machine learning models, and intelligent automation tools that help brokers, carriers, and shippers price freight accurately, source capacity, reduce risk, and operate more efficiently. With nearly 700 teammates across offices in Denver, CO; Portland, OR; Seattle, WA; Springfield, MO; Toronto, ON; and Bangalore, India, DAT combines the credibility of a multi-decade market leader with the drive of a company that is not done disrupting the industry it helped build.

For more information, visit www.DAT.com

Job Application Deadline: 10/31/2026

The Opportunity

As a Staff AI Platform Engineer, you'll build and own the platform that every AI and machine learning workload at DAT runs on. Freight is an uncertain business, and the models we ship reduce that uncertainty: rate forecasts, load-to-truck matching, document extraction, fraud signals, and the agentic workflows our brokers and carriers use to move freight faster. None of that reaches a customer without a platform that makes training, serving, evaluating, and monitoring models routine instead of heroic.

You’ll set the architectural direction for that platform, from the model gateway and inference layer through feature and vector storage, evaluation harnesses, and production observability. This is a highly visible role for an engineer who wants their work multiplied across every AI team in the company.

What You'll Do
  • Platform Ownership: Design, build, and operate the shared services that engineers and business users use to ship models: a model gateway for LLM access, inference endpoints for real-time and batch scoring, feature storage, vector search, and a common SDK.
  • Technical Leadership: Lead architecture for large-scale AI systems, write the design documents, and drive alignment across Product, and Engineering and Business users on how models get built and shipped at DAT.
  • Cloud Architecture: Architect and run scalable, reliable AI infrastructure on AWS, including Bedrock, SageMaker, EKS, Redpanda, MSK (Kafka), Lambda, and S3, all defined in Pulumi.
  • Evaluation and Quality: Build the offline and online evaluation systems that tell us whether a model or prompt change is an improvement, including regression suites, LLM-as-judge pipelines, A/B and shadow testing, and drift detection.
  • Cost and Performance: Own inference cost and latency as first-class metrics. Right-size GPU and serverless capacity, tune batching, caching, and quantization, and give teams clear visibility into what their workloads cost.
  • Safety and Governance: Implement guardrails, prompt and output logging, PII handling, access controls, and model and dataset lineage so AI systems meet our security and customer data commitments.
  • Best Practices: Drive the adoption of modern software development practices across AI work, including automated testing, code reviews, CI/CD pipelines, and infrastructure-as-code.
  • Mentorship: Mentor engineers and data scientists on production ML and distributed systems, and raise the operational bar of every team that builds on the platform.
  • Incident Management: Lead the response and resolution for complex production incidents involving AI services, perform root cause analysis, and implement preventative measures.
The Skills and Experience You'll Bring
  • 10+ years of experience in software engineering, including significant time as a senior or staff engineer owning production infrastructure that other engineering teams depend on.
  • 4+ years building and operating machine learning or AI platforms, with models you've taken from notebook to production traffic and then kept healthy.
  • Strong development language skills ideally using TypeScript/Node.js, Java, or Go, with the software engineering discipline to build libraries other teams adopt willingly.
  • Extensive AWS experience with production systems, ideally including EKS/Kubernetes, Lambda, Redpanda, MSK/Kafka, S3, and Secrets Manager, plus infrastructure-as-code with Pulumi, Terraform or similar.
  • Hands-on experience serving models at scale: containerized inference, GPU scheduling, autoscaling, batching, and the tradeoffs between real-time, streaming, and batch scoring.
  • Practical experience with LLM systems in production, including retrieval-augmented generation, vector databases, prompt and context management, tool-calling or agent frameworks, and managed model APIs such as Amazon Bedrock.
  • Proven experience building evaluation and monitoring for models, not just services: offline eval sets, online metrics, drift and data quality checks, and a clear definition of what "regression" means for a model.
  • Strong background in event-driven and distributed systems using Redpanda, AWS MSK or Kafka
  • Fluency with observability and on-call health, designing dashboards, APM traces, logs, and alerts, defining SLOs, and using these tools to drive down incident frequency and MTTR.
  • Experience establishing and raising engineering standards for a team: PR and testing guidelines, deployment validation checklists, model release criteria, and post-incident review practices.
  • Demonstrated technical leadership: leading cross-team designs in ambiguous problem spaces, writing clear system design documents and operations runbooks, and driving alignment across Engineering, Product, and Operations.
  • Proven mentoring track record, especially helping engineers and data scientists grow in system design, observability, and operational excellence.
  • Excellent communication skills, with the ability to explain trade-offs and system behavior clearly to engineers, product managers, and non-technical stakeholders.
We'd be Extra Excited if You Have
  • Experience with model fine-tuning, distillation, or parameter-efficient training, and a clear point of view on when it beats prompting.
  • Background in marketplace, pricing, forecasting, or document extraction problems.
  • Freight, logistics, or supply chain domain experience.
Why DAT?

DAT is an award winning employer of choice.

For starters, we have a hybrid work environment, but we also know what makes a great workplace. We have a time-tested and resolute set of operating values predicated on integrity, mutual respect, open communication, and executing with excellence. These values inform our strategic vision as much as any one of our products does. We’ve been an employer of choice in the Portland metropolitan area for four decades, and within one year of opening our Denver office, DAT was #26 on Built In Colorado’s 100 Best Places to Work In Colorado.

  • Medical, Dental, Vision, Life, and AD&D insurance
  • Parental Leave
  • Flexible Vacation Time (FVT)
  • An additional 10 holidays of paid time off per calendar year
  • 401k matching (immediately vested)Employee Stock Purchase Plan
  • Short- and Long-term disability sick leave
  • Flexible Spending Accounts
  • Health Savings Accounts
  • Employee Assistance Program
  • Additional programs - Employee Referral, Internal Recognition, and Wellness
  • Free TriMet transit pass (Beaverton Office)
  • Competitive salary and benefits package
  • Work on impactful projects in a cutting-edge environment
  • Collaborative and supportive team culture
  • Opportunity to make a real difference in the trucking industry
  • Employee Resource Groups

For Washington-based candidates, in compliance with the Washington State Pay Transparency Law, the salary range for this role is $198,000.00 - $246,000.00 + target bonus. DAT considers factors such as scope and responsibilities of the position, candidate's work experience, education and training, core skills, internal equity, and market and business elements when extending an offer.

DAT embraces the value of a diverse workforce, and believes it is a core strength of our company that we encourage those values in every DAT employee, at every level of our organization, regardless of tenure or rank. We provide equal employment opportunities (EEO) to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, disability, genetic information, marital status, amnesty, or status as a covered veteran in accordance with applicable federal, state, and local laws.

Equal Opportunity Employer/Protected Veterans/Individuals with Disabilities

The contractor will not discharge or in any other manner discriminate against employees or applicants because they have inquired about, discussed, or disclosed their own pay or the pay of another employee or applicant. However, employees who have access to the compensation information of other employees or applicants as a part of their essential job functions cannot disclose the pay of other employees or applicants to individuals who do not otherwise have access to compensation information, unless the disclosure is (a) in response to a formal complaint or charge, (b) in furtherance of an investigation, proceeding, hearing, or action, including an investigation conducted by the employer, or (c) consistent with the contractor’s legal duty to furnish information. 41 CFR 60-1.35(c)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Business Operations Associate
Senior Business Operations Associate

DAT Freight & Analytics • Seattle (WA)

Hybrid
USD 141,000 - 190,000
Medical, Dental, Vision
401k matching
Hybrid work environment
Engineering Manager
Engineering Manager

DAT Freight & Analytics • Portland (OR)

Hybrid
USD 192,000 - 261,000
Medical, Dental, Vision, Life, and AD&
Parental Leave
Flexible Vacation Time
+4
Staff Data Engineer
Staff Data Engineer

DAT Freight & Analytics • Denver (CO)

Hybrid
USD 134,000 - 175,000
Medical, Dental, Vision, Life, AD&D
Parental Leave
Flexible Vacation Time
+8
Engineering Manager
Engineering Manager

DAT Freight Solutions • Seattle (WA)

On-site
USD 192,000 - 261,000
Medical, Dental, Vision
Parental Leave
Flexible Vacation Time
+5
Staff Data Engineer
Staff Data Engineer

DAT Freight Solutions • Denver (CO)

Hybrid
USD 134,000 - 175,000
Hybrid work environment
401k matching
Health, Dental, Vision
+2
Software Engineer II
Software Engineer II

DAT Freight & Analytics • Denver (CO)

Hybrid
USD 113,000 - 144,000
Medical, Dental, Vision, Life
Parental Leave
Up to 20 days PTO + 10 holidays
+3
Senior Engineering Manager, Fintech
Senior Engineering Manager, Fintech

DAT Freight Solutions • Seattle (WA)

Hybrid
USD 237,000 - 297,000
Medical/Dental/Vision
401(k) matching
Paid time off
+2
Senior Software Engineer, Frontend
Senior Software Engineer, Frontend

DAT Freight & Analytics • Denver (CO)

Hybrid
USD 143,000 - 183,000
Medical, Dental, Vision insurance
401(k) matching
Up to 20 days of paid time off
+1
Site Reliability Engineer II
Site Reliability Engineer II

Worky • Denver (CO)

Hybrid
USD 95,000 - 134,000
Medical, Dental, Vision
401k matching
Employee Stock Purchase Plan
+3
Staff Software Engineer, Salesforce
Staff Software Engineer, Salesforce

DAT Freight & Analytics • Seattle (WA)

Hybrid
USD 195,000 - 230,000
Medical Insurance
Dental Insurance
Vision Insurance
+10