Platform Engineer, Model Shaping

Together AI

San Francisco (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive health insurance plans
401(k) plan
Flexible time off policy
Monthly team lunches

Job summary

Together AI is seeking a Platform Engineer to design and build the foundational layers of the platform for model customization and evaluation. The role involves collaboration with various engineering teams and requires expertise in infrastructure automation and software engineering.

You will participate in improving reliability, creating internal tooling, and managing complex deployments across a hybrid cloud environment. Ideal candidates will have experience in Python or Go along with strong communication skills.

Qualifications

  • 3+ years of experience in building infrastructure or backend components.
  • Experience with large-scale production systems with high reliability requirements.

Responsibilities

  • Design the infrastructure and backend services for model customization.
  • Collaborate with other engineering teams to integrate services.
  • Contribute to reliability improvements and participate in on-call rotation.
  • Build a job orchestration platform across multiple data centers.

Skills

Infrastructure automation tools (Terraform, Ansible)
Monitoring/observability stacks (Prometheus, Grafana)
CI/CD pipelines (GitHub Actions, ArgoCD)
Python
Go
Linux environments
Container/orchestration stacks (Docker, Kubernetes)
Cloud environment administration (AWS/GCP/Azure)

Job description

  • As a Platform Engineer at Model Shaping, you will work on the foundational layers of Together’s platform for model customization and evaluation
  • You will design the infrastructure and backend services that will allow us to sustainably and reliably scale the systems powering production workflows launched by our users, as well as internal research experiments
  • You will operate in a cross‑functional environment, collaborating with other engineers and researchers in the team to improve the infrastructure based on the needs of projects they work on
  • You will also interact with other engineering teams at Together (such as Commerce, Data Engineering, and Cloud Infrastructure) to integrate the services developed by Model Shaping with systems developed by those teams
  • Design and build Together’s systems and infrastructure for model customization, including user‑facing features and internal improvements
  • Contribute to reliability improvements for the platform, participating in an on‑call rotation and improving processes for incident response
  • Create and improve internal tooling for deployment, continuous integration, and observability
  • Build a job orchestration platform spanning multiple data centers, supporting a highly heterogeneous hardware landscape
  • Partner with teams developing internal services, co‑designing these services and incorporating them in systems built by Model Shaping
Benefits
  • Competitive health insurance plans
  • Pre‑tax flexible spending accounts
  • Dental and vision insurance
  • Income protection & retirement
  • Mental health support and services
  • Life insurance
  • AD&D insurance
  • 401(k) plan
  • STD & LTD insurance
  • Monthly commuting stipend + pre‑tax bene
  • Flexible time off policy
  • Monthly team lunches
  • Team‑driven celebrations and events
Qualifications
  • Experienced with infrastructure automation tools (Terraform, Ansible), monitoring/observability stacks (Prometheus, Grafana), and CI/CD pipelines (GitHub Actions, ArgoCD)
  • 3+ years of experience in building infrastructure or backend components of production services
  • Skilled with analyzing non‑trivial issues of complex software systems and documenting your findings
  • Strong software engineering background in Python or Go
  • Strong communication skills, willing to document systems and processes and collaborate with peers of varying technical expertise
  • Have cloud environment (e.g., AWS/GCP/Azure) administration experience, preferably with a hybrid bare‑metal/cloud environment
  • Comfortable with the fundamentals of Linux environments and modern container/orchestration stacks (e.g., Docker and Kubernetes)
  • Developing large‑scale production systems with high reliability requirements
  • Pipeline orchestration frameworks (e.g., Kubeflow, Argo Workflows, Flyte)
  • Deployment of services for AI training or inference
  • Managing GPU workloads on HPC clusters, ideally with hands‑on experience in operating NVIDIA’s networking stack (e.g., NCCL, Mellanox firmware, GPUDirect RDMA)
  • Maintaining or contributing to open‑source projects
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Engineer, Model Shaping
Platform Engineer, Model Shaping

Togetherai • San Francisco (CA)

Hybrid
USD 200,000 - 290,000
Competitive compensation
Startup equity
Health insurance
+1
Platform Engineer, Model Shaping
Platform Engineer, Model Shaping

Together AI • San Francisco (CA)

Hybrid
USD 200,000 - 290,000
Startup equity
Health insurance
Flexible remote work
Staff Platform Engineer, Service Infrastructure
Staff Platform Engineer, Service Infrastructure

Togetherai • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance
Competitive compensation
Staff Platform Engineer, Service Infrastructure
Staff Platform Engineer, Service Infrastructure

AI Chopping Block • San Francisco (CA)

On-site
USD 240,000 - 280,000
Health insurance
Startup equity
Comprehensive benefits
Research Engineer, Large-Scale Training
Research Engineer, Large-Scale Training

Together AI • San Francisco (CA)

On-site
USD 200,000 - 290,000
Health insurance
Startup equity
Other benefits
Staff Platform Engineer, Service Infrastructure
Staff Platform Engineer, Service Infrastructure

Together AI • San Francisco (CA)

On-site
USD 240,000 - 280,000
Research Engineer, Large-Scale Training
Research Engineer, Large-Scale Training

Togetherai • San Francisco (CA)

On-site
USD 200,000 - 290,000
Competitive compensation
Startup equity
Health insurance
+1
Senior Software Engineer - Together Cloud Platform
Senior Software Engineer - Together Cloud Platform

Togetherai • San Francisco (CA)

Hybrid
USD 160,000 - 230,000
Startup equity
Health insurance
Flexible remote work options
Senior Software Engineer - Together Cloud Platform
Senior Software Engineer - Together Cloud Platform

Together • San Francisco (CA)

Hybrid
USD 160,000 - 230,000
Equity
Health insurance
Remote work flexibility
Senior Software Engineer - Together Cloud Infrastructure
Senior Software Engineer - Together Cloud Infrastructure

Togetherai • San Francisco (CA)

Hybrid
USD 160,000 - 230,000
Equity
Health insurance
Flexible remote work options