Senior Computer Vision Engineer (Egocentric), Data Foundry

Stord

United States

On-site

USD 140,000 - 210,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Stord is building a new physical AI data business from the ground up and seeks an experienced Computer Vision Engineer. You will own the data product, capture operations, and perception stack, working closely with leadership to shape the technical and commercial direction of the venture.

You will collect real-world warehouse data, develop scalable CV pipelines, and deploy models that perform reliably in production, not just in notebooks. A builder mindset is essential.

Qualifications

  • 8+ years building and shipping production computer vision systems.
  • Experience building and scaling an egocentric perception or video data stack end to end.
  • Deep expertise in detection, tracking, segmentation, depth, 2D/3D pose estimation, and vision transformers.
  • Strong command of geometric computer vision, including camera calibration, multi-view geometry, synchronization, and 3D reconstruction.
  • Proven ownership of complex perception problems from data and model design through evaluation, optimization, and deployment.
  • Experience with large, unstructured video and multimodal datasets.
  • Track record of setting technical direction and raising the bar for engineers.
  • Ability to take ambiguous 0→1 problems from concept to working system with limited resources and no playbook.
  • Expert-level Python and strong software engineering fundamentals; C++ where performance demands it.
  • Builder mindset across hardware, data, models, infrastructure, and operations.

Responsibilities

  • Define and build the data product. Own the product across quality tiers—from RGB egocentric video to depth-enhanced and multimodal capture with hand/body pose.
  • Stand up the capture operation with camera rigs, hardware setup, enrollment, edge processing, and data pipelines.
  • Build the perception stack: detection, tracking, segmentation, depth/3D reconstruction, 6DoF, and multi-view pose estimation.
  • Automate labeling with VLM-assisted workflows and human-in-the-loop QA.
  • Own the hardware-vision intersection: calibration, multi-view geometry, synchronization, and 3D pose triangulation.
  • Train and ship models: design, fine-tune, and deploy CV/Multimodal models on large datasets.

Skills

Python
C++
Computer vision
Machine learning
Deep learning
Robotics
Data pipelines
Software engineering

Education

MS/PhD in computer vision, ML, or robotics

Job description

Stord is The Consumer Experience Company, powering seamless checkout through delivery for today's leading brands. Stord is rapidly growing and is on track to double our revenue in the next 18 months. To meet and exceed this target, Stord is strategically scaling teams across the entire company, and seeking energetic experts to help us achieve our mission.

By combining comprehensive commerce-enablement technology with high-volume fulfillment services, Stord provides brands a platform to compete with retail giants. Stord manages over $10 billion of commerce annually through its fulfillment, warehousing, transportation, and operator-built software suite including OMS, Pre- and Post-Purchase, and WMS platforms. Stord is leveling the playing field for all brands to deliver the best consumer experience at scale.

With Stord, brands can increase cart conversion, improve unit economics, and drive sustained customer loyalty. Stord’s end-to-end commerce solutions combine best-in-class omnichannel fulfillment and shipping with leading technology to ensure fast shipping, reliable delivery promises, easy access to more channels, and improved margins on every order.

Hundreds of leading DTC and B2B companies like AG1, True Classic, Native, Seed Health, quip, goodr, Sundays for Dogs, and more trust Stord to deliver industry-leading consumer experiences on every order. Stord is headquartered in Atlanta with facilities across the United States, Canada, and Europe. Stord is backed by top-tier investors including Kleiner Perkins, Franklin Templeton, Founders Fund, Strike Capital, Baillie Gifford, and Salesforce Ventures.

Stord operates one of the largest independent e-commerce fulfillment networks in the U.S., with 20+ fulfillment centers, 4,000+ warehouse associates, and nearly 100 million packages shipped annually.

We’re building a new business line that transforms this real-world operational infrastructure into high-value training data for the next generation of physical AI.

We’re looking for an experienced Computer Vision Engineer and technologist to help build and scale this business from the ground up. You’ll work at the intersection of computer vision, robotics, data, and warehouse operations, turning real-world environments and workflows into high-quality datasets that enable smarter, more capable AI systems.

Why This Role

This is an opportunity to build a new physical AI data business from the ground up, with access to an operating environment that would be difficult to replicate anywhere else.

You’ll Have
  • A structural data advantage. Direct access to real-world warehouse environments, workflows, and human activity at significant scale.
  • A massive and rapidly growing market. Build data products for companies developing the next generation of robotics and physical AI.
  • True 0→1 ownership. Shape the product, technology, team, and operating model from the beginning.
  • Direct partnership with the CTO & Co-Founder. Work closely with company leadership to define the technical and commercial direction of the business.
What You’ll Do

You will own the early egocentric video and perception stack—from data collection and camera rigs through vision models, processing pipelines, and dataset delivery. This is a hands‑on builder‑operator role: you’ll define what needs to be built, build the critical pieces yourself, and work with a small team to operationalize and scale them.

You Will
  • Define and build the data product. Own the product across quality tiers—from RGB egocentric video to depth‑enhanced and multimodal capture with hand pose, body pose, and annotations. Prioritize what gets built based on customer demand and hold a high bar for data quality.
  • Stand up the capture operation. Own camera and rig selection, hardware setup, enrollment, edge processing, data ingestion, and the pipelines that turn raw capture into production‑ready datasets. Partner closely with warehouse operations, engineering, and customers to deliver on spec and on schedule.
  • Build the perception stack. Develop detection, tracking, segmentation, depth/3D reconstruction, 6DoF, and multi‑view 3D hand/body pose estimation across egocentric and fixed‑camera systems.
  • Automate labeling at scale. Build VLM‑assisted and automated labeling workflows with human‑in‑the‑loop QA, reducing the cost of annotation while maintaining rigorous quality standards.
  • Own the hardware‑vision intersection. Drive camera calibration, epipolar and multi‑view geometry, frame‑accurate synchronization, and 3D pose triangulation across multi‑camera and egocentric rigs.
  • Train and ship models. Design, fine‑tune, optimize, and deploy computer vision and multimodal models against large, unstructured video datasets. Build reproducible systems that perform reliably in production—not models that live in a notebook.
Basic Qualifications
  • 8+ years building and shipping production computer vision/perception systems, or an MS/PhD in computer vision, ML, or robotics with 6+ years of hands‑on industry experience.
  • Experience building and scaling an egocentric perception or video data stack end to end, ideally within robotics, physical AI, or an AI data company.
  • Deep expertise in computer vision, including detection, tracking, segmentation, depth, 2D/3D pose estimation, and vision transformers.
  • Strong command of geometric computer vision, including camera calibration, multi‑view geometry, synchronization, and 3D reconstruction.
  • Proven ownership of complex perception problems from data and model design through evaluation, optimization, and deployment, with measurable improvements in accuracy and reliability.
  • Experience working with large, unstructured video and multimodal datasets, with a disciplined approach to evaluation and quality measurement.
  • A track record of setting technical direction and raising the bar for other engineers.
  • Proven ability to take ambiguous 0→1 problems from concept to working system with limited resources and no established playbook.
  • Expert‑level Python and strong software engineering fundamentals; C++ experience where performance demands it.
  • Most importantly, you’re a builder. You’re comfortable moving between hardware, data, models, infrastructure, and operations to make the system work.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Computer Vision Engineer (Egocentric), Data Foundry Software
Senior Computer Vision Engineer (Egocentric), Data Foundry Software

Front Door Defense • Atlanta (GA)

On-site
USD 160,000 - 230,000
Senior Computer Vision Engineer (Egocentric), Data Foundry
Senior Computer Vision Engineer (Egocentric), Data Foundry

Stord-Warehous • United States

Remote
USD 140,000 - 210,000
Staff Technical Program Manager - Vision AI
Staff Technical Program Manager - Vision AI

Stord • United States

On-site
USD 150,000 - 210,000
Staff Technical Program Manager
Staff Technical Program Manager

Stord • United States

On-site
USD 180,000 - 240,000
Senior Computer Vision Engineer — Build Ground‑Up AI Data
Senior Computer Vision Engineer — Build Ground‑Up AI Data

Stord • United States

On-site
USD 140,000 - 210,000
Senior Egocentric Vision Engineer - Remote
Senior Egocentric Vision Engineer - Remote

Stord • Atlanta (GA)

Remote
USD 170,000 - 210,000
Senior Computer Vision Engineer - Build Perception Stack
Senior Computer Vision Engineer - Build Perception Stack

Stord-Warehous • Atlanta (GA)

On-site
USD 140,000 - 210,000
Senior Computer Vision/Machine Learning Engineer
Senior Computer Vision/Machine Learning Engineer

One Way Ventures • Mountain View (CA)

Hybrid
USD 120,000 - 150,000
Staff Data Scientist
Staff Data Scientist

B Capital • United States

Hybrid
USD 130,000 - 160,000
Sr. Computer Vision / Machine Learning Engineer
Sr. Computer Vision / Machine Learning Engineer

Corvus Robotics, Inc. • Mountain View (CA)

Hybrid
USD 120,000 - 160,000