Role: Computer Vision Engineer
Type of Employment : Full time with Zillion Technologies
Location: Ashburn, VA or Bethesda MD (Hybrid)
The Computer Vision Engineer role owns all computer vision engineering effort. You will work on edge-deployed CV pipelines running on NVIDIA Jetson hardware, and GPU servers - building, training, optimizing, and deploying models that handle real-world conditions in public and commercial spaces.
You will build systems with person and intent detection, multi-camera tracking, track package placement and removal events at shelf zones. You will also build model training and deployment pipelines and perform edge deployment and performance optimization.
Required
- 5+ years of hands-on computer vision engineering experience, with at least 2 years deploying models to production edge hardware (not just cloud or research environments)
- Deep practical experience with the YOLO family of detectors - training, fine-tuning, hyperparameter tuning, and understanding failure modes in real-world conditions
- Proficiency with PyTorch for model training and ONNX / TensorRT for inference optimization; hands-on experience with INT8 or FP16 post-training quantization
- Experience building multi-object tracking pipelines - SORT, DeepSORT, BoT-SORT, or equivalent - and understanding the tradeoffs between tracker accuracy, computational cost, and track stability
- Solid Python and C++ skills for pipeline development; comfort reading and modifying GStreamer pipeline graphs
- Experience with NVIDIA GPU tooling: CUDA, cuDNN, TensorRT, and the JetPack / Jetson SDK ecosystem
- Experience building annotation pipelines and managing training datasets for custom object detection tasks - not just using pre-trained models on standard benchmarks
- Comfort working with RTSP IP camera streams in Linux environments; understanding of H.264/H.265 codec pipeline and hardware decode
- Experience with cross-camera person re-identification - OSNet, FastReID, or equivalent architectures; homography-based multi-camera fusion
- Experience with zone-based spatial analytics - polygon intersection, floor-plane projection, homography calibration from camera to world coordinates
Strongly preferred
- Experience building CV systems for retail, logistics, or security environments where the camera network covers a physical space and detections must be spatially anchored
- Familiarity with Roboflow or CVAT for dataset management and annotation workflow automation
- Experience with the NVIDIA Metropolis or DeepStream framework - even if ultimately not used, understanding where these add value vs a custom open-source stack
- Prior work on privacy-preserving CV pipelines - on-device inference, derived-data-only architectures, anonymization techniques