Training data for robots that learn from humans

Real-world manipulation demonstrations — towel folding, t-shirt folding, envelope insertion, assembly — captured in real homes and annotated for Vision-Language-Action model training.

4,000+
Episodes
Up to 4K · 60fps
Capture
5
Task Categories
Annotated towel-folding demonstration — bounding boxes over hands and towel
EP-0001 · 4K
Real demonstrations

Every dataset is footage you can actually watch.

Everyday manipulation tasks — folding, insertion, assembly — captured top-down in real homes and labelled with bounding boxes for VLA training. Pick a clip to watch it play.

Towel Folding

1080p · 24fps · bounding boxes

View dataset

Every clip previews a real dataset — full episodes live in the catalog.

Task categories

Simple, everyday manipulation tasks.

We focus on a small, growing set of household manipulation tasks — captured cleanly and labelled for VLA training.

  • Towel Folding — annotated sample frame

    Towel Folding

    800eps

    Folding towels of varied size and color on a flat surface.

  • T-Shirt Folding — annotated sample frame

    T-Shirt Folding

    650eps

    Cotton t-shirt folding, captured top-down from a flat start.

  • Envelope Insertion — annotated sample frame

    Envelope Insertion

    500eps

    Precision insertion — aligning and placing cards into envelopes.

How the data is made

One thread, from a real home to your training run.

Four deliberate steps turn raw moments of everyday life into annotation-ready episodes — captured, labelled, checked, and shipped in the format you train on.

  1. 01

    Capture

    Everyday household tasks recorded in real homes — never staged labs.

    GoPro · up to 4K·60fps

  2. 02

    Annotate

    Each episode labelled with bounding boxes and natural-language task descriptions.

    bounding boxes + language

  3. 03

    Validate

    Automated quality checks plus human review before anything leaves the pipeline.

    QC + human review

  4. 04

    Deliver

    Shipped annotation-ready for Vision-Language-Action models, in the format you train on.

    MP4 · RLDS · HDF5 · ROSBag

1,000+ episodes·3 task categories·Human-validated

Browse Datasets
Why Yoata

Real homes, validated by hand, at a price that scales — not staged lab data.

Real Environments

Captured in real homes — never staged labs. The lighting, clutter and variation your policies actually face.

Validated Pipeline

Every episode passes automated checks for occlusion and motion blur, then a manual annotation audit — before anything ships.

Cost-Effective

Real-world manipulation data at a fraction of lab-collection cost — collected efficiently, in-region, at scale.

Ready to put real-world data to work?

Order a custom collection or start from a ready-made dataset — captured in real homes, delivered VLA-ready.

Annotated assembly demonstration — hands fastening a bolt and nut