Skip to main content
Hand activity answers one question: how much of this footage shows someone actually doing something? Video can be perfectly sharp, well exposed and completely idle. A hundred hours of a camera-wearer walking between workstations is a hundred hours of undamaged video and very little work. These measurements separate the two. Every percentage on this page is taken over assessed_seconds, so multiplying any of them by that duration turns it into hours of footage.
These fields describe a human camera-wearer’s hands in egocentric footage. They do not apply to robot grippers.

Active manipulation

What it means. The percentage of assessed footage where a visible hand is purposefully acting on a physical object or material — picking something up, turning it, assembling, wiping, cutting. Read this one first. It is the closest single number to how much of the footage is work. A hundred-hour dataset that reads 15% active manipulation contains roughly fifteen hours of activity and eighty-five of walking, waiting and standing around. When it is low: long transit between tasks, a lot of setup and teardown, observation-heavy work, or a camera angle that misses the hands. What it does not mean: that the work was done well. This is a measure of purposeful activity, not of task success, skill, or whether the demonstration achieved anything. What to do. Convert it to hours. Recorded duration and useful duration are different numbers, and this is the one that tells you how much of a dataset is worth annotating, training on, or looking through.

Hand visibility

What it means. How much of the footage shows no hands, one hand, or both. The three are mutually exclusive and sum to 100%, give or take the rounding — each is rounded independently to two decimal places, so a sum can land on 99.99. When no_hands_percent is high: the camera is aimed too high or too far forward, the mount has shifted, or the task genuinely does not involve the hands for long stretches. What it does not mean: that nothing is happening. A hand can be doing work just outside the frame. This measures framing, not activity — which is exactly why it is worth reading separately. What to do. This is the earliest warning you get about a capture setup problem, and it is the cheapest one to fix. Persistent no_hands_percent on work you expected to be manual usually means the mount or the framing is wrong, and every hour recorded after that point carries the same flaw. Checking it early in a collection is worth far more than checking it at the end. A high two_hands_percent tells you the footage is suited to bimanual work. If you are collecting for two-armed manipulation, this is a coverage measurement rather than a defect one.

Hand-object contact

What it means. When hands are visible, what they are touching. Contact against manipulation. any_contact_percent is never below the other: a hand resting on a bench is in contact but is not manipulating anything, and manipulation without contact is not counted at all. The gap between them is passive contact, and a wide gap on a task you expected to be busy is worth a look. The two are equal when every moment of contact is active. Portable against non-portable is the one that describes what kind of work the footage contains. Contact dominated by fixed surfaces suggests leaning, steadying, and operating machinery in place; contact dominated by portable objects suggests handling and transport. Neither is better — but if you are looking for one and the footage is mostly the other, this is where it shows. The two can occur at the same moment, so they do not sum to any_contact_percent. What it does not mean: contact is not grasping, and it is not force. A hand brushing an object registers the same as a hand gripping one.

Reading the response

Response (abridged)
assessed_seconds is the denominator. 255,600 seconds is 71 hours, so 41.8% active manipulation is 29.7 hours of work. coverage_percent is a different thing and is easy to confuse with the percentages above it: it reports how much of the planned analysis produced a usable result, weighting each planned clip equally rather than each second. A low value means these percentages describe less of the dataset than you might assume. It is not a quality score. null values appear when no footage produced a complete, applicable observation. A 0 means the opposite — we looked and the event did not occur.

How this is measured

Footage is assessed by a vision model, which reports whether the wearer’s hands are visible, whether they are in contact with something, and whether that contact is purposeful. Footage that does not produce a determinate answer is left out of assessed_seconds rather than guessed at, which is why the denominator is published alongside the percentages. The three stay separate because they come apart in practice: a visible hand may touch nothing, and contact need not be purposeful. Visible hands, contact state and the object being acted on are annotated as distinct labels in egocentric vision research such as 100 Days of Hands and VISOR.