Skip to main content
Content diversity answers one question: what is actually in this footage? Not whether it is damaged, and not whether anyone is working — what settings, objects, actions and activity contexts the dataset contains, and how often each one turns up. Each is measured independently. The default response returns all three axes with their category summaries. Ask for one axis with ?axis=objects. Exact activity contexts are available as task_contexts on each row of the paginated quality-clip response.

The unit is clips, not time

Everything here counts clips, not seconds. clip_count is the number of distinct analyzed clips a concept appears in, counted once per clip no matter how often it occurs inside one. clip_percent divides that by analyzed_clip_count. Concepts co-occur — one clip contains a kitchen, a kettle and a cup — so percentages within an axis do not sum to 100 and are not meant to.

Environments

What it means. The kinds of settings the footage was shot in. Environment concepts are also grouped under the broad setting categories used by Lightwheel EgoSuite, with other retained for settings that do not fit one category or cannot be identified reliably. The open environment concepts remain the detailed evidence; the category is only a roll-up. What to use it for. Checking that a dataset spans the range of places you expected. Footage collected entirely in one room generalises differently from the same number of hours across twenty. What it does not mean: it does not identify specific locations, and it cannot tell two different kitchens apart. Twenty clips reading kitchen might be twenty kitchens or one kitchen twenty times — a distinction that matters and that this axis cannot make for you.

Objects

What it means. Identifiable physical objects and materials observed in the footage, including things that are simply present in the background. Visible against interacted. Each object carries two distributions:
  • its ordinary clip_count and clip_percent — every clip the object appears in;
  • an interaction block — the subset of those clips where the object was actually involved in an action.
The gap between them is the useful part. A workbench appears in every clip and is handled in none; a screwdriver that appears in 40% of clips but is interacted with in 4% is mostly sitting on the table. For manipulation work the interaction numbers are the ones that matter, and the plain counts tell you about context and scene composition. What to do. Check the objects you care about appear at all, and in enough clips to be worth having. A named object with clip_count: 2 across a thousand-clip dataset is present but not represented.

Action verbs

What it means. What is being done, as base-form verbs — pick up, pour, wipe. What to use it for. The fastest check that footage contains the activity you are after. If you are looking for assembly work and the verb distribution is dominated by walk and look, you can see that in one glance without opening a clip. Actions that involve no object still count here, which is why this axis and object_actions do not line up exactly.

Task contexts

What it means. Each observed action together with the full set of objects the inventory related to it and the setting reported for that same window — for example, fold · towel @ laundry room. A repeated fold · towel @ laundry room context counts once for that clip. Contexts preserve observed relationships without assigning purpose categories. The action-verb axis retains its separate dexterity grouping. Exact contexts are available on each row of the paginated quality-clip response. They are omitted from content-diversity summaries and the generic episode enrichment-term facet because the combinations can produce a large vocabulary.

Object-actions

What it means. Observed combinations of a verb and an object — pour + kettle — with their own clip counts. These come from the same observations as the separate axes. The API does not build combinations by pairing verbs and objects that merely appeared in the same clip, so a combination in this list was actually seen. What to use it for. Inspecting the observed relationships between actions and objects. objects tells you a kettle was there and action_verbs tells you something was poured; object_actions tells you the kettle was the thing poured.

Reading the distribution

The length of the concepts array is a breadth number and is easy to over-read on its own. The array is ordered most frequent first and tells you whether that breadth is real. A dataset spanning 130 objects sounds broad. If the top three appear in 80% of clips, it is a narrow dataset with a long tail of incidental background items. The head of the distribution describes what the footage is about; the tail describes what happened to be in shot. Two checks worth automating: Coverage. Every concept you care about appears, with a clip_count above some floor you set. Concentration. The percentage of clips taken by the top few concepts. Rising concentration across successive collections usually means the footage is becoming repetitive.
Response (abridged, axis=objects)
The workbench is scenery: present in 86% of clips, handled in 2%. The screwdriver is the work: present in 40%, handled in 38%.

Categories

Each axis also returns categories — stable groupings with example concepts, useful for rolling a long distribution up into something readable in a report. Each row carries a stable machine category and an exact human-facing label. Each carries two counts, and they answer different questions: concept_count tells you the vocabulary has, say, forty-six deformable-object words in it. clip_count tells you how much of the footage actually involves a deformable object, which is usually the question being asked. A dataset can name forty-six of them and handle one in 2% of clips. A clip is counted once per category however many of the category’s concepts it contains — a clip holding a towel and a cable is one deformable clip, not two. Categories can also overlap, since one clip can hold a deformable object and a non-deformable one, so clip_percent does not sum to 100 across an axis, in exactly the way concept percentages do not. They are descriptive buckets, not scores. A category with many concepts is not better than one with few, and concepts with no category still count toward the axis.

How this is measured

Clips are described by a vision model, and the resulting terms are reduced to a canonical vocabulary so that cup, mug and coffee cup do not appear as three separate concepts inflating the count. Verbs, objects and their combinations come from the same descriptions, which is what lets the combination list be observations rather than inferences. A concept is published only once it has been corroborated, so a dataset can come back with empty axes and a non-zero analyzed_clip_count. That means it has not been analysed enough times yet, not that nothing was found in it. Describing a dataset by its distributions rather than by one diversity score is the norm in the field: Ego4D and EPIC-KITCHENS report scenario, noun and verb distributions, and DROID and Open X-Embodiment report task, object and scene distributions. None defines a threshold for “diverse enough”, and neither do we.