Each is measured independently. The default response returns all three axes
with their category summaries. Ask for one axis with
?axis=objects. Exact
activity contexts are available as task_contexts on each row of the paginated
quality-clip response.
The unit is clips, not time
Everything here counts clips, not seconds.clip_count is the number of
distinct analyzed clips a concept appears in, counted once per clip no matter how
often it occurs inside one. clip_percent divides that by analyzed_clip_count.
Concepts co-occur — one clip contains a kitchen, a kettle and a cup — so
percentages within an axis do not sum to 100 and are not meant to.
Environments
What it means. The kinds of settings the footage was shot in. Environment concepts are also grouped under the broad setting categories used by Lightwheel EgoSuite, withother retained for settings that do not fit one
category or cannot be identified reliably. The open environment concepts remain
the detailed evidence; the category is only a roll-up.
What to use it for. Checking that a dataset spans the range of places you
expected. Footage collected entirely in one room generalises differently from the
same number of hours across twenty.
What it does not mean: it does not identify specific locations, and it
cannot tell two different kitchens apart. Twenty clips reading kitchen might be
twenty kitchens or one kitchen twenty times — a distinction that matters and that
this axis cannot make for you.
Objects
What it means. Identifiable physical objects and materials observed in the footage, including things that are simply present in the background. Visible against interacted. Each object carries two distributions:- its ordinary
clip_countandclip_percent— every clip the object appears in; - an
interactionblock — the subset of those clips where the object was actually involved in an action.
interaction
numbers are the ones that matter, and the plain counts tell you about context and
scene composition.
What to do. Check the objects you care about appear at all, and in enough
clips to be worth having. A named object with clip_count: 2 across a
thousand-clip dataset is present but not represented.
Action verbs
What it means. What is being done, as base-form verbs —pick up, pour,
wipe.
What to use it for. The fastest check that footage contains the activity you
are after. If you are looking for assembly work and the verb distribution is
dominated by walk and look, you can see that in one glance without opening a
clip.
Actions that involve no object still count here, which is why this axis and
object_actions do not line up exactly.
Task contexts
What it means. Each observed action together with the full set of objects the inventory related to it and the setting reported for that same window — for example,fold · towel @ laundry room.
A repeated fold · towel @ laundry room context counts once for that clip.
Contexts preserve observed relationships without assigning purpose categories.
The action-verb axis retains its separate dexterity grouping.
Exact contexts are available on each row of the paginated quality-clip response.
They are omitted from content-diversity summaries and the generic episode
enrichment-term facet because the combinations can produce a large vocabulary.
Object-actions
What it means. Observed combinations of a verb and an object —pour +
kettle — with their own clip counts.
These come from the same observations as the separate axes. The API does not
build combinations by pairing verbs and objects that merely appeared in the same
clip, so a combination in this list was actually seen.
What to use it for. Inspecting the observed relationships between actions
and objects.
objects tells you a kettle was there and action_verbs tells you something
was poured; object_actions tells you the kettle was the thing poured.
Reading the distribution
The length of theconcepts array is a breadth number and is easy to over-read
on its own. The array is ordered most frequent first and tells you whether that
breadth is real.
A dataset spanning 130 objects sounds broad. If the top three appear in 80% of
clips, it is a narrow dataset with a long tail of incidental background items.
The head of the distribution describes what the footage is about; the tail
describes what happened to be in shot.
Two checks worth automating:
Coverage. Every concept you care about appears, with a clip_count above
some floor you set.
Concentration. The percentage of clips taken by the top few concepts. Rising
concentration across successive collections usually means the footage is becoming
repetitive.
Response (abridged, axis=objects)
Categories
Each axis also returnscategories — stable groupings with example concepts,
useful for rolling a long distribution up into something readable in a report.
Each row carries a stable machine category and an exact human-facing label.
Each carries two counts, and they answer different questions:
concept_count tells you the vocabulary has, say, forty-six deformable-object
words in it. clip_count tells you how much of the footage actually involves a
deformable object, which is usually the question being asked. A dataset can name
forty-six of them and handle one in 2% of clips.
A clip is counted once per category however many of the category’s concepts it
contains — a clip holding a towel and a cable is one deformable clip, not two.
Categories can also overlap, since one clip can hold a deformable object and a
non-deformable one, so clip_percent does not sum to 100 across an axis, in
exactly the way concept percentages do not.
They are descriptive buckets, not scores. A category with many concepts is not
better than one with few, and concepts with no category still count toward the
axis.
How this is measured
Clips are described by a vision model, and the resulting terms are reduced to a canonical vocabulary so thatcup, mug and coffee cup do not appear as three
separate concepts inflating the count. Verbs, objects and their combinations come
from the same descriptions, which is what lets the combination list be
observations rather than inferences.
A concept is published only once it has been corroborated, so a dataset can come
back with empty axes and a non-zero analyzed_clip_count. That means it has not
been analysed enough times yet, not that nothing was found in it.
Describing a dataset by its distributions rather than by one diversity score is
the norm in the field:
Ego4D
and EPIC-KITCHENS
report scenario, noun and verb distributions, and
DROID and Open
X-Embodiment report task, object
and scene distributions. None defines a threshold for “diverse enough”, and
neither do we.