Skip to main content
The API returns measurements you can use in your own QC policy, not a universal quality score or pass/fail decision.

Shared response conventions

Measurement coverage

coverage_percent is the endpoint’s planned-work completeness, expressed from 0 to 100, not the percentage of footage that was good. On video characteristics it covers the deterministic FFmpeg pass only; each duration metric’s assessed_hours is its actual denominator. null means a trustworthy planned denominator is unavailable.

Counts, percentages, and unavailable values

Each metric publishes the denominator its percentages are taken over, and that denominator is not the same in all three:
  • Video-characteristics percentages are weighted by duration, and each measurement carries its own assessed_hours — they are not all the same, because each covers the footage it could be taken over. The exception is unreadable, which counts clips and carries assessed_clips: footage that never decoded has no measured duration to weight it by.
  • Hand-activity percentages are weighted by duration and reported over assessed_seconds.
  • Content-diversity percentages count clips, over analyzed_clip_count. A concept counts at most once per clip.
Multiplying a duration-weighted percentage by its own denominator converts it into hours of footage. coverage_percent is a separate quantity from all of these: it describes how much of the planned measurement produced a usable result, not how much footage was covered, so it is not the assessed duration divided by the footage you hold. Each endpoint defines it over its own planned work. A present value of 0 means the metric was measured and the event was not observed. A nullable value is null when its direct denominator is zero, or when nothing produced a usable observation.

Measurement tokens

Fields such as analysis, video_characteristics, and content_diversity are opaque measurement tokens. Store and compare the full value without parsing it; omit token query parameters to retrieve the latest result. For video characteristics, the returned tokens pin the clip-selection analysis and deterministic FFmpeg measurement. The current response does not expose tokens for its model judgement or camera-motion auxiliary runs, so blurry and unstable can move when newer auxiliary evidence is published.

Video characteristics

How much of a dataset is visibly damaged, and in what way. Each measurement is a percentage of the footage it was taken over: Each carries a percent and the footage it was taken over. Pass include=distribution for per-clip histograms and duration-weighted percentiles on every available duration metric. unreadable counts clips and has no distribution. A measurement is null when it could not be taken at all, which is not the same as a clean zero. These can overlap, so read them one at a time rather than adding them into a total. See Video characteristics for what each measurement means, when it is high, what it does not mean, and how to act on it. Comparisons are safest within one capture setup. Different cameras, resolutions and lighting conditions shift what “normal” looks like, so a bar calibrated on one collection may not transfer to another.

Hand activity

How much of a dataset shows someone actually doing something. Every percentage is taken over assessed_seconds, so multiplying converts it into hours. The three hand_visibility percentages are mutually exclusive and sum to 100, give or take the rounding — every percentage is rounded independently to two decimal places, so a sum can land on 99.99. Contact can be inactive, so any_contact_percent is never below active manipulation. Portable and non-portable contact can occur at the same moment, so they do not sum to any_contact_percent. These fields describe a human camera-wearer’s hands, not robot grippers, and active manipulation is not a task-success judgment. See Hand activity for what each measurement means, when it is high or low, what it does not mean, and how to act on it.

Content diversity

What actually appears in a dataset. Three independent axes and one relationship distribution, all counted in clips rather than seconds: Exact action, object, and setting combinations are available as task_contexts on the paginated quality-clip response, without purpose categories. clip_count counts distinct analyzed clips containing a concept, at most once per clip; clip_percent divides it by analyzed_clip_count. Concepts co-occur, so percentages within an axis do not sum to 100. Objects additionally carry an interaction block — the subset of clips where the object was involved in an action — which separates scenery from the work. categories are descriptive groupings, not diversity scores. Each carries a concept_count — how many distinct concepts fall in the bucket — alongside a clip_count and clip_percent for how much footage contains at least one of them, which is usually the question being asked. A clip counts once per category however many of its concepts it holds, and categories can overlap, so they do not sum to 100. Uncategorized concepts still contribute to their axis. A dataset can come back with empty axes and a non-zero analyzed_clip_count. That means it has not been analysed enough times for anything to be published yet, not that nothing was found in it. See Content diversity for how to read the distribution.

Duplicates

How much of a dataset is the same footage more than once. Reported as a share of compared_clip_count, which counts the clips that were compared rather than every clip in the dataset: clips the comparison has not covered are excluded, not assumed unique. exact compares media bytes over the same span. Two clips match only when the underlying file is byte-for-byte identical and they cover the same window, so the same episode packed into one video alongside others is not a duplicate of its neighbours. This needs no threshold and makes no judgment: a group is a fact about the collection. Because it compares bytes, it does not find footage that was re-encoded, resized, or re-exported. Two visually identical clips at different bitrates are not duplicates here. That is deliberate, and it is why the block is named for the notion rather than the endpoint: a comparison that tolerates re-encoding answers a different question and would arrive as a sibling block in the same response. redundant_clip_count is what could be removed without losing any footage, which is every group’s members except the one that survives it. Expect 0 on most datasets. A non-zero value usually means the same footage was included twice.