Skip to main content
This page works through one situation: footage arrives, and you have to decide whether to accept it, look at it, or send it back, without watching it. The accept-and-reject vocabulary belongs to that scenario rather than to the API. Three questions, in order. Later ones are only worth asking if earlier ones pass, because damaged footage cannot be rescued by being on-topic.

1. Is the footage damaged?

Read the measurements in this order:
  1. unreadable. This is a percentage of planned clips with no usable deterministic video-characteristics reading, not a percentage of hours. Check those clips before interpreting the deterministic metrics; independent blur or stability observers may still cover them.
  2. frozen and black. Footage that carries nothing, and usually fixable by trimming rather than rejecting.
  3. clipped_highlights and crushed_shadows. Matters of degree.
  4. blurry and unstable. Also matters of degree, and the two most worth sampling by eye before you set a bar.
Two things to know before acting on them:
  • null means the measurement was not taken. It is not a clean zero.
  • They can overlap. Read them one at a time; adding them into a damage total can count footage twice.

Trim before you reject

Request include=distribution and check the histogram before rejecting on black or frozen. 3% black spread thinly across every clip is a different problem from 3% concentrated in a few clips, and only the second can be handled by reviewing or excluding a small subset. The histogram shows concentration across clips but not position in time. Inspect the source or clip-level data to distinguish intentional head/tail black from a mid-recording fault.

Setting your bars

We do not publish recommended thresholds, because one that suits your cameras and lighting will not suit someone else’s. Deriving your own takes one batch:
  1. Run the measurements against footage you already consider good.
  2. Record what each one reads. That is your baseline, not zero.
  3. Set each bar above its baseline, with enough margin for normal variation.
unreadable and frozen often sit near zero on healthy footage, so those bars can be tight. The other metrics can also be exactly zero; derive their normal levels from a trusted delivery rather than assuming either zero or a universal nonzero baseline. Err loose to begin with. A false rejection means re-shooting good footage, while a false acceptance means one noisy batch among many.

An example check

Return findings rather than a verdict, so the routing stays yours:
Whether a finding means “trim it” or “send it back” is a distribution question:
A concentrated tail says a small set of severe clips is worth finding in your clip-level data; the distribution does not identify them. Black spread evenly is a problem throughout the dataset.

2. Is anything happening in it?

Undamaged footage of nothing is still not worth much. active_manipulation_percent is the one to read. 100 hours reading 15% is about fifteen hours of work and eighty-five of walking and waiting. The visibility split tells you whether the camera is pointed at the work at all. Persistent no_hands_percent on work you expected to be manual is worth catching early, because it affects everything shot afterwards.
These fields describe a human camera-wearer’s hands. They do not apply to robot grippers, and they do not judge whether a task was done well.

3. Is it the content you were expecting?

Two checks worth automating: Presence. The objects and actions you expected appear at all, and in enough clips to be worth having. A named object with clip_count: 2 across a thousand-clip batch did not really get filmed. Concentration. 130 objects sounds broad, but if the top three appear in 80% of clips it is narrower than the count suggests. Content diversity describes composition. It does not detect duplicate footage, and it does not judge whether a demonstration succeeded.

Turn a percentage into a conversation

Percentages are easier to act on as footage:
8.5% crushed shadows across 71 assessed hours is 6 hours reaching the dark-end gate, with 80% of the assessed footage in clips below 10% affected and a p95 clip value of 31%.
That is specific enough for someone to act on, which “quality score 62” is not. Store the dataset, the response, the bars you applied and the outcome together, so a decision can be re-examined later against a bar you have since changed.