Skip to main content
Video characteristics answers one question: how much of this dataset is visibly damaged, and in what way? Seven measurements. Six are percentages of assessed footage; unreadable is a percentage of planned clips. Each answers its own question, so read them one at a time. Do not add them together. Defects can overlap, so a total can count the same footage twice. None of these is a verdict. Each is a measurement you set your own bar against — see Build a quality-control check.

Black

What it means. The percentage of footage with no usable picture — frames that are black or near-black. When it is high: the lens was covered, the recording started before the camera was in position or ran after it was put down, or the sensor dropped out. What it does not mean: that the footage is dark. A dimly lit room with visible content is not black. This measures the absence of a picture, not a low-light scene. What to do. Black footage carries nothing, so it is safe to trim or exclude. The aggregate response cannot locate it in time; inspect the source or clip-level data before trimming so intentional head/tail black is not confused with a mid-recording fault.

Frozen

What it means. The percentage of footage where the picture has stopped changing — consecutive frames that are identical or nearly so. When it is high: capture stalled, an encoder hung, or frames were duplicated during a bad transfer or re-encode. What it does not mean: that the scene is still. A stationary camera watching a motionless subject still produces sensor noise that varies frame to frame; a genuine freeze does not. Long static shots will not register here. What to do. Frozen footage contains no information beyond its first frame, so there is nothing to recover. A non-trivial value is worth tracing to its cause rather than silently discarding, since a fault that duplicated frames once is probably still in the path.

Blurry

What it means. The percentage of footage persistently too soft to identify what is being handled, excluding blur caused only by the camera moving, turning, or shaking. When it is high: autofocus hunting, a dirty, smeared, or fogged lens, heavy denoising or upscaling, or persistent defocus. What it does not mean: deliberate shallow depth of field, or a soft background behind a sharp subject. What to do. Your own footage is the best reference here. Run this across a dataset you already consider good to see where your normal sits before deciding what counts as too much.

Unstable

What it means. The percentage of footage where the picture jumps between consecutive frames — the whole scene vibrating — rather than moving with the person carrying the camera. When it is high: running or rapid reversals, an unsecured mount, hand-held capture without stabilisation, or vibration transmitted from a vehicle or machine. What it does not mean: camera motion by itself. Smooth low-frequency walking, turning, and looking around is largely removed; high-frequency reversals or vibration can still count. Egocentric footage is constantly in motion by nature, and that alone is not a defect. Nor is scene activity itself camera motion: a fixed camera can remain steady while watching a crowded workspace. On a moving mount. A camera carried on a robot arm, or on any other machine that changes direction quickly, reads high here. The mount really is moving that way, so the number is doing its job, but it describes the mount more than the capture. Compare this figure between datasets only when the camera is mounted the same way in both. What to do. Sample some by eye before setting a bar. How much instability is acceptable depends on what you intend to do with the footage, so this is a measurement worth seeing for yourself rather than taking a number for.

Clipped highlights and crushed shadows

What it means. The percentage of footage whose robust luma percentile crosses the fixed bright- or dark-end gate. This is an endpoint-saturation indicator; it does not prove where clipping or crushing occurred in the capture pipeline. When it is high: the sun or a window in frame, an underlit interior, or exposure settings fixed for the wrong conditions. What it does not mean: that the footage is merely bright or dark, or that it is an EBU R 103 conformance measurement. What it will not catch. Footage that is badly exposed without reaching an endpoint. A backlit subject, a scene metered for the wrong part of the frame, mid-tones sitting too low or too high: these look wrong while both numbers can stay at zero. Luma-only measurement can also miss isolated channel clipping, and letterbox or pillarbox bars can affect the reading. Why two numbers. They are separate defects with separate fixes — one wants less light or less gain, the other wants more. They also cannot be averaged into a single “exposure” figure: footage that is too dark for half its duration and too bright for the other half would average out to “correct”, which is the opposite of the truth.

Unreadable

What it means. The percentage of planned clips that produced no usable measurement — for example, corrupt, truncated, unsupported, or too short to measure. What it does not mean: that the footage is bad. It means the deterministic video-characteristics pass produced no usable reading. Independent blur or camera-motion observations may still cover the clip. It is counted in clips, not hours. Every other measurement here is weighted by duration. This one cannot be: footage that never decoded has no measured duration to weight it by, and estimating one from the clips that did decode would be a number we made up. It reports assessed_clips where the others report assessed_hours. It has no distribution because there is no assessed duration for unreadable clips to weight one by. A clip that decoded but produced too little for the deterministic pass is counted here too. What to do. This is often a storage or transfer problem that resolves by fetching the files again. If the files open, check whether the clips are simply too short to measure. Read this metric first: if it is high, the other deterministic percentages describe a smaller set of clips than the plan.

Reading the response

The default response carries each percentage and its denominator. Add include=distribution when you need chart data for the duration-weighted measurements:
Response (abridged)
percent is the percentage of assessed footage affected, weighted by duration. assessed_hours on a measurement is the footage that measurement was taken over. Multiply the two to get hours: 8.5% of 71.0 hours is just over six hours crossing the dark-end gate. Each measurement carries its own, and they are not all the same — a measurement covers the footage it could be taken over. Use the number sitting next to the percentage rather than the one at the top of the response, which is there to tell you how much footage was looked at overall. unreadable reports assessed_clips instead, because it counts clips. It never has distribution; the other measurements omit distribution unless it was requested. histogram puts each clip assessed for that defect into one affected-percentage bin. Plot the bin minimums on the x-axis and footage_percent on the y-axis. footage_percent is weighted by assessed duration, while clip_count is deliberately unweighted. This lets you see whether a bin represents a few long clips or many short ones. All ten bins are returned, including empty ones; exact boundaries start the next bin, and 100 is included in the last. percentiles are duration-weighted nearest-rank values over the same per-clip percentages. They can feed a weighted box plot directly: p0 and p100 are the whiskers, p25 and p75 the box, and p50 the median. p90 and p95 expose the tail without inventing a service threshold.
An average alone cannot distinguish a defect spread lightly through most clips from one concentrated in a small severe tail. The histogram shows how much footage sits in that tail, and the weighted box plot summarizes its shape. Neither identifies the clips or positions in time. Affected and footage percentages are rounded independently to two decimal places, so the published histogram may total slightly above or below 100.

How these are measured

Black, frozen, and endpoint saturation are measured directly from the footage. unreadable instead compares the planned clips with the clips that produced a usable measurement. Blur is judged by a model. Instability uses camera-motion measurements when they are available; a complete per-clip model assessment fills spans that measurement cannot classify, and the model also acts as a fallback when camera motion has not been measured. Both are graded by degree, so borderline footage stays borderline in the percentage instead of being rounded to a yes or a no. A metric is null when neither of its assessment paths has run. The deterministic values reproduce when the footage, filter graph, and FFmpeg build are held fixed. Blur and fallback stability judgements are model outputs and are not bit-identical across runs. Their auxiliary run tokens are not exposed or pinnable today, so even a request with both returned primary tokens can change when newer auxiliary evidence is published.