unreadable is a
percentage of planned clips.
Each answers its own question, so read them one at a time. Do not add them
together. Defects can overlap, so a total can count the same footage twice.
None of these is a verdict. Each is a measurement you set your own bar against —
see Build a quality-control check.
Black
What it means. The percentage of footage with no usable picture — frames that are black or near-black. When it is high: the lens was covered, the recording started before the camera was in position or ran after it was put down, or the sensor dropped out. What it does not mean: that the footage is dark. A dimly lit room with visible content is not black. This measures the absence of a picture, not a low-light scene. What to do. Black footage carries nothing, so it is safe to trim or exclude. The aggregate response cannot locate it in time; inspect the source or clip-level data before trimming so intentional head/tail black is not confused with a mid-recording fault.Frozen
What it means. The percentage of footage where the picture has stopped changing — consecutive frames that are identical or nearly so. When it is high: capture stalled, an encoder hung, or frames were duplicated during a bad transfer or re-encode. What it does not mean: that the scene is still. A stationary camera watching a motionless subject still produces sensor noise that varies frame to frame; a genuine freeze does not. Long static shots will not register here. What to do. Frozen footage contains no information beyond its first frame, so there is nothing to recover. A non-trivial value is worth tracing to its cause rather than silently discarding, since a fault that duplicated frames once is probably still in the path.Blurry
What it means. The percentage of footage persistently too soft to identify what is being handled, excluding blur caused only by the camera moving, turning, or shaking. When it is high: autofocus hunting, a dirty, smeared, or fogged lens, heavy denoising or upscaling, or persistent defocus. What it does not mean: deliberate shallow depth of field, or a soft background behind a sharp subject. What to do. Your own footage is the best reference here. Run this across a dataset you already consider good to see where your normal sits before deciding what counts as too much.Unstable
What it means. The percentage of footage where the picture jumps between consecutive frames — the whole scene vibrating — rather than moving with the person carrying the camera. When it is high: running or rapid reversals, an unsecured mount, hand-held capture without stabilisation, or vibration transmitted from a vehicle or machine. What it does not mean: camera motion by itself. Smooth low-frequency walking, turning, and looking around is largely removed; high-frequency reversals or vibration can still count. Egocentric footage is constantly in motion by nature, and that alone is not a defect. Nor is scene activity itself camera motion: a fixed camera can remain steady while watching a crowded workspace. On a moving mount. A camera carried on a robot arm, or on any other machine that changes direction quickly, reads high here. The mount really is moving that way, so the number is doing its job, but it describes the mount more than the capture. Compare this figure between datasets only when the camera is mounted the same way in both. What to do. Sample some by eye before setting a bar. How much instability is acceptable depends on what you intend to do with the footage, so this is a measurement worth seeing for yourself rather than taking a number for.Clipped highlights and crushed shadows
What it means. The percentage of footage whose robust luma percentile crosses the fixed bright- or dark-end gate. This is an endpoint-saturation indicator; it does not prove where clipping or crushing occurred in the capture pipeline. When it is high: the sun or a window in frame, an underlit interior, or exposure settings fixed for the wrong conditions. What it does not mean: that the footage is merely bright or dark, or that it is an EBU R 103 conformance measurement. What it will not catch. Footage that is badly exposed without reaching an endpoint. A backlit subject, a scene metered for the wrong part of the frame, mid-tones sitting too low or too high: these look wrong while both numbers can stay at zero. Luma-only measurement can also miss isolated channel clipping, and letterbox or pillarbox bars can affect the reading. Why two numbers. They are separate defects with separate fixes — one wants less light or less gain, the other wants more. They also cannot be averaged into a single “exposure” figure: footage that is too dark for half its duration and too bright for the other half would average out to “correct”, which is the opposite of the truth.Unreadable
What it means. The percentage of planned clips that produced no usable measurement — for example, corrupt, truncated, unsupported, or too short to measure. What it does not mean: that the footage is bad. It means the deterministic video-characteristics pass produced no usable reading. Independent blur or camera-motion observations may still cover the clip. It is counted in clips, not hours. Every other measurement here is weighted by duration. This one cannot be: footage that never decoded has no measured duration to weight it by, and estimating one from the clips that did decode would be a number we made up. It reportsassessed_clips where the others report
assessed_hours. It has no distribution because there is no assessed duration
for unreadable clips to weight one by.
A clip that decoded but produced too little for the deterministic pass is
counted here too.
What to do. This is often a storage or transfer problem that resolves by
fetching the files again. If the files open, check whether the clips are simply
too short to measure. Read this metric first: if it is high, the other
deterministic percentages describe a smaller set of clips than the plan.
Reading the response
The default response carries each percentage and its denominator. Addinclude=distribution when you need chart data for the duration-weighted
measurements:
Response (abridged)
percent is the percentage of assessed footage affected, weighted by
duration.
assessed_hours on a measurement is the footage that measurement was taken
over. Multiply the two to get hours: 8.5% of 71.0 hours is just over six hours
crossing the dark-end gate.
Each measurement carries its own, and they are not all the same — a measurement
covers the footage it could be taken over. Use the number sitting next to the
percentage rather than the one at the top of the response, which is there to
tell you how much footage was looked at overall.
unreadable reports assessed_clips instead, because it counts clips. It never
has distribution; the other measurements omit distribution unless it was
requested.
histogram puts each clip assessed for that defect into one
affected-percentage bin. Plot the bin minimums on the x-axis and
footage_percent on the y-axis. footage_percent is weighted
by assessed duration, while clip_count is deliberately unweighted. This lets
you see whether a bin represents a few long clips or many short ones. All ten
bins are returned, including empty ones; exact boundaries start the next bin,
and 100 is included in the last.
percentiles are duration-weighted nearest-rank values over the same
per-clip percentages. They can feed a weighted box plot directly: p0 and p100
are the whiskers, p25 and p75 the box, and p50 the median. p90 and p95 expose the
tail without inventing a service threshold.
How these are measured
Black, frozen, and endpoint saturation are measured directly from the footage.unreadable instead
compares the planned clips with the clips that produced a usable measurement.
Blur is judged by a model. Instability uses camera-motion measurements when they
are available; a complete per-clip model assessment fills spans that measurement
cannot classify, and the model also acts as a fallback when camera motion has not
been measured. Both are graded by degree, so
borderline footage stays borderline in the percentage instead of being rounded
to a yes or a no. A metric is null when neither of its assessment paths has
run.
The deterministic values reproduce when the footage, filter graph, and FFmpeg
build are held fixed. Blur and fallback stability judgements are model outputs
and are not bit-identical across runs. Their auxiliary run tokens are not
exposed or pinnable today, so even a request with both returned primary tokens
can change when newer auxiliary evidence is published.