Skip to content

Kernel Density Estimate (KDE)

Definition

A non-parametric smooth probability density function fitted to a set of measurements by placing a kernel (usually Gaussian) over each data point and summing them. Used in glass and other trace evidence LR models to convert a set of reference measurements into a continuous density that can be evaluated at any measurement value.

Type
Non-parametric probability density estimate
Typical kernel
Gaussian
Common use
Glass and trace evidence likelihood ratio models
Input
A set of reference measurements

Common questions

Why choose a KDE over fitting a normal distribution to reference data?+

A KDE makes no assumption that the underlying data follows a particular shape such as a normal curve, so it can capture skewed or multi-modal patterns that a single parametric distribution would misrepresent. This matters for trace evidence populations, such as refractive index measurements, that do not always follow a simple bell curve.

What choice most affects the accuracy of a KDE-based likelihood ratio?+

The bandwidth, which controls how much smoothing is applied around each data point, has the largest effect. Too narrow a bandwidth overfits to the specific reference sample and produces a jagged, unreliable density; too wide a bandwidth oversmooths and can understate how distinctive a measurement really is.

Related terms

Bandwidth
The smoothing parameter in a kernel density estimate. A small bandwidth produces a jagged curve that follows every data point; a large...
Box Plot (Box-and-Whisker Plot)
A graphic showing the five-number summary of a distribution: the minimum, first quartile (Q1), median (Q2), third quartile (Q3), and maximum. Whiskers...
Defence Proposition (Hd)
The alternative proposition, typically asserting that someone else is the source (e.g., 'the crime-scene DNA came from an unknown, unrelated person'). It...
Histogram
A bar chart in which data are grouped into contiguous equal-width intervals (bins) and the bar height represents the count or relative...
Interquartile Range (IQR)
The difference between the third quartile and the first quartile: IQR = Q3 - Q1. It measures the spread of the central...
Likelihood Ratio (LR)
The ratio of two conditional probabilities: the probability of the observed evidence given the prosecution's hypothesis (same source), divided by the probability...
Log-LR (Log Likelihood Ratio)
The natural or base-10 logarithm of the LR. Log-LRs are additive for independent evidence types, making them convenient for combining across disciplines....
Prosecution Proposition (Hp)
The proposition advanced by the prosecution, typically asserting that the defendant is the source of the questioned material (e.g., 'the crime-scene DNA...
Random Match Probability (RMP)
The probability that a randomly chosen unrelated person from the relevant population would match the evidence profile by chance. A very small...
Scatter Plot
A two-dimensional graph in which each observation is plotted as a point at coordinates (x, y), where x and y are two...

Explained in these topics

Your journey to becoming a forensic professional starts here.

Practice with mock tests, learn from structured notes, and get your questions answered by a global forensic community, all in one place.