Free account · your comment posts right after signup
In statistics the Cramér–von Mises criterion is a criterion used for judging the goodness of fit of a cumulative distribution function (CDF) compared to a given empirical distribution function , or for comparing two empirical distributions. It is also used as a part of other algorithms, such as minimum distance estimation. It is defined as , where
Article
In statistics the Cramér–von Mises criterion is a criterion used for judging the goodness of fit of a cumulative distribution function (CDF)
F
∗
{\displaystyle F^{*}}
compared to a given empirical distribution function
F
n
{\displaystyle F_{n}}
, or for comparing two empirical distributions. It is also used as a part of other algorithms, such as minimum distance estimation. It is defined as
is the empirically observed distribution. Alternatively the two distributions can both be empirically estimated ones; this is called the two-sample case.
The criterion is named after Harald Cramér and Richard Edler von Mises who first proposed it in 1928–1930. The generalization to two samples is due to Anderson.
The Cramér–von Mises test is an alternative to the Kolmogorov–Smirnov test (1933).
Contents
Cramér–von Mises test (one sample)
Let
x
1
,
x
2
,
…
,
x
n
{\displaystyle x_{1},x_{2},\ldots ,x_{n}}
be the observed values, in increasing order. Then the test statistic is
T
=
n
ω
2
=
1
12
n
+
∑
i
=
1
n
[
2
i
−
1
2
n
Watson test
A modified version of the Cramér–von Mises test is the Watson test which uses the statistic U2, where
If the value of T is larger than the tabulated values, the hypothesis that the two samples come from the same distribution can be rejected. (Some books give critical values for U, which is more convenient, as it avoids the need to compute T via the expression above. The conclusion will be the same.)
The above assumes there are no duplicates in the
x
{\displaystyle x}
,
y
{\displaystyle y}
, and
r
{\displaystyle r}
sequences. So
x
i
{\displaystyle x_{i}}
is unique, and its rank is
i
{\displaystyle i}
in the sorted list
x
1
,
…
,
x
n
{\displaystyle x_{1},\ldots ,x_{n}}
. If there are duplicates, and
x
i
{\displaystyle x_{i}}
through
x
j
{\displaystyle x_{j}}
are a run of identical values in the sorted list, then one common approach is the midrank method: assign each duplicate a "rank" of
a metric on the space of such distributions. Note that some sources define the Cramér distance as
ℓ
2
2
{\displaystyle \ell _{2}^{2}}
, but this fails the triangle inequality and so cannot be properly defined as a distance. The Cramér distance is the one-dimensional case of the energy distance via the relationship
2
ℓ
2
=
D
{\displaystyle {\sqrt {2}}\ell _{2}=D}
, and when
G
{\displaystyle G}
represents a single observation
y
{\displaystyle y}
with cumulative distribution
G
(
x
)
=
1
{
x
≥
y
}
{\displaystyle G(x)=\mathbf {1} \{x\geq y\}}
,
ℓ
2
2
(
F
,
G
)
{\displaystyle \ell _{2}^{2}(F,G)}
is equivalent to the continuous ranked probability score, a strictly proper scoring rule.
Under the probability integral transform (PIT), the plot of the empirical distribution of the transformed values
F
∗
(
x
1
)
,
…
,
F
∗
(
x
n
)
{\displaystyle F^{*}(x_{1}),\ldots ,F^{*}(x_{n})}
and the uniform distribution on
[
0
,
1
]
{\displaystyle [0,1]}
creates a PIT reliability diagram. The Cramér distance
ℓ
2
{\displaystyle \ell _{2}}
between these two distributions equals
ω
{\displaystyle \omega }
, the square root of the criterion, and serves as a numerical score of the calibration error of
F
∗
{\displaystyle F^{*}}
. This may also be referred to as the Root Mean Square Calibration Error (RMSCE).
For a deterministic (point) forecast at
μ
{\displaystyle \mu }
, the PIT degenerates to a Bernoulli random variable on
{
0
,
1
}
{\displaystyle \{0,1\}}
with success probability
p
=
Pr
(
y
>
μ
)
{\displaystyle p=\Pr(y>\mu )}
, so in the population limit the Cramér distance between the PIT CDF and the uniform distribution evaluates in closed form to
, establishing a calibration-error floor that no point forecast can fall below regardless of how accurate its central value is. In contrast, a well-calibrated probabilistic forecast can approach 0. Similarly, this quantity is maximized at the bias extremes