The Nyquist–Shannon sampling theorem, or the sampling theorem, is a theorem in the field of signal processing which serves as a fundamental bridge between continuous-time signals and discrete-time signals. In the case of uniformly spaced (periodic) sampling, it establishes a sufficient condition on the sample rate that permits a discrete sequence of samples to capture all the information from a continuous-time signal of finite bandwidth, such that the original signal can be reconstructed exactly from those samples.
Strictly speaking, the theorem only applies to a class of mathematical functions having a Fourier transform that is zero outside of a finite region of frequencies. Intuitively we expect that when one reduces a continuous function to a discrete sequence and interpolates back to a continuous function, the fidelity of the result depends on the density (or sample rate) of the original samples. The sampling theorem introduces the concept of a sample rate that is sufficient for perfect fidelity for the class of functions that are band-limited to a given bandwidth, such that no actual information is lost in the sampling process. It expresses the sufficient sample rate in terms of the bandwidth for the class of functions. The theorem also leads to a formula for perfectly reconstructing the original continuous-time function from the samples.
Perfect reconstruction may still be possible when the sample-rate criterion is not satisfied, provided other constraints on the signal are known (see § Sampling of non-baseband signals below and compressed sensing). In some cases (when the sample-rate criterion is not satisfied), utilizing additional constraints allows for approximate reconstructions. The fidelity of these reconstructions can be verified and quantified utilizing Bochner's theorem.
An important consequence of the sampling theorem is the concept of Nyquist frequency, which holds that in order to reconstruct a bandlimited signal free of aliasing, the sampling rate must be at least twice the signal's bandwidth.
The name Nyquist–Shannon sampling theorem honours Harry Nyquist and Claude Shannon, but the theorem was first proven in a 1920 paper by Kinnosuke Ogura, according to a historical review in 2011.
The theorem is also known by the names Whittaker–Shannon sampling theorem, Whittaker–Shannon, or Whittaker–Nyquist–Shannon to include E. T. Whittaker since his 1915 paper is sometimes regarded as the first paper on the sampling theorem, though it was actually about an interpolation formula and its proof was unclear and contained errors.
The theorem may also be referred to as the cardinal theorem of interpolation.
Contents
Introduction
Sampling is a process of converting a signal (for example, a function of continuous time or space) into a sequence of values (a function of discrete time or space). Shannon's version of the theorem states:
A sufficient sample-rate is therefore anything larger than
2
B
{\displaystyle 2B}
samples per second. Equivalently, for a given sample rate
f
s
{\displaystyle f_{s}}
, perfect reconstruction is guaranteed possible for a bandlimit
B
<
f
s
/
2
{\displaystyle B<f_{s}/2}
.
When the bandlimit is too high (or there is no bandlimit), the reconstruction exhibits imperfections known as aliasing. Modern statements of the theorem are sometimes careful to explicitly state that
x
(
t
)
{\displaystyle x(t)}
must contain no sinusoidal component at exactly frequency
Aliasing
When
x
(
t
)
{\displaystyle x(t)}
is a function with a Fourier transform
X
(
f
)
{\displaystyle X(f)}
:
X
(
f
)
≜
∫
−
∞
∞
x
(
t
)
e
−
i
2
π
f
t
d
t
,
{\displaystyle X(f)\ \triangleq \ \int _{-\infty }^{\infty }x(t)\ e^{-i2\pi ft}\ {\rm {d}}t,}
Derivation as a special case of Poisson summation
When there is no overlap of the copies (also known as "images") of
X
(
f
)
{\displaystyle X(f)}
, the
k
=
0
{\displaystyle k=0}
term of Eq.1 can be recovered by the product:
X
(
f
)
=
H
(
f
)
⋅
X
1
/
T
(
f
)
,
{\displaystyle X(f)=H(f)\cdot X_{1/T}(f),}
where:
Shannon's original proof
Poisson shows that the Fourier series in Eq.1 produces the periodic summation of
X
(
f
)
{\displaystyle X(f)}
, regardless of
f
s
{\displaystyle f_{s}}
and
B
{\displaystyle B}
. Shannon, however, only derives the series coefficients for the case
f
s
=
2
B
{\displaystyle f_{s}=2B}
. Virtually quoting Shannon's original paper:
Let
X
(
ω
)
{\displaystyle X(\omega )}
be the spectrum of
x
(
Application to multivariable signals and images
The sampling theorem is usually formulated for functions of a single variable. Consequently, the theorem is directly applicable to time-dependent signals and is normally formulated in that context. However, the sampling theorem can be extended in a straightforward way to functions of arbitrarily many variables. Grayscale images, for example, are often represented as two-dimensional arrays (or matrices) of real numbers representing the relative intensities of pixels (picture elements) located at the intersections of row and column sample locations. As a result, images require two independent variables, or indices, to specify each pixel uniquely—one for the row, and one for the column.
Color images typically consist of a composite of three separate grayscale images, one to represent each of the three primary colors—red, green, and blue, or RGB for short. Other colorspaces using 3-vectors for colors include HSV, CIELAB, XYZ, etc. Some colorspaces such as cyan, magenta, yellow, and black (CMYK) may represent color by four dimensions. All of these are treated as vector-valued functions over a two-dimensional sampled domain.
Similar to one-dimensional discrete-time signals, images can also suffer from aliasing if the sampling resolution, or pixel density, is inadequate. For example, a digital photograph of a striped shirt with high frequencies (in other words, the distance between the stripes is small), can cause aliasing of the shirt when it is sampled by the camera's image sensor. The aliasing appears as a moiré pattern. The "solution" to higher sampling in the spatial domain for this case would be to move closer to the shirt, use a higher resolution sensor, or to optically blur the image before acquiring it with the sensor using an optical low-pass filter.
Another example is shown here in the brick patterns. The top image shows the effects when the sampling theorem's condition is not satisfied. When software rescales an image (the same process that creates the thumbnail shown in the lower image) it, in effect, runs the image through a low-pass filter first and then downsamples the image to result in a smaller image that does not exhibit the moiré pattern. The top image is what happens when the image is downsampled without low-pass filtering: aliasing results.
Critical frequency
To illustrate the necessity of
f
s
>
2
B
,
{\displaystyle f_{s}>2B,}
consider the family of sinusoids generated by different values of
θ
{\displaystyle \theta }
in this continuous-time signal:
x
(
t
)
=
cos
(
2
π
B
t
+
θ
)
cos
(
θ
)
,
−
Sampling of non-baseband signals
As discussed by Shannon:
A similar result is true if the band does not start at zero frequency but at some higher value, and can be proved by a linear translation (corresponding physically to single-sideband modulation) of the zero-frequency case. In this case the elementary pulse is obtained from
sin
(
x
)
/
x
{\displaystyle \sin(x)/x}
by single-side-band modulation.
That is, a sufficient no-loss condition for sampling signals that do not have baseband components exists that involves the width of the non-zero frequency interval as opposed to its highest frequency component. See sampling for more details and examples.
For example, in order to sample FM radio signals in the frequency range of 100–102 MHz, it is not necessary to sample at 204 MHz (twice the upper frequency), but rather it is sufficient to sample at 4 MHz (twice the width of the frequency interval). (Reconstruction is not usually the goal with sampled IF or RF signals. Rather, the sample sequence can be treated as ordinary samples of the signal frequency-shifted to near baseband, and digital demodulation can proceed on that basis.)
Using the bandpass condition, where
X
(
f
)
=
Nonuniform sampling
The sampling theory of Shannon can be generalized for the case of nonuniform sampling, that is, samples not taken equally spaced in time. The Shannon sampling theory for non-uniform sampling states that a band-limited signal can be perfectly reconstructed from its samples if the average sampling rate satisfies the Nyquist condition. Therefore, although uniformly spaced samples may result in easier reconstruction algorithms, it is not a necessary condition for perfect reconstruction.
The general theory for non-baseband and nonuniform samples was developed in 1967 by Henry Landau. He proved that the average sampling rate (uniform or otherwise) must be twice the occupied bandwidth of the signal, assuming it is a priori known what portion of the spectrum was occupied.
In the late 1990s, this work was partially extended to cover signals for which the amount of occupied bandwidth is known but the actual occupied portion of the spectrum is unknown. In the 2000s, a complete theory was developed
(see the section Sampling below the Nyquist rate under additional restrictions below) using compressed sensing. In particular, the theory, using signal processing language, is described in a 2009 paper by Mishali and Eldar. They show, among other things, that if the frequency locations are unknown, then it is necessary to sample at least at twice the Nyquist criteria; in other words, you must pay at least a factor of 2 for not knowing the location of the spectrum. Note that minimum sampling requirements do not necessarily guarantee stability.
Sampling below the Nyquist rate under additional restrictions
The Nyquist–Shannon sampling theorem provides a sufficient condition for the sampling and reconstruction of a band-limited signal. When reconstruction is done via the Whittaker–Shannon interpolation formula, the Nyquist criterion is also a necessary condition to avoid aliasing, in the sense that if samples are taken at a slower rate than twice the band limit, then there are some signals that will not be correctly reconstructed. However, if further restrictions are imposed on the signal, then the Nyquist criterion may no longer be a necessary condition.
A non-trivial example of exploiting extra assumptions about the signal is given by the recent field of compressed sensing, which allows for full reconstruction with a sub-Nyquist sampling rate. Specifically, this applies to signals that are sparse (or compressible) in some domain. As an example, compressed sensing deals with signals that may have a low overall bandwidth (say, the effective bandwidth
E
B
{\displaystyle EB}
) but the frequency locations are unknown, rather than all together in a single band, so that the passband technique does not apply. In other words, the frequency spectrum is sparse. Traditionally, the necessary sampling rate is thus
2
B
.
{\displaystyle 2B.}
Using compressed sensing techniques, the signal could be perfectly reconstructed if it is sampled at a rate slightly lower than
2
E
B
.
Historical background
According to a 2011 historical review by Butzer et al., which investigated the history of the sampling theorem, the theorem was first proven in a 1920 paper by Kinnosuke Ogura. Although a paper published by E. T. Whittaker in 1915—prior to Ogura’s work—is sometimes regarded as the first paper on the sampling theorem, the paper actually presented an interpolation formula (the converse of the sampling theorem) and, moreover, the proof was unclear and contained errors.
Harry Nyquist in his 1924 and 1928 papers pointed out the importance of the frequency 1/“time unit” as well as its connection with the speed of transmission, but these two papers did not state the sampling theorem explicitly anywhere nor contain a statement about sampling at twice the maximum frequency of the input signal, although these two papers are often cited as sources of the sampling theorem. The terms Nyquist interval or Nyquist rate "very likely had their origin" in Shannon’s paper of 1949.
About the same time, Karl Küpfmüller showed a similar result and discussed the sinc-function impulse response of a band-limiting filter, via its integral, the step-response sine integral; this bandlimiting and reconstruction filter that is so central to the sampling theorem is sometimes referred to as a Küpfmüller filter (but seldom so in English).
Claude E. Shannon is often considered as an inventor of the sampling theorem due to his papers from 1948 and 1949, although "he never claimed any credit for originating it". In "Theorem 13" of the paper "A Mathematical Theory of Communication" of 1948, the sampling theorem is formulated as:
f
(
t
)
=
∑
n
=
−
∞
Other discoverers
Others who have independently discovered or played roles in the development of the sampling theorem have been discussed in several historical articles, for example, by Jerri and by Lüke. For example, Lüke points out that Herbert Raabe, an assistant to Küpfmüller, proved the theorem in his 1939 Ph.D. dissertation; the term Raabe condition came to be associated with the criterion for unambiguous representation (sampling rate greater than twice the bandwidth). Meijering mentions several other discoverers and names in a paragraph and pair of footnotes:
As pointed out by Higgins, the sampling theorem should really be considered in two parts, as done above: the first stating the fact that a bandlimited function is completely determined by its samples, the second describing how to reconstruct the function using its samples. Both parts of the sampling theorem were given in a somewhat different form by J. M. Whittaker and before him also by Ogura. They were probably not aware of the fact that the first part of the theorem had been stated as early as 1897 by Borel. As we have seen, Borel also used around that time what became known as the cardinal series. However, he appears not to have made the link. In later years it became known that the sampling theorem had been presented before Shannon to the Russian communication community by Kotel'nikov. In more implicit, verbal form, it had also been described in the German literature by Raabe. Several authors have mentioned that Someya introduced the theorem in the Japanese literature parallel to Shannon. In the English literature, Weston introduced it independently of Shannon around the same time.
In Russian literature it is called Kotelnikov's theorem, named after Vladimir Kotelnikov, who independently discovered it in 1933.
Why Nyquist?
Exactly how, when, or why Harry Nyquist had his name attached to the sampling theorem remains obscure. The term Nyquist Sampling Theorem (capitalized thus) appeared as early as 1959 in a book from his former employer, Bell Labs, and appeared again in 1963, and not capitalized in 1965. It had been called the Shannon Sampling Theorem as early as 1954, but also just the sampling theorem by several other books in the early 1950s.
In 1958, Blackman and Tukey cited Nyquist's 1928 article as a reference for the sampling theorem of information theory, even though that article does not treat sampling and reconstruction of continuous signals as others did. Their glossary of terms includes these entries:
Sampling theorem (of information theory)
Nyquist's result that equi-spaced data, with two or more points per cycle of highest frequency, allows reconstruction of band-limited functions. (See Cardinal theorem.)
Cardinal theorem (of interpolation theory)
A precise statement of the conditions under which values given at a doubly infinite set of equally spaced points can be interpolated to yield a continuous band-limited function with the aid of the function
sin
(
x
−
x
i
)
x
−
x
i
.
{\displaystyle {\frac {\sin(x-x_{i})}{x-x_{i}}}.}