1 Introduction
1.1 Measurement scales
We recall from Stats 101 and Stevens (1946) that measured variables can have one of four scales:
- nominal, qualitative; values are names that can not be ordered in a meaningful way (land use: urban, forest, water)
- ordinal, qualitative; values are names that can be ordered unambiguously (terrain: flat, undulating, hilly, mountainous)
- interval, numeric; values without a meaningful (absolute) zero, where ratios do not make sense (degrees Celcius, hour-of-day)
- ratio, numeric; values with an absolute zero, where ratios make sense (mass, duration, distance)
These scales are useful, because different scales allow different operations (“is equal”, “is greater than”, subtract, divide).
In addition, for numeric we can distinguish between discrete and continuous variables:
- discrete variables can take on a finite set of values, e.g. the natural numbers ℕ (0,1,2,…),
- continuous variables can take any value over a given interval; e.g. the real numbers ℝ
Further, special variable types are e.g.
- circular variables (hour-of-day, day-of-week, wind direction), or
- bounded variables such as relative frequencies and probabilities.
1.2 Fields and coverages
One can think of space and time as dimensions with coordinates that should be real numbers (e.g. for space a location \(s\) is defined in \(s \in ℝ^2\) or \(s \in ℝ^3\), for time instances \(t\) with \(t \in ℝ^1\), for space-time possibly \((s,t) \in ℝ^4\)), as, in theory, we can move in space, or pick moments in time, continuously, meaning with infinitely small steps. These are obviously theoretical (but useful) models, as in practice we for instance only measure anything at a finite number of locations and times. Also, when we consider something located the Earth surface, the space considered is not \(ℝ^2\) (flat, unbounded) or \(ℝ^3\) (“space”, unbounded) but closer to \(S^2\), the bounded, finite surface of a sphere.
We will use \(s\) for spatial coordinates (two- or three-dimensional), and \(t\) for the time coordinate (one-dimensional).
When we observe or measure some kind of variable, which we will call \(y\), it must happens at some location \(s\) and time \(t\), and we can think of this as the tuple \(\{y,s,t\}\).
When the variable \(y\) can only take on a single value for each specific pair \(s,t\), we can think of it as a function,
\[y = f(s,t)\]
with \(s\) and \(t\) continuous variables (the domain of the function), and \(y\) the range of the function. Examples could be air temperature (in three spatial dimensions, and varying over time) or surface elevation with respect to sea level (in two dimensions, or of the Earth surface; often static in time).
The variable measured can also be qualitative (a category, e.g. land use, or bedrock or soil type). In that case we have \(y(s,t)\) with:
- \(y\) being categorical, and
- \(s\) and \(t\) continuous
Variables that are functions of continuous space and/or time may, in geographic context, be called fields, or coverages, or geostatistical variables.
1.3 Measurements
Measurements of continuous functions \(y(s,t)\) are necessarily discrete, but may appear continuous:
- dense time series \(y(s)\) collected with high frequency may be shown on a graph with a time line, suggesting we measured continuously
- imagery, such as sattelite imagery, may appear so detailed that we can no longer see individual pixels, and suggest we measured continuousl over time
On the other hand, measurement will be sparse when measurement devices are expensive or measurement is harmful, examples being
- stationary LANUV air quality sensors, which are expensive and need maintenance
- subsurface oil reservoir estimation needs deep drillings, which cost millions
- where body scans (CT, MRT) are “cheap”, sampling of body tissue (biopsy) is something doctors will try to minimize (if not avoid).
This leads to sparse sampling, and the problem of estimating (“predicting”) the variable at unsampled locations, where it exists but was not measured; this problem may be called spatial interpolation or spatial prediction. The locations of the samples are not of primary interest, of primary interest is the value of the variable at all locations (or some aggregate properties: were there any cancer sells in the tissue sampled? what was the average temperature?).
1.4 Variables with discrete time and space
Many phenomena do not vary continuously over space and/or time, but are discrete of their nature. Examples include:
- lightning strikes
- traffic accidents
- disease cases
- car engine failures
All these examples are discrete in space and time: at arbitrary locations and times, they do not happen, or occur; we call measurements on such phenomena spatiotemporal point patterns.
In contrast to geostatistical data, where the location of the sensors is not of primary interest, for point patterns the locations of points (objects, events) are of primary interest: where (and when) the traffic accident happened tells us something about accident risk, and about whether intervention (changing the traffic situation) may help preventing future accidents.
Phenomena may also be discrete in space, but continuous in time; examples include wind turbines generating energy, power plants emitting \(CO_2\), or concert halls having a number of visitors. Other phenomena may be continuous in space but discrete in time, but examples are harder to find (elections?).
Timely persisting objects may also move around (people, cars, birds, hurricanes) and generate data about their state (resp. heart beat rate, engine temperature, altitude, extent); this leads to trajectory data (Oueslati et al. (2023)).
1.5 Aggregations
In addition to geostatistical data and point pattern data, in many cases datasets contain aggregated values, where aggregation took place over some region, time period, or both. We may think of:
- health or socio-economic data, aggregated for administrative regions
- GDP, a monetary amount summed over (usually) a country and a year
- population, or population density, aggregated over a region (grid cell or administrative region)
Aggregated values “inherit” some of the properties of the original (non-aggregated) data, but some variation also gets lost, depending on the nature of the non-aggregated data (smooth / noisy) and the size of the aggregation area/period.
1.6 Further reading
Pebesma and Bivand (2023) has additional discussion on spatial data science, practical problems and many examples and exercises including code. The second edition of this book, subtitled with applications in R and Python is work in progress and also contains Python code.