diff --git a/_quarto.yml b/_quarto.yml index 3de5853..e95d794 100644 --- a/_quarto.yml +++ b/_quarto.yml @@ -133,16 +133,19 @@ website: contents: - lessons/R-time-series-modeling/material.qmd - lessons/R-time-series-modeling/r_tutorial.qmd + - lessons/R-time-series-modeling/equation_review.qmd - section: 'R: Time series modeling 2' href: lessons/R-time-series-modeling-2/index.qmd contents: - lessons/R-time-series-modeling-2/material.qmd - lessons/R-time-series-modeling-2/r_tutorial.qmd + - lessons/R-time-series-modeling-2/equation_review.qmd - section: 'R: Time series modeling 3' href: lessons/R-time-series-modeling-3/index.qmd contents: - lessons/R-time-series-modeling-3/material.qmd - lessons/R-time-series-modeling-3/r_tutorial.qmd + - lessons/R-time-series-modeling-3/equation_review.qmd - section: Intro to Ecological Forecasting href: lessons/intro-to-forecasting/index.qmd contents: @@ -229,11 +232,13 @@ website: contents: - lessons/R-complex-time-series-models-1/material.qmd - lessons/R-complex-time-series-models-1/r_tutorial.qmd + - lessons/R-complex-time-series-models-1/equation_review.qmd - section: 'R: Complex Time-Series Models 2' href: lessons/R-complex-time-series-models-2/index.qmd contents: - lessons/R-complex-time-series-models-2/material.qmd - lessons/R-complex-time-series-models-2/r_tutorial.qmd + - lessons/R-complex-time-series-models-2/equation_review.qmd - section: 'R: State Space Models 1' href: lessons/R-state-space-models-1/index.qmd contents: diff --git a/lessons/R-complex-time-series-models-1/equation_review.qmd b/lessons/R-complex-time-series-models-1/equation_review.qmd new file mode 100644 index 0000000..71f628b --- /dev/null +++ b/lessons/R-complex-time-series-models-1/equation_review.qmd @@ -0,0 +1,71 @@ +--- +title: Annotated Equations +order: 4 +--- + +A quick reference for the models introduced in the [R Tutorial](r_tutorial.qmd). Each equation is broken down term by term so you can check your understanding of what every symbol is doing. + +## Baseline model: AR + covariate + Gaussian error + +The starting point in `mvgam` looks like the ARIMAX models from previous lessons — a covariate, an autoregressive term, and Gaussian error — just written and fit a different way (via `mvgam()`'s Bayesian MCMC fitting, instead of `fable`'s `ARIMA()`). + +$$ +y_t = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{constant}}} + \;+\; \underbrace{\color{#319795}{\beta_1 x_{1,t}}}_{\textstyle\color{#319795}{\text{covariate}}} + \;+\; \underbrace{\color{#c05621}{\beta_2 y_{t-1}}}_{\textstyle\color{#c05621}{\text{AR: lag-1}}} + \;+\; \underbrace{\color{#718096}{\mathcal{N}(0,\sigma^2)}}_{\textstyle\color{#718096}{\text{Gaussian noise}}} +$$ + +- $y_t$ — the observed value at time $t$ (e.g., desert pocket mouse abundance) +- $c$ — a constant +- $\beta_1 x_{1,t}$ — the covariate effect (e.g., minimum temperature) +- $\beta_2 y_{t-1}$ — the AR part: the effect of the previous value +- $\mathcal{N}(0,\sigma^2)$ — Gaussian noise with mean 0 and variance $\sigma^2$ + +**How it's fit:** `mvgam` uses Bayesian methods (MCMC via STAN) rather than the exact-likelihood approach `fable` uses, but for this baseline model the equation itself is the same idea — the AR term is specified separately, via `trend_model = AR(p = 1)`, rather than inside the model formula. + +## The same model, decomposed + +The additive form above can equivalently be written as an explicit observation distribution around a modeled mean — a small notational shift that matters once we start changing the error distribution. + +$$ +y_t = \mathcal{N}( + \underbrace{\color{#d53f8c}{\mu_t}}_{\textstyle\color{#d53f8c}{\text{process mean}}},\sigma^2 +) +$$ + +$$ +\color{#d53f8c}{\mu_t} = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{constant}}} + \;+\; \underbrace{\color{#319795}{\beta_1 x_{1,t}}}_{\textstyle\color{#319795}{\text{covariate}}} + \;+\; \underbrace{\color{#c05621}{\beta_2 y_{t-1}}}_{\textstyle\color{#c05621}{\text{AR: lag-1}}} +$$ + +- $y_t$ — the observed value at time $t$, drawn from a normal distribution +- $\mu_t$ — the process mean at time $t$: everything the model can explain +- $\sigma^2$ — the variance of the observation noise around that mean +- $c$, $\beta_1 x_{1,t}$, $\beta_2 y_{t-1}$ — exactly the same constant, covariate, and AR terms as the baseline model, just relocated into $\mu_t$ + +**Why bother rewriting it:** splitting "what generates $y_t$" ($\mathcal{N}(\mu_t, \sigma^2)$) from "what determines $\mu_t$" makes it straightforward to swap in a different observation distribution without changing how $\mu_t$ is modeled — which is exactly what happens next. + +## Modeling count data: Poisson error + +Abundance counts are non-negative integers, which a Gaussian error model doesn't respect. Swapping in a Poisson distribution fixes that. + +$$ +y_t = \mathrm{Pois}( + \underbrace{\color{#d53f8c}{\lambda_t}}_{\textstyle\color{#d53f8c}{\text{rate parameter}}} +) +$$ + +$$ +\underbrace{\color{#d69e2e}{\log(\lambda_t)}}_{\textstyle\color{#d69e2e}{\text{log link}}} = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{constant}}} + \;+\; \underbrace{\color{#319795}{\beta_1 x_{1,t}}}_{\textstyle\color{#319795}{\text{covariate}}} + \;+\; \underbrace{\color{#c05621}{\beta_2 y_{t-1}}}_{\textstyle\color{#c05621}{\text{AR: lag-1}}} +$$ + +- $y_t$ — the observed value at time $t$, now constrained to non-negative integers +- $\lambda_t$ — the Poisson rate parameter: both the mean and the variance of the distribution, so it must be positive +- $\log(\lambda_t)$ — the log link: instead of modeling $\lambda_t$ directly, the model predicts $\log(\lambda_t)$, which can be any real number, then exponentiates it back — guaranteeing $\lambda_t$ stays positive +- $c$, $\beta_1 x_{1,t}$, $\beta_2 y_{t-1}$ — the same constant, covariate, and AR terms as before, now predicting $\log(\lambda_t)$ instead of $y_t$ directly + +**A side effect of the log link:** because the model is linear on the log scale, the relationship between the covariate and the untransformed abundance becomes exponential — worth checking with `plot_predictions()` rather than assuming it looks like the linear relationship on the link scale. diff --git a/lessons/R-complex-time-series-models-1/index.qmd b/lessons/R-complex-time-series-models-1/index.qmd index ac5dbbb..4af03bb 100644 --- a/lessons/R-complex-time-series-models-1/index.qmd +++ b/lessons/R-complex-time-series-models-1/index.qmd @@ -6,6 +6,7 @@ listing: contents: - material.qmd - r_tutorial.qmd + - equation_review.qmd type: table fields: - title diff --git a/lessons/R-complex-time-series-models-2/equation_review.qmd b/lessons/R-complex-time-series-models-2/equation_review.qmd new file mode 100644 index 0000000..71f628b --- /dev/null +++ b/lessons/R-complex-time-series-models-2/equation_review.qmd @@ -0,0 +1,71 @@ +--- +title: Annotated Equations +order: 4 +--- + +A quick reference for the models introduced in the [R Tutorial](r_tutorial.qmd). Each equation is broken down term by term so you can check your understanding of what every symbol is doing. + +## Baseline model: AR + covariate + Gaussian error + +The starting point in `mvgam` looks like the ARIMAX models from previous lessons — a covariate, an autoregressive term, and Gaussian error — just written and fit a different way (via `mvgam()`'s Bayesian MCMC fitting, instead of `fable`'s `ARIMA()`). + +$$ +y_t = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{constant}}} + \;+\; \underbrace{\color{#319795}{\beta_1 x_{1,t}}}_{\textstyle\color{#319795}{\text{covariate}}} + \;+\; \underbrace{\color{#c05621}{\beta_2 y_{t-1}}}_{\textstyle\color{#c05621}{\text{AR: lag-1}}} + \;+\; \underbrace{\color{#718096}{\mathcal{N}(0,\sigma^2)}}_{\textstyle\color{#718096}{\text{Gaussian noise}}} +$$ + +- $y_t$ — the observed value at time $t$ (e.g., desert pocket mouse abundance) +- $c$ — a constant +- $\beta_1 x_{1,t}$ — the covariate effect (e.g., minimum temperature) +- $\beta_2 y_{t-1}$ — the AR part: the effect of the previous value +- $\mathcal{N}(0,\sigma^2)$ — Gaussian noise with mean 0 and variance $\sigma^2$ + +**How it's fit:** `mvgam` uses Bayesian methods (MCMC via STAN) rather than the exact-likelihood approach `fable` uses, but for this baseline model the equation itself is the same idea — the AR term is specified separately, via `trend_model = AR(p = 1)`, rather than inside the model formula. + +## The same model, decomposed + +The additive form above can equivalently be written as an explicit observation distribution around a modeled mean — a small notational shift that matters once we start changing the error distribution. + +$$ +y_t = \mathcal{N}( + \underbrace{\color{#d53f8c}{\mu_t}}_{\textstyle\color{#d53f8c}{\text{process mean}}},\sigma^2 +) +$$ + +$$ +\color{#d53f8c}{\mu_t} = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{constant}}} + \;+\; \underbrace{\color{#319795}{\beta_1 x_{1,t}}}_{\textstyle\color{#319795}{\text{covariate}}} + \;+\; \underbrace{\color{#c05621}{\beta_2 y_{t-1}}}_{\textstyle\color{#c05621}{\text{AR: lag-1}}} +$$ + +- $y_t$ — the observed value at time $t$, drawn from a normal distribution +- $\mu_t$ — the process mean at time $t$: everything the model can explain +- $\sigma^2$ — the variance of the observation noise around that mean +- $c$, $\beta_1 x_{1,t}$, $\beta_2 y_{t-1}$ — exactly the same constant, covariate, and AR terms as the baseline model, just relocated into $\mu_t$ + +**Why bother rewriting it:** splitting "what generates $y_t$" ($\mathcal{N}(\mu_t, \sigma^2)$) from "what determines $\mu_t$" makes it straightforward to swap in a different observation distribution without changing how $\mu_t$ is modeled — which is exactly what happens next. + +## Modeling count data: Poisson error + +Abundance counts are non-negative integers, which a Gaussian error model doesn't respect. Swapping in a Poisson distribution fixes that. + +$$ +y_t = \mathrm{Pois}( + \underbrace{\color{#d53f8c}{\lambda_t}}_{\textstyle\color{#d53f8c}{\text{rate parameter}}} +) +$$ + +$$ +\underbrace{\color{#d69e2e}{\log(\lambda_t)}}_{\textstyle\color{#d69e2e}{\text{log link}}} = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{constant}}} + \;+\; \underbrace{\color{#319795}{\beta_1 x_{1,t}}}_{\textstyle\color{#319795}{\text{covariate}}} + \;+\; \underbrace{\color{#c05621}{\beta_2 y_{t-1}}}_{\textstyle\color{#c05621}{\text{AR: lag-1}}} +$$ + +- $y_t$ — the observed value at time $t$, now constrained to non-negative integers +- $\lambda_t$ — the Poisson rate parameter: both the mean and the variance of the distribution, so it must be positive +- $\log(\lambda_t)$ — the log link: instead of modeling $\lambda_t$ directly, the model predicts $\log(\lambda_t)$, which can be any real number, then exponentiates it back — guaranteeing $\lambda_t$ stays positive +- $c$, $\beta_1 x_{1,t}$, $\beta_2 y_{t-1}$ — the same constant, covariate, and AR terms as before, now predicting $\log(\lambda_t)$ instead of $y_t$ directly + +**A side effect of the log link:** because the model is linear on the log scale, the relationship between the covariate and the untransformed abundance becomes exponential — worth checking with `plot_predictions()` rather than assuming it looks like the linear relationship on the link scale. diff --git a/lessons/R-complex-time-series-models-2/index.qmd b/lessons/R-complex-time-series-models-2/index.qmd index 8de8267..f2134d4 100644 --- a/lessons/R-complex-time-series-models-2/index.qmd +++ b/lessons/R-complex-time-series-models-2/index.qmd @@ -6,6 +6,7 @@ listing: contents: - material.qmd - r_tutorial.qmd + - equation_review.qmd type: table fields: - title diff --git a/lessons/R-time-series-modeling-2/equation_review.qmd b/lessons/R-time-series-modeling-2/equation_review.qmd new file mode 100644 index 0000000..4f476ae --- /dev/null +++ b/lessons/R-time-series-modeling-2/equation_review.qmd @@ -0,0 +1,118 @@ +--- +title: Annotated Equations +order: 4 +--- + +A quick reference for the models introduced in the [R Tutorial](r_tutorial.qmd). + +## Moving average (MA) model + +Instead of using past *values* like an AR model, an MA model uses past *errors* to predict the current observation. + +$$ +y_t = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{constant}}} + \;+\; \underbrace{\color{#7c3aed}{\theta_1 \epsilon_{t-1}}}_{\textstyle\color{#7c3aed}{\text{MA: lag-1}}} + \;+\; \underbrace{\color{#718096}{\epsilon_t}}_{\textstyle\color{#718096}{\text{random error}}} +$$ + +- $y_t$ — the observed value at time $t$ +- $c$ — a constant, analogous to the intercept in a regression +- $\theta_1 \epsilon_{t-1}$ — the lag-1 MA effect: $\theta_1$ scales how much last time step's *error* (not value) carries over +- $\epsilon_t$ — random error at time $t$ + +**Reading the coefficient:** if $\theta_1 > 0$, an observation that came in above the mean last time step nudges this time step above the mean too. Positive MA components say an outlier tends to be followed by another outlier in the same direction — unlike an AR model, it's the *surprise*, not the *value*, that carries forward. + +## ARMA model + +AR and MA structure can both be present in the same time series — an ARMA model combines them. + +$$ +y_t = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{constant}}} + \;+\; \underbrace{\color{#c05621}{b_1 y_{t-1}}}_{\textstyle\color{#c05621}{\text{AR: lag-1}}} + \;+\; \underbrace{\color{#7c3aed}{\theta_1 \epsilon_{t-1}}}_{\textstyle\color{#7c3aed}{\text{MA: lag-1}}} + \;+\; \underbrace{\color{#718096}{\epsilon_t}}_{\textstyle\color{#718096}{\text{random error}}} +$$ + +- $y_t$ — the observed value at time $t$ +- $c$ — a constant +- $b_1 y_{t-1}$ — the AR part: the effect of the previous *value* +- $\theta_1 \epsilon_{t-1}$ — the MA part: the effect of the previous *error* +- $\epsilon_t$ — random error at time $t$ + +## Differencing + +ARIMA-family models assume the series is "stationary" — no overall trend. Differencing removes a trend by modeling the *change* from one time step to the next instead of the raw values. + +$$ +y_t' = \underbrace{\color{#38a169}{y_t}}_{\textstyle\color{#38a169}{\text{current value}}} + \;-\; \underbrace{\color{#92400e}{y_{t-1}}}_{\textstyle\color{#92400e}{\text{previous value}}} +$$ + +- $y_t'$ — the differenced series: how much the value changed since last time step +- $y_t$ — the current value +- $y_{t-1}$ — the previous value + +**Why this matters:** once the data is differenced, the model is no longer fit to the raw values but to the *changes* between them, which removes a long-term trend before the AR/MA structure is estimated. + +## A fitted example: MA(3) + +Fitting `ARIMA()` with no structure specified can return a model like this one — a 3rd-order MA model with no differencing and no AR terms. + +$$ +y_t = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{constant}}} + \;+\; \underbrace{ + \color{#7c3aed}{\theta_1 \epsilon_{t-1}} + \color{black}{\,+\,} + \color{#9333ea}{\theta_2 \epsilon_{t-2}} + \color{black}{\,+\,} + \color{#c026d3}{\theta_3 \epsilon_{t-3}} + }_{\textstyle\color{#9333ea}{\text{MA terms: lags 1-3}}} + \;+\; \underbrace{\color{#718096}{\epsilon_t}}_{\textstyle\color{#718096}{\text{random error}}} +$$ + +- $y_t$ — the observed value at time $t$ +- $c$ — a constant +- $\theta_1 \epsilon_{t-1}$, $\theta_2 \epsilon_{t-2}$, $\theta_3 \epsilon_{t-3}$ — the MA(3) terms: this model uses the last three months' errors, shown in the `Coefficients` table as `ma1`, `ma2`, `ma3` +- $\epsilon_t$ — random error at time $t$ + +**Reading it:** if NDVI was above average over the last three months, this model expects it to be above average now too — consistent with the idea that a good year (an MA-style "outlier" in errors) tends to persist for a while. + +## Seasonal AR term + +ARIMA models can also capture a seasonal cycle directly, by using a lag equal to the length of one full cycle — here, 12 months. + +$$ +y_t = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{constant}}} + \;+\; \underbrace{\color{#dd6b20}{b_{12} y_{t-12}}}_{\textstyle\color{#dd6b20}{\text{seasonal effect (1 year back)}}} +$$ + +- $y_t$ — the observed value at time $t$ +- $c$ — a constant +- $b_{12} y_{t-12}$ — the seasonal AR term: how strongly the value from exactly one year ago (`sar1` in the `Coefficients` table) predicts the current value + +**Ecological connection:** it makes sense that greenness one year ago is informative — ecosystems are reliably greener in summer than winter, so "how green was it this time last year" is a good predictor of "how green is it now." + +## Full seasonal ARIMA model + +Like the MA(3) model above but with a seasonal AR term added. + +$$ +y_t = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{constant}}} + \;+\; \underbrace{ + \color{#7c3aed}{\theta_1 \epsilon_{t-1}} + \color{black}{\,+\,} + \color{#9333ea}{\theta_2 \epsilon_{t-2}} + \color{black}{\,+\,} + \color{#c026d3}{\theta_3 \epsilon_{t-3}} + }_{\textstyle\color{#9333ea}{\text{MA terms: lags 1-3}}} + \;+\; \underbrace{\color{#dd6b20}{b_{12} y_{t-12}}}_{\textstyle\color{#dd6b20}{\text{seasonal effect}}} + \;+\; \underbrace{\color{#718096}{\epsilon_t}}_{\textstyle\color{#718096}{\text{random error}}} +$$ + +- $y_t$ — the observed value at time $t$ +- $c$ — a constant +- $\theta_1 \epsilon_{t-1}$, $\theta_2 \epsilon_{t-2}$, $\theta_3 \epsilon_{t-3}$ — the MA(3) terms from the last three months' errors +- $b_{12} y_{t-12}$ — the seasonal AR term from one year back +- $\epsilon_t$ — random error at time $t$ + +**Putting it together:** this model explains the current value using recent surprises (MA), last year's value (seasonal AR), and nothing else — no direct AR terms on the immediately preceding months were needed once the seasonal and MA structure were included. diff --git a/lessons/R-time-series-modeling-2/index.qmd b/lessons/R-time-series-modeling-2/index.qmd index 140b958..a96a031 100644 --- a/lessons/R-time-series-modeling-2/index.qmd +++ b/lessons/R-time-series-modeling-2/index.qmd @@ -6,6 +6,7 @@ listing: contents: - material.qmd - r_tutorial.qmd + - equation_review.qmd type: table fields: - title diff --git a/lessons/R-time-series-modeling-3/equation_review.qmd b/lessons/R-time-series-modeling-3/equation_review.qmd new file mode 100644 index 0000000..f0b78fc --- /dev/null +++ b/lessons/R-time-series-modeling-3/equation_review.qmd @@ -0,0 +1,104 @@ +--- +title: Annotated Equations +order: 4 +--- + +A quick reference for the models introduced in the [R Tutorial](r_tutorial.qmd). + +## Time-series linear model (TSLM) + +A TSLM is similar to a standard regression. + +$$ +y_t = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{constant}}} + \;+\; \underbrace{\color{#319795}{\beta_1 x_t}}_{\textstyle\color{#319795}{\text{covariate}}} + \;+\; \underbrace{\color{#718096}{\epsilon_t}}_{\textstyle\color{#718096}{\text{random error}}} +$$ + +- $y_t$ — the observed value at time $t$ +- $c$ — a constant, the intercept +- $\beta_1 x_t$ — the covariate effect: how strongly the predictor $x_t$ is related to $y_t$ +- $\epsilon_t$ — random error at time $t$ + +**Limitation:** Regression assumes independent observations, so autocorrelation in the residuals limits the model's statistical inferences. + +## Adding trends + +We can explicitly model trends that are not captured by the covariates by fitting them directly. + +$$ +y_t = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{constant}}} + \;+\; \underbrace{\color{#319795}{\beta_1 x_{1,t}}}_{\textstyle\color{#319795}{\text{covariate 1}}} + \;+\; \underbrace{\color{#059669}{\beta_3 t}}_{\textstyle\color{#059669}{\text{trend}}} + \;+\; \underbrace{\color{#718096}{\epsilon_t}}_{\textstyle\color{#718096}{\text{random error}}} +$$ + +- $y_t$ — the observed value at time $t$ +- $c$ — a constant, the intercept +- $\beta_1 x_{1,t}$ — the covariate's effect +- $\beta_3 t$ — the trend term: $t$ is time and serves as a predictor for long-term increases or decreases the other covariates don't capture +- $\epsilon_t$ — random error at time $t$ + +**Reading it:** adding `trend()` doesn't explain *why* there's a long-term change — it just accounts for the fact that a long-term change exists + +## ARIMAX: ARIMA with external predictors + +`ARIMA()` can include exogenous covariates directly, combining regression with AR and MA structure in one model. + +$$ +y_t = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{constant}}} + \;+\; \underbrace{\color{#319795}{\beta_1 x_{1,t}}}_{\textstyle\color{#319795}{\text{covariate}}} + \;+\; \underbrace{\color{#c05621}{\beta_2 y_{t-1}}}_{\textstyle\color{#c05621}{\text{AR: lag-1}}} + \;+\; \underbrace{\color{#7c3aed}{\theta_1 \epsilon_{t-1}}}_{\textstyle\color{#7c3aed}{\text{MA: lag-1}}} + \;+\; \underbrace{\color{#718096}{\epsilon_t}}_{\textstyle\color{#718096}{\text{random error}}} +$$ + +- $y_t$ — the observed value at time $t$ +- $c$ — a constant +- $\beta_1 x_{1,t}$ — the covariate effect +- $\beta_2 y_{t-1}$ — AR: the effect of the previous value +- $\theta_1 \epsilon_{t-1}$ — MA: the effect of the previous error +- $\epsilon_t$ — random error at time $t$ + +**Why this helps:** AR and MA terms capture autocorrelation that a plain TSLM leaves behind, while the covariates allow inclusion of external drivers. + +## The differenced ARIMAX model + +Once `ARIMA()` decides differencing is needed, every term in the model — including the covariate — gets differenced too, and the constant drops out (as in the [differencing section](../R-time-series-modeling-2/equation_review.qmd#differencing) of the previous lesson). + +$$ +y_t' = \underbrace{\color{#319795}{\beta_1 x_{1,t}'}}_{\textstyle\color{#319795}{\text{covariate (differenced)}}} + \;+\; \underbrace{\color{#c05621}{\beta_2 y_{t-1}'}}_{\textstyle\color{#c05621}{\text{AR: lag-1}}} + \;+\; \underbrace{\color{#7c3aed}{\theta_1 \epsilon_{t-1}}}_{\textstyle\color{#7c3aed}{\text{MA: lag-1}}} + \;+\; \underbrace{\color{#718096}{\epsilon_t}}_{\textstyle\color{#718096}{\text{random error}}} +$$ + +- $y_t'$ — the differenced response +- $\beta_1 x_{1,t}'$ — the differenced covariate effect +- $\beta_2 y_{t-1}'$ — AR: the effect of the previous differenced value +- $\theta_1 \epsilon_{t-1}$ — MA: the effect of the previous error +- $\epsilon_t$ — random error at time $t$ + +## Regression with ARIMA errors + +`fable` doesn't literally fit the equations above — it fits a plain regression on the covariate, then models the *leftover* structure in that regression's error with an ARIMA model. + +$$ +y_t = \underbrace{\color{#319795}{\beta_1 x_{1,t}}}_{\textstyle\color{#319795}{\text{covariate effect}}} + \;+\; \underbrace{\color{#d53f8c}{\eta_t}}_{\textstyle\color{#d53f8c}{\text{ARIMA-structured error}}} +$$ + +$$ +\eta_t = \underbrace{\color{#c05621}{\beta_2 \eta_{t-1}}}_{\textstyle\color{#c05621}{\text{AR: lag-1 on the error}}} + \;+\; \underbrace{\color{#7c3aed}{\theta_1 \epsilon_{t-1}}}_{\textstyle\color{#7c3aed}{\text{MA: lag-1}}} + \;+\; \underbrace{\color{#718096}{\epsilon_t}}_{\textstyle\color{#718096}{\text{random error}}} +$$ + +- $y_t$ — the response +- $\beta_1 x_{1,t}$ — the covariate's effect modeled as a linear regression term +- $\eta_t$ — the time-series structured regression error +- $\beta_2 \eta_{t-1}$ — the AR part of that structured error +- $\theta_1 \epsilon_{t-1}$ — the MA part of that structured error +- $\epsilon_t$ — the remaining Normally distributed error left after the time-series components + +**Why split it this way:** it keeps the covariate's coefficient directly interpretable (a plain regression slope), while letting the ARIMA machinery handle whatever autocorrelation the regression alone couldn't. diff --git a/lessons/R-time-series-modeling-3/index.qmd b/lessons/R-time-series-modeling-3/index.qmd index 7c85c80..fcc3f3d 100644 --- a/lessons/R-time-series-modeling-3/index.qmd +++ b/lessons/R-time-series-modeling-3/index.qmd @@ -6,6 +6,7 @@ listing: contents: - material.qmd - r_tutorial.qmd + - equation_review.qmd type: table fields: - title diff --git a/lessons/R-time-series-modeling/equation_review.qmd b/lessons/R-time-series-modeling/equation_review.qmd new file mode 100644 index 0000000..7328919 --- /dev/null +++ b/lessons/R-time-series-modeling/equation_review.qmd @@ -0,0 +1,83 @@ +--- +title: Annotated Equations +order: 4 +--- + +A quick reference for the models introduced in the [R Tutorial](r_tutorial.qmd). + +## White noise model + +The simplest possible time-series model: +every observation is a random number drawn from a normal distribution (aka White Noise) + +$$ +y_t = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{mean}}} + \;+\; \underbrace{\color{#718096}{\epsilon_t}}_{\textstyle\color{#718096}{\text{random error}}} +$$ + +$$ +\epsilon_t \sim \underbrace{\color{#6b46c1}{N}}_{\textstyle\color{#6b46c1}{\text{normal distribution}}}\!\left( + \underbrace{\color{#2c7a7b}{0}}_{\textstyle\color{#2c7a7b}{\text{mean = 0}}},\; \quad + \underbrace{\color{#b7791f}{\sigma^2}}_{\textstyle\color{#b7791f}{\text{variance}}} +\right) +$$ + +- $y_t$ — the observed value at time $t$ (e.g., NDVI in a given month) +- $c$ — a constant: the mean of the observed ($y$) values +- $\epsilon_t$ — the error at time $t$ +- $\sigma^2$ — the variance of the error distribution, i.e., how wide is the Normal distribution we are drawing random numbers from + +**Key assumption:** each $\epsilon_t$ is drawn independently — knowing $\epsilon_{t-1}$ doesn't inform $\epsilon_t$. +The model is just "mean plus noise" + +## AR(1) model + +The last observed value ($y_{t-1}$) influences the current value ($y_t$) + +$$ +y_t = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{constant}}} + \;+\; \underbrace{\color{#c05621}{b_1 y_{t-1}}}_{\textstyle\color{#c05621}{\text{lag-1 effect}}} + \;+\; \underbrace{\color{#718096}{\epsilon_t}}_{\textstyle\color{#718096}{\text{random error}}} +$$ + +$$ +\epsilon_t \sim \underbrace{\color{#6b46c1}{N}}_{\textstyle\color{#6b46c1}{\text{normal distribution}}}\!\left( + \underbrace{\color{#2c7a7b}{0}}_{\textstyle\color{#2c7a7b}{\text{mean = 0}}},\; \quad + \underbrace{\color{#b7791f}{\sigma^2}}_{\textstyle\color{#b7791f}{\text{variance}}} +\right) +$$ + +- $y_t$ — the observed value at time $t$ +- $c$ — a constant, analogous to the intercept in a regression +- $b_1 y_{t-1}$ — the lag-1 effect: $b_1$ is the AR1 coefficient, determining how strongly (and in which direction) the previous value, $y_{t-1}$, influences $y_t$ +- $\epsilon_t$ — random error at time $t$, independent of past errors + +**Reading the coefficient:** if $b_1$ is large and positive, a high value last time step predicts a high value this time step (values persist). If $b_1$ were negative, high values would tend to be followed by low values. + +**Ecological connection:** if $y$ is $\log(N)$, this is essentially a Gompertz population model — the current abundance depends on the previous abundance. + +## AR(2) model + +Like the AR(1) model but with 2 lags. + +$$ +y_t = \underbrace{\color{#2b6cb0}{c}}_{\textstyle\color{#2b6cb0}{\text{constant}}} + \;+\; \underbrace{\color{#c05621}{b_1 y_{t-1}}}_{\textstyle\color{#c05621}{\text{lag-1 effect}}} + \;+\; \underbrace{\color{#2f855a}{b_2 y_{t-2}}}_{\textstyle\color{#2f855a}{\text{lag-2 effect}}} + \;+\; \underbrace{\color{#718096}{\epsilon_t}}_{\textstyle\color{#718096}{\text{random error}}} +$$ + +$$ +\epsilon_t \sim \underbrace{\color{#6b46c1}{N}}_{\textstyle\color{#6b46c1}{\text{normal distribution}}}\!\left( + \underbrace{\color{#2c7a7b}{0}}_{\textstyle\color{#2c7a7b}{\text{mean = 0}}},\; \quad + \underbrace{\color{#b7791f}{\sigma^2}}_{\textstyle\color{#b7791f}{\text{variance}}} +\right) +$$ + +- $y_t$ — the observed value at time $t$ +- $c$ — a constant, analogous to the intercept in a regression +- $b_1$ — the AR1 coefficient: the effect of the value one time step back ($y_{t-1}$) +- $b_2$ — the AR2 coefficient: the effect of the value two time steps back ($y_{t-2}$) +- $\epsilon_t$ — random error at time $t$, independent of past errors + +**Reading the coefficients:** $b_1$ and $b_2$ don't have to point the same direction. For example, a large positive $b_1$ combined with a smaller negative $b_2$ means a high value last time step predicts a high value now, while a high value two time steps back predicts a slightly *lower* value now — the two lags can partially offset each other. diff --git a/lessons/R-time-series-modeling/index.qmd b/lessons/R-time-series-modeling/index.qmd index 4543a99..ca27bed 100644 --- a/lessons/R-time-series-modeling/index.qmd +++ b/lessons/R-time-series-modeling/index.qmd @@ -6,6 +6,7 @@ listing: contents: - material.qmd - r_tutorial.qmd + - equation_review.qmd type: table fields: - title