Article Views: 25
Accurate forecasting of mean sea level (MSL) is critical for coastal risk management, infrastructure planning, and climate adaptation, particularly along the low-lying Mid-Atlantic coast of the United States. This study presents a comprehensive comparative evaluation of four data-driven deep learning and neural network architectures, including Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), one-dimensional Convolutional Neural Network (CNN), and a lagged feed-forward neural network (FFNN), applied to seasonally adjusted monthly MSL records at the NOAA tide gauge station in Ocean City, Maryland (Station 8570283), spanning August 2002 through February 2025. All models share a standardized 24-month lookback window (lags 1–13 for the FFNN), 80/20 train/test split, Adam optimization, and early stopping, enabling controlled comparison. Models are evaluated using Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), and the coefficient of determination (R²) on held-out test data. Results indicate that the GRU achieves the best predictive performance among the neural architectures (RMSE = 0.0704 m, MAE = 0.0587 m, R² = 0.099), followed by the LSTM (RMSE = 0.0734 m, R² = 0.020). Even the GRU explained only about 10% of the test-period variance, and its RMSE was only about 3% lower than that of a persistence forecast. The CNN and FFNN both yield negative R² values (−0.219 and −0.907 respectively), which on this single test period placed the recurrent models ahead of the convolutional and feed-forward networks; this ranking was not stable across test periods (see below). A 24-month ensemble forecast is presented for all models through February 2027. The FFNN ensemble spread (±1σ across 20 independently initialized networks) is reported only as a measure of initialization variability; calibrated prediction intervals are derived from the ARIMA model and from rolling-origin forecast errors. When benchmarked under the identical protocol against persistence, seasonal-naïve, random walk with drift, linear trend, ARIMA, SARIMA and exponential smoothing (ETS) baselines, however, all three statistical time-series models outperformed every neural architecture: ARIMA(3,0,0) with a linear trend achieved the lowest test error (RMSE = 0.0642 m, R² = 0.250), and the GRU improved on a simple persistence forecast by only 3% in RMSE. These findings indicate that, for this univariate monthly record, the neural architectures tested did not add predictive value beyond well-specified linear time-series models; whether this holds at other stations has not yet been tested. An expandingwindow rolling-origin evaluation over 49 forecast origins (December 2012 – December 2024), with every model re-trained or re-estimated at each origin, confirmed this ranking at forecast horizons of 1, 3, 6, 12 and 24 months: ARIMA had the lowest or near-lowest RMSE at every horizon (0.056–0.069 m), whereas the best neural model was 6–15% less accurate, and the single-split ranking of the neural architectures was not stable across test periods. All models predict the seasonally adjusted monthly MSL level in meters; training the networks on residuals from a training-period trend improved their accuracy and stabilised their forecasts but did not change their ranking relative to the statistical models. The NNAR model of R’s nnetar(), added as a separate benchmark, was far more accurate than the FFNN on the single split (RMSE = 0.0708 m) but was unstable at multi-step horizons and did not outperform ARIMA. Backtesting of the recursive multi-step forecasts showed that the neural models increasingly under-forecast with lead time, and that ARIMA prediction intervals were close to nominal coverage, whereas the FFNN ensemble spread and the simulated nnetar() intervals were far too narrow to serve as prediction intervals.
mean sea level; time series forecasting; deep learning; NNAR; LSTM; GRU; CNN; Ocean City; Mid-Atlantic