<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">ASCMO</journal-id><journal-title-group>
    <journal-title>Advances in Statistical Climatology, Meteorology and Oceanography</journal-title>
    <abbrev-journal-title abbrev-type="publisher">ASCMO</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Adv. Stat. Clim. Meteorol. Oceanogr.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">2364-3587</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/ascmo-12-221-2026</article-id><title-group><article-title>Comparative evaluation of statistical and deep learning methods for high-frequency radar surface current forecasting in a narrow tropical strait</article-title><alt-title>Statistical and deep learning for HF radar current forecasting</alt-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Gautama</surname><given-names>Dhava</given-names></name>
          <email>dhava.gautama@bmkg.go.id</email>
        <ext-link>https://orcid.org/0009-0004-8578-8470</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1 aff2">
          <name><surname>Putra</surname><given-names>Alifficionaldo Agpri</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>Indonesia Meteorology, Climatology, and Geophysical Agency (BMKG), Jakarta, Indonesia</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Department of Meteorology, University of Reading, Reading, United Kingdom</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Dhava Gautama (dhava.gautama@bmkg.go.id)</corresp></author-notes><pub-date><day>15</day><month>September</month><year>2026</year></pub-date>
      
      <volume>12</volume>
      <issue>2</issue>
      <fpage>221</fpage><lpage>242</lpage>
      <history>
        <date date-type="received"><day>7</day><month>April</month><year>2026</year></date>
           <date date-type="rev-recd"><day>1</day><month>August</month><year>2026</year></date>
           <date date-type="accepted"><day>7</day><month>September</month><year>2026</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2026 Dhava Gautama</copyright-statement>
        <copyright-year>2026</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026.html">This article is available from https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026.html</self-uri><self-uri xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026.pdf">The full text article is available as a PDF file from https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d2e92">Accurate prediction of ocean surface currents is essential for maritime safety in narrow waterways with strong tidal forcing. This study presents a three-phase evaluation of forecasting methods for high-frequency (HF) radar surface currents in the Sunda Strait, Indonesia, where tidal currents regularly exceed 150 cm s<sup>−1</sup>. Phase 1 compares twelve methods – persistence, tidal harmonics, classical time series, a reduced-rank spatio-temporal statistical model, shallow machine learning, and deep learning – for one-step-ahead prediction; evaluated consistently over all 291 grid cells and the full test period, the deep learning models achieve the lowest root-mean-square errors (RMSE) for both components, with the convolutional neural network (CNN) reaching 11.31 cm s<sup>−1</sup> for the zonal component (skill score 0.50 relative to persistence) and the convolutional neural network–gated recurrent unit (CNN-GRU) reaching 15.44 cm s<sup>−1</sup> for the meridional component (skill score 0.42); the pointwise classical statistical models (autoregressive integrated moving average, ARIMA; exponential smoothing) fall well below, while an empirical-orthogonal-function vector autoregression (EOF-VAR) that models spatial correlation closes most of the gap to deep learning, indicating that the neglect of spatial structure – not of nonlinearity – is the principal limitation of the classical baselines. Phase 2 examines lookback window length (<inline-formula><mml:math id="M4" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>, 6, 12 h); CNN-gated recurrent unit (CNN-GRU) improves monotonically with longer lookback, reaching 11.21 and 15.41 cm s<sup>−1</sup> at <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">12</mml:mn></mml:mrow></mml:math></inline-formula>, while standalone CNN and GRU show no benefit. Phase 3 extends prediction to six-hour nowcasting using direct multi-step and autoregressive (ConvLSTM-ED, BiEF) architectures; under a controlled comparison (five seeds, identical training budget on one GPU), accuracy is governed by model capacity and plateaus near 1 M parameters (the 1.00 M CNN-GRU-MS-Small reaching 18.39 and 22.07 cm s<sup>−1</sup>; skill scores 0.77 and 0.72), and at matched capacity the direct and autoregressive architectures are statistically indistinguishable. The direct multi-step model is nonetheless preferred for operational nowcasting because it trains <inline-formula><mml:math id="M8" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">2.5</mml:mn></mml:mrow></mml:math></inline-formula> times faster and runs <inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">4.5</mml:mn></mml:mrow></mml:math></inline-formula> times faster at inference (a single forward pass versus sequential decoding). Diurnal error analysis, corroborated by concurrent wind observations from three coastal weather stations, reveals that zonal prediction errors correlate positively with afternoon sea-breeze wind enhancement (Spearman <inline-formula><mml:math id="M10" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.48</mml:mn></mml:mrow></mml:math></inline-formula> across the 24 diurnal-mean hours), while meridional errors follow a distinct diurnal pattern not explained by wind forcing. Relative errors of 3 %–7 % of the observed speed range are of the same order as those reported in calmer environments – though this normalisation partly reflects the strait's very large speed range – suggesting that deep learning does not break down in energetic strait settings.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>Lembaga Pengelola Dana Pendidikan</funding-source>
<award-id>2025061211202566</award-id>
</award-group>
</funding-group>
</article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d2e230">The prediction of ocean surface currents is a fundamental challenge in operational oceanography, with applications spanning maritime navigation, search-and-rescue operations, oil spill trajectory forecasting, and coastal environmental management <xref ref-type="bibr" rid="bib1.bibx1" id="paren.1"/>. In narrow strait environments, where tidal amplification can produce surface currents exceeding 150 cm s<sup>−1</sup> <xref ref-type="bibr" rid="bib1.bibx30" id="paren.2"/>, accurate short-term forecasts are especially critical for shipping safety and port operations.</p>
      <p id="d2e251">High-frequency (HF) radar provides spatially continuous, real-time surface current measurements at hourly or sub-hourly temporal resolution <xref ref-type="bibr" rid="bib1.bibx24 bib1.bibx26" id="paren.3"/>. The resulting spatiotemporal data present a natural prediction problem: given the observed current field over a lookback window of <inline-formula><mml:math id="M12" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula> timesteps, forecast the field one or more steps ahead. This problem admits solutions from a hierarchy of statistical and computational methods, ranging from simple persistence and autoregressive models to deep neural networks designed for spatiotemporal sequence prediction.</p>
      <p id="d2e264">Several studies have applied deep learning to HF radar current prediction. <xref ref-type="bibr" rid="bib1.bibx32" id="text.4"/> introduced a CNN-GRU architecture combining convolutional neural networks (CNN) for spatial feature extraction with gated recurrent units (GRU; <xref ref-type="bibr" rid="bib1.bibx5" id="altparen.5"/>) for temporal modelling, reporting RMSE improvements of 11 % and 27 % over standalone models in the Gulf of Thailand. Follow-up work incorporated attention mechanisms <xref ref-type="bibr" rid="bib1.bibx31" id="paren.6"/>, and <xref ref-type="bibr" rid="bib1.bibx4" id="text.7"/> proposed a spatiotemporal attention-based GRU. For multi-step forecasting, convolutional long short-term memory (ConvLSTM) encoder–decoder architectures <xref ref-type="bibr" rid="bib1.bibx28" id="paren.8"/>, which extend the long short-term memory cell <xref ref-type="bibr" rid="bib1.bibx15" id="paren.9"/> with convolutional state transitions, have been applied to sea surface temperature prediction <xref ref-type="bibr" rid="bib1.bibx36 bib1.bibx37 bib1.bibx33" id="paren.10"/> and four-dimensional ocean temperature forecasting <xref ref-type="bibr" rid="bib1.bibx21" id="paren.11"/>, but their application to surface current nowcasting remains limited. <xref ref-type="bibr" rid="bib1.bibx39" id="text.12"/> demonstrated a residual-learning CNN with attention for offshore current field forecasting, and <xref ref-type="bibr" rid="bib1.bibx38" id="text.13"/> presented a physics-informed model for surface current prediction. Earlier neural network approaches include <xref ref-type="bibr" rid="bib1.bibx17" id="text.14"/> and <xref ref-type="bibr" rid="bib1.bibx25" id="text.15"/>, who applied feedforward and recurrent networks to coastal currents. Recent advances include deep learning for ocean current estimation from satellite altimetry and sea surface temperature <xref ref-type="bibr" rid="bib1.bibx6" id="paren.16"/> and end-to-end neural global ocean forecasting <xref ref-type="bibr" rid="bib1.bibx11" id="paren.17"/>. Reviews by <xref ref-type="bibr" rid="bib1.bibx10" id="text.18"/> and <xref ref-type="bibr" rid="bib1.bibx40" id="text.19"/> document the growing role of deep learning in physical oceanography.</p>
      <p id="d2e317">Despite these advances, several gaps remain. First, most studies have been conducted in open-sea environments where current speeds are moderate (typically below 50 cm s<sup>−1</sup>); performance in energetic strait environments has not been systematically evaluated. Second, the relationship between lookback window length and prediction skill – particularly the role of tidal cycle coverage – has not been explicitly investigated. Third, the transition from one-step to multi-step prediction introduces a design choice between autoregressive and direct prediction strategies, the trade-offs of which have not been studied for HF radar currents in tidally dominated settings.</p>
      <p id="d2e333">This study addresses these gaps through a three-phase evaluation using over three years of hourly HF radar data from the Sunda Strait, Indonesia: <list list-type="order"><list-item>
      <p id="d2e338">Phase 1 compares twelve forecasting methods for one-step-ahead prediction, including UTide tidal harmonic predictions and a reduced-rank spatio-temporal statistical model (EOF-VAR) alongside persistence;</p></list-item><list-item>
      <p id="d2e342">Phase 2 examines the sensitivity of the three best deep learning architectures to lookback window length (<inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>, 6, 12 h);</p></list-item><list-item>
      <p id="d2e360">Phase 3 extends prediction to six-hour nowcasting using ConvLSTM encoder–decoder (ConvLSTM-ED), bidirectional ConvLSTM encoder–forecaster (BiEF), and direct multi-step CNN-GRU (CNN-GRU-MS).</p></list-item></list></p>
      <p id="d2e363">The main contributions are: (i) a comprehensive benchmark spanning persistence, classical and spatio-temporal statistical models, and deep learning (twelve methods) for HF radar currents in a narrow tropical strait with strong tidal forcing; (ii) demonstration that tidal cycle coverage in the lookback window is the primary driver of skill improvement for hybrid architectures; (iii) a controlled multi-seed, capacity-matched comparison showing that direct and autoregressive multi-step architectures achieve comparable accuracy, with the direct model preferred operationally for its <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">2.5</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math></inline-formula> faster training and <inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">4.5</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math></inline-formula> faster single-pass inference; and (iv) diurnal and seasonal error analysis linked to concurrent meteorological observations.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Study area and data</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>The Sunda Strait</title>
      <p id="d2e405">The Sunda Strait connects the Java Sea to the Indian Ocean between the islands of Java and Sumatra, Indonesia (Fig. <xref ref-type="fig" rid="F1"/>). The strait is approximately 24 km wide at its narrowest point and experiences strong tidal forcing dominated by the semidiurnal M2 constituent (<inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">12.42</mml:mn></mml:mrow></mml:math></inline-formula> h period). Surface currents regularly exceed 150 cm s<sup>−1</sup> <xref ref-type="bibr" rid="bib1.bibx30" id="paren.20"/>. The circulation is modulated by two monsoon seasons: the northwest (NW) monsoon (December–February) and the southeast (SE) monsoon (June–August) <xref ref-type="bibr" rid="bib1.bibx35 bib1.bibx20" id="paren.21"/>. The strait serves as a secondary pathway for the Indonesian Throughflow (ITF) <xref ref-type="bibr" rid="bib1.bibx29" id="paren.22"/>. <xref ref-type="bibr" rid="bib1.bibx22" id="text.23"/> previously demonstrated the value of HF radar observations in this region.</p>

      <fig id="F1" specific-use="star"><label>Figure 1</label><caption><p id="d2e449"><bold>(a)</bold> Regional map showing the Sunda Strait between Java and Sumatra, Indonesia (map data: Natural Earth). <bold>(b)</bold> HF radar grid (<inline-formula><mml:math id="M19" display="inline"><mml:mrow><mml:mn mathvariant="normal">21</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">21</mml:mn></mml:mrow></mml:math></inline-formula>) showing 291 observed sea points (blue), 137 blind-zone sea points (orange), and 13 land points (brown) at approximately 1 km spacing.</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f01.png"/>

        </fig>

</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>HF radar observations and quality control</title>
      <p id="d2e483">The BADA (Bidirectional Array Direction of Arrival) HF radar system, a Coastal Ocean Dynamics Applications Radar (CODAR) type direction-finding radar operating at 13.5 MHz, is operated by the Indonesia Meteorology, Climatology, and Geophysical Agency (BMKG). It provides hourly measurements of the eastward (zonal, <inline-formula><mml:math id="M20" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>) and northward (meridional, <inline-formula><mml:math id="M21" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>) surface current components across 291 active grid points on a <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:mn mathvariant="normal">21</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">21</mml:mn></mml:mrow></mml:math></inline-formula> spatial grid (<inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> km resolution). The dataset spans December 2022–February 2026, providing 28 321 hourly timesteps.</p>
      <p id="d2e522">Quality control follows established HF radar protocols <xref ref-type="bibr" rid="bib1.bibx14 bib1.bibx26" id="paren.24"/>: (1) speed thresholding at 200 cm s<sup>−1</sup>, (2) temporal consistency checks requiring less than 100 cm s<sup>−1</sup> change between consecutive hours, and (3) spatial median absolute deviation (MAD) filtering with a factor of three. After quality control, temporal coverage is approximately 85 %. Missing values at each grid point are filled using tidal harmonic predictions from UTide <xref ref-type="bibr" rid="bib1.bibx7" id="paren.25"/> fitted to the available record at that point, ensuring gap-free lookback windows for all models.</p>
      <p id="d2e555">The data are split chronologically: training (December 2022–December 2024; 18 216 steps), validation (January–June 2025; 4344 steps), and test (July 2025–February 2026; 5761 steps). The training set spans approximately two years, covering at least three complete monsoon seasons (SE 2023, NW 2023/24, SE 2024) and portions of the NW monsoon at both ends of the period.</p>
      <p id="d2e558">For diurnal error analysis, concurrent wind observations are obtained from three BMKG automatic weather stations (AWS) at Merak, Ciwandan, and Bakauheni ports surrounding the Sunda Strait, providing 1 min wind speed and direction over January 2025–April 2026, of which the July 2025–February 2026 test period is used for the diurnal error analysis. These are averaged to hourly resolution with basic quality control (removal of wind speeds exceeding 40 m s<sup>−1</sup>).</p>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Methods</title>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Problem formulation</title>
      <p id="d2e589">The prediction task is formulated at two levels. For one-step prediction (Phases 1–2), given the observed current field over a lookback window of <inline-formula><mml:math id="M27" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula> timesteps, the goal is to predict the field at <inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>:

            <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M29" display="block"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold">X</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi>f</mml:mi><mml:mo mathsize="1.1em">(</mml:mo><mml:mi mathvariant="bold">X</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="bold">X</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">X</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>T</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo><mml:mo mathsize="1.1em">)</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:mi mathvariant="bold">X</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">21</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">21</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> contains the <inline-formula><mml:math id="M31" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M32" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> fields at time <inline-formula><mml:math id="M33" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>. Of the <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:mn mathvariant="normal">21</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">21</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">441</mml:mn></mml:mrow></mml:math></inline-formula> grid cells, 291 are observed sea points, 137 are coastal blind-zone sea points (outside the radar footprint), and 13 are land; blind-zone and land cells are set to zero in the input and excluded from the loss function and evaluation metrics.</p>
      <p id="d2e759">For multi-step nowcasting (Phase 3), the prediction extends over a horizon <inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:mi>H</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:math></inline-formula> h:

            <disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M36" display="block"><mml:mrow><mml:mo mathsize="1.1em">[</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold">X</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold">X</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>H</mml:mi><mml:mo>)</mml:mo><mml:mo mathsize="1.1em">]</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi>g</mml:mi><mml:mo mathsize="1.1em">(</mml:mo><mml:mi mathvariant="bold">X</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">X</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>T</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo><mml:mo mathsize="1.1em">)</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e864">The neural-network inputs and targets are normalised to <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> using min–max scaling computed on the training set only; the training-set extremes are <inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:mi>U</mml:mi><mml:mo>∈</mml:mo><mml:mo>[</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">185.4</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">153.3</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup> and <inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:mi>V</mml:mi><mml:mo>∈</mml:mo><mml:mo>[</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">197.3</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">196.4</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup>. This <inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> scaling is applied <italic>only</italic> to the neural networks: the classical statistical baselines (UTide, moving average, exponential smoothing, and ARIMA) and the EOF-VAR model are fitted directly to the raw current series in cm s<sup>−1</sup> (Sect. <xref ref-type="sec" rid="Ch1.S3"/>), so the ARIMA error model in particular operates on the unbounded real line rather than on bounded <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> data. All neural networks use mean squared error (MSE) over the 291 observed sea cells as the loss function and default PyTorch weight initialisation (Kaiming uniform for convolutional layers, uniform for linear and recurrent layers).</p>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Phase 1: one-step prediction methods</title>
      <p id="d2e1009">Twelve methods are compared, spanning six categories (Table <xref ref-type="table" rid="T1"/>). All methods use a default lookback of <inline-formula><mml:math id="M45" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>. The methods were chosen to span the full ladder relevant to operational HF-radar nowcasting – a naive persistence reference, a deterministic-physical baseline (tidal harmonics), classical univariate time-series models, a reduced-rank spatio-temporal statistical model that accounts for spatial correlation (EOF-VAR), shallow machine learning, and spatio-temporal deep learning – and to include the specific architectures previously applied to HF-radar currents (CNN-GRU; ConvLSTM, Phase 3) so that our energetic-strait results are directly comparable to earlier open-sea studies. Approaches deliberately left for future work include attention/transformer sequence models and graph neural networks.</p>

<table-wrap id="T1" specific-use="star"><label>Table 1</label><caption><p id="d2e1031">Summary of all forecasting methods evaluated in this study, grouped by category and phase.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Category</oasis:entry>
         <oasis:entry colname="col2">Method</oasis:entry>
         <oasis:entry colname="col3">Phase</oasis:entry>
         <oasis:entry colname="col4">Params</oasis:entry>
         <oasis:entry colname="col5">Description</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Baseline</oasis:entry>
         <oasis:entry colname="col2">Persistence</oasis:entry>
         <oasis:entry colname="col3">1, 3</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold">X</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="bold">X</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">UTide</oasis:entry>
         <oasis:entry colname="col3">1</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
         <oasis:entry colname="col5">Tidal harmonic prediction</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Classical</oasis:entry>
         <oasis:entry colname="col2">Moving average</oasis:entry>
         <oasis:entry colname="col3">1</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
         <oasis:entry colname="col5">Mean of <inline-formula><mml:math id="M47" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula> lookback frames</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">time series</oasis:entry>
         <oasis:entry colname="col2">Exp. smoothing (SES)</oasis:entry>
         <oasis:entry colname="col3">1</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
         <oasis:entry colname="col5">Per-cell, optimised level</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">ARIMA</oasis:entry>
         <oasis:entry colname="col3">1</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
         <oasis:entry colname="col5">Per-cell, low-order: (1,0,1) and (2,0,2)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Spatio-temporal</oasis:entry>
         <oasis:entry colname="col2">EOF-VAR</oasis:entry>
         <oasis:entry colname="col3">1</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
         <oasis:entry colname="col5">VAR on leading EOFs (reduced-rank DSTM)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">statistical</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Shallow</oasis:entry>
         <oasis:entry colname="col2">Temporal <inline-formula><mml:math id="M48" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>NN</oasis:entry>
         <oasis:entry colname="col3">1</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> neighbours in lookback space</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">machine learning</oasis:entry>
         <oasis:entry colname="col2">Perceptron</oasis:entry>
         <oasis:entry colname="col3">1</oasis:entry>
         <oasis:entry colname="col4">1.81 M</oasis:entry>
         <oasis:entry colname="col5">Single layer, 512 units</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">MLP (multilayer perceptron)</oasis:entry>
         <oasis:entry colname="col3">1</oasis:entry>
         <oasis:entry colname="col4">0.97 M</oasis:entry>
         <oasis:entry colname="col5">Two hidden layers, 256 units each</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Deep learning</oasis:entry>
         <oasis:entry colname="col2">CNN</oasis:entry>
         <oasis:entry colname="col3">1, 2</oasis:entry>
         <oasis:entry colname="col4">1.34 M</oasis:entry>
         <oasis:entry colname="col5">2 <inline-formula><mml:math id="M50" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> Conv2d(32) <inline-formula><mml:math id="M51" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> 2 <inline-formula><mml:math id="M52" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> Conv2d(64)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">(one step)</oasis:entry>
         <oasis:entry colname="col2">GRU</oasis:entry>
         <oasis:entry colname="col3">1, 2</oasis:entry>
         <oasis:entry colname="col4">1.10 M</oasis:entry>
         <oasis:entry colname="col5">256 hidden units, one layer</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">CNN-GRU</oasis:entry>
         <oasis:entry colname="col3">1, 2</oasis:entry>
         <oasis:entry colname="col4">1.72 M</oasis:entry>
         <oasis:entry colname="col5">CNN encoder <inline-formula><mml:math id="M53" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula> GRU <inline-formula><mml:math id="M54" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula> fully connected</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Deep learning</oasis:entry>
         <oasis:entry colname="col2">ConvLSTM-ED</oasis:entry>
         <oasis:entry colname="col3">3</oasis:entry>
         <oasis:entry colname="col4">0.30 M</oasis:entry>
         <oasis:entry colname="col5">ConvLSTM encoder–decoder</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">(multi-step)</oasis:entry>
         <oasis:entry colname="col2">BiEF</oasis:entry>
         <oasis:entry colname="col3">3</oasis:entry>
         <oasis:entry colname="col4">0.47 M</oasis:entry>
         <oasis:entry colname="col5">Bidirectional ConvLSTM encoder–forecaster</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">CNN-GRU-MS</oasis:entry>
         <oasis:entry colname="col3">3</oasis:entry>
         <oasis:entry colname="col4">2.85 M</oasis:entry>
         <oasis:entry colname="col5">Direct multi-step CNN-GRU</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">CNN-GRU-MS-Small</oasis:entry>
         <oasis:entry colname="col3">3</oasis:entry>
         <oasis:entry colname="col4">1.00 M</oasis:entry>
         <oasis:entry colname="col5">Reduced GRU (90 units)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">CNN-GRU-MS-Matched</oasis:entry>
         <oasis:entry colname="col3">3</oasis:entry>
         <oasis:entry colname="col4">0.48 M</oasis:entry>
         <oasis:entry colname="col5">Reduced GRU (40 units), matched to BiEF</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e1484">The <italic>persistence</italic> model predicts the most recent observation: <inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold">X</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="bold">X</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The <italic>UTide</italic> baseline predicts using tidal harmonic analysis <xref ref-type="bibr" rid="bib1.bibx7" id="paren.26"/> fitted independently at each grid point using eight major constituents (M2, S2, K1, O1, N2, K2, P1, M4), representing the deterministic tidal component. The <italic>moving average</italic> predicts the mean of the <inline-formula><mml:math id="M56" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula> lookback frames, effectively a low-pass filter that smooths both signal and noise. <italic>Simple exponential smoothing</italic> (SES) forecasts the next value as an exponentially weighted average of the past, with the smoothing level optimised per cell on the training set. <italic>ARIMA</italic> <xref ref-type="bibr" rid="bib1.bibx3" id="paren.27"/> is fitted independently to each grid cell on the raw current series (cm s<sup>−1</sup>); only the neural networks use the <inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> min–max scaling, so the ARIMA error model operates on the unbounded real line. Because min–max scaling is an invertible affine transform under which ARIMA is equivariant, this choice does not affect the reported errors – it only keeps the Gaussian error model on the natural unbounded scale. The AR term captures the strong one-step autocorrelation of tidal currents and the MA term smooths observation noise. No differencing (<inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:mi>d</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>) is applied: augmented Dickey–Fuller tests reject the unit root in all 30 cells of a representative sample for both components (median <inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">2.4</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> for <inline-formula><mml:math id="M61" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:mn mathvariant="normal">7.0</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">7</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> for <inline-formula><mml:math id="M63" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>), confirming stationarity. The sample autocorrelation and partial-autocorrelation functions (Fig. <xref ref-type="fig" rid="F2"/>) show the strongly oscillatory structure of a deterministic semidiurnal tide – not the rapidly decaying signature of an autocorrelated-noise process – with significant partial autocorrelations extending to lags well beyond those an ARMA(1,1) can represent. Evaluating AIC across low orders <inline-formula><mml:math id="M64" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>d</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi><mml:mo>)</mml:mo><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:msup><mml:mo mathvariant="italic">}</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>×</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula> favours (2,0,2) in 54 of 60 cell-series. Because this optimum lies on the boundary of the search set, we extended the search to <inline-formula><mml:math id="M65" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula> (<inline-formula><mml:math id="M66" display="inline"><mml:mrow><mml:mi>d</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>): the AIC optimum then moves beyond the original boundary in <italic>all</italic> 60 cell-series – most frequently (3,0,4) (28 series) and (4,0,4) (14 series), again often on the boundary of the extended set – and adding a seasonal component at the 12 h semidiurnal period reduces the AIC further in all 20 series tested. This boundary behaviour confirms that no finite low-order ARMA representation is adequate for these series, consistent with the quasi-deterministic tidal harmonics visible in Fig. <xref ref-type="fig" rid="F2"/>, and the selected order varies from cell to cell, so no single higher order would be defensible as a uniform specification. Crucially, however, the higher-order and seasonal specifications do not materially improve <italic>out-of-sample</italic> forecasts: evaluated over the full 5761-step test period on the same cell sample, the per-series AIC-best orders (<inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula>) reduce the one-step RMSE relative to (2,0,2) by a median of only 0.29 cm s<sup>−1</sup> (<inline-formula><mml:math id="M69" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> %) while the mean change is a slight <italic>degradation</italic> of 0.18 cm s<sup>−1</sup> (individual cells improve or degrade), and the best seasonal specification improves the mean RMSE by only <inline-formula><mml:math id="M71" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">0.6</mml:mn></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup> (<inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula> %). We therefore retain ARIMA as an intentionally <italic>low-order empirical baseline</italic>: it is a data-driven model that encodes no tidal physics (the deterministic tidal prior is instead carried by the UTide baseline), and it is not expected to capture the underlying physical process in full. We report both the parsimonious ARIMA(1,0,1) and the low-order AIC-optimal ARIMA(2,0,2), and the comparison with the spatio-temporal models should be read with this qualification in mind. Ljung–Box tests show that residuals retain significant autocorrelation at all orders examined, including the seasonal specifications (<inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> in essentially all cells; excess kurtosis <inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:math></inline-formula>, so the Gaussian likelihood acts as a quasi-maximum-likelihood estimator of the conditional mean), reflecting tidal harmonic structure that a per-cell low-order ARMA cannot fully capture. ARIMA, SES, and temporal <inline-formula><mml:math id="M76" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>NN are evaluated on all 291 sea cells over the full test period; fitting ARIMA at all 291 cells is computationally inexpensive (about 17 min in total, far less than the time required to train the neural networks). The <italic>temporal kNN</italic> <xref ref-type="bibr" rid="bib1.bibx16" id="paren.28"/> (<inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula>) fits a per-cell <inline-formula><mml:math id="M78" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-nearest-neighbours regressor on the lookback feature vector (Euclidean distance) and predicts the mean of the five nearest neighbours. As per-cell univariate models, the moving average, SES, ARIMA, and <inline-formula><mml:math id="M79" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>NN have no spatial coupling and are therefore confined to Phase 1; the lookback-sensitivity (Phase 2) and multi-step spatio-temporal (Phase 3) experiments concern the joint spatial–temporal modelling these methods cannot perform.</p>

      <fig id="F2" specific-use="star"><label>Figure 2</label><caption><p id="d2e1899">Sample autocorrelation (ACF) and partial autocorrelation (PACF) of the training-period <inline-formula><mml:math id="M80" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M81" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> series at a representative cell. The oscillatory ACF reflects the deterministic semidiurnal tide; the PACF retains significant terms to lags far beyond the reach of an ARMA(1,1), explaining why AIC pushes the selected order to the boundary of the search set and why ARIMA residuals remain autocorrelated at all orders examined. Red dashed lines mark the <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">1.96</mml:mn><mml:mo>/</mml:mo><mml:msqrt><mml:mi>n</mml:mi></mml:msqrt></mml:mrow></mml:math></inline-formula> significance band.</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f02.png"/>

        </fig>

      <p id="d2e1937">To provide a <italic>statistical</italic> model that does account for spatial correlation – bridging the pointwise time-series methods and the spatio-temporal neural networks – we include an empirical-orthogonal-function vector autoregression (<italic>EOF-VAR</italic>), a standard reduced-rank dynamic spatio-temporal model <xref ref-type="bibr" rid="bib1.bibx8 bib1.bibx34" id="paren.29"/>. The combined <inline-formula><mml:math id="M83" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi>U</mml:mi><mml:mo>,</mml:mo><mml:mi>V</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> field over the 291 sea cells is centred on the training mean and decomposed by truncated singular-value decomposition into empirical orthogonal functions (EOFs); the leading <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:mi>K</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">11</mml:mn></mml:mrow></mml:math></inline-formula> modes capture 95 % of the training variance. A vector autoregression of order <inline-formula><mml:math id="M85" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> (selected by AIC, giving <inline-formula><mml:math id="M86" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:math></inline-formula>) is fitted to the <inline-formula><mml:math id="M87" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula> principal-component time series, jointly modelling the temporal evolution of the dominant spatial modes. One-step-ahead forecasts of the principal components are reconstructed to the full field and scored identically to the other methods. EOF-VAR thus captures spatial covariance (through the shared EOFs) and temporal dynamics (through the VAR) within a linear statistical framework.</p>
      <p id="d2e2008">The <italic>perceptron</italic> is a single hidden layer of 512 units with rectified linear unit (ReLU) activation, taking the flattened lookback (<inline-formula><mml:math id="M88" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mo>×</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">21</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">21</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">2646</mml:mn></mml:mrow></mml:math></inline-formula> features) as input. The <italic>MLP</italic> uses two hidden layers of 256 units each with ReLU activation. Both use a sigmoid output layer to predict the normalised <inline-formula><mml:math id="M89" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> current field.</p>
      <p id="d2e2059">The <italic>CNN</italic> <xref ref-type="bibr" rid="bib1.bibx19" id="paren.30"/> uses two blocks of Conv2d layers (32 and 64 filters, <inline-formula><mml:math id="M90" display="inline"><mml:mrow><mml:mn mathvariant="normal">3</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula> kernels, exponential linear unit (ELU) activation) with <inline-formula><mml:math id="M91" display="inline"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> max-pooling and dropout (0.2) after each block. All <inline-formula><mml:math id="M92" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula> lookback frames are stacked along the channel dimension (<inline-formula><mml:math id="M93" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mo>×</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:math></inline-formula> input channels for <inline-formula><mml:math id="M94" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>). The feature maps are flattened and passed through a fully connected layer of 512 units (ReLU) followed by a sigmoid output layer. Figure <xref ref-type="fig" rid="F3"/> illustrates the CNN architecture.</p>

      <fig id="F3" specific-use="star"><label>Figure 3</label><caption><p id="d2e2136">CNN architecture. All <inline-formula><mml:math id="M95" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula> lookback frames are stacked along the channel dimension and processed by two convolutional blocks (32 and 64 filters), followed by flattening and a fully connected output layer (1.34 M parameters).</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f03.png"/>

        </fig>

      <p id="d2e2153">The <italic>GRU</italic> <xref ref-type="bibr" rid="bib1.bibx5" id="paren.31"/> processes the flattened frame sequence through a 256-unit recurrent layer; at each timestep, the <inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">21</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">21</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">882</mml:mn></mml:mrow></mml:math></inline-formula> values are flattened into a single vector, and the final hidden state is decoded via a fully connected layer with sigmoid activation. Figure <xref ref-type="fig" rid="F4"/> illustrates the GRU architecture.</p>

      <fig id="F4" specific-use="star"><label>Figure 4</label><caption><p id="d2e2188">GRU architecture. Each lookback frame is flattened to an 882-dimensional vector and fed sequentially to a 256-unit GRU layer. The final hidden state is decoded via a fully connected layer (1.10 M parameters).</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f04.png"/>

        </fig>

      <p id="d2e2197">The <italic>CNN-GRU</italic> <xref ref-type="bibr" rid="bib1.bibx32" id="paren.32"/> applies the CNN encoder independently to each lookback frame, extracting a spatial feature vector per timestep. The resulting feature sequence is passed to the GRU, and the final hidden state is decoded via a fully connected output layer with sigmoid activation. Figure <xref ref-type="fig" rid="F5"/> illustrates the CNN-GRU architecture.</p>

      <fig id="F5" specific-use="star"><label>Figure 5</label><caption><p id="d2e2210">CNN-GRU architecture. A shared CNN encoder extracts spatial features from each lookback frame independently; the resulting feature sequence is processed by the GRU, and the final hidden state is decoded to the output field (1.72 M parameters).</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f05.png"/>

        </fig>

      <p id="d2e2219">All neural networks are trained with Adam <xref ref-type="bibr" rid="bib1.bibx18" id="paren.33"/> (lr <inline-formula><mml:math id="M97" display="inline"><mml:mrow><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>), early stopping (patience 10), batch size 64, and a maximum of 50 epochs.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Phase 2: lookback sensitivity</title>
      <p id="d2e2251">Phase 2 carries forward the three deep learning models from Phase 1 (CNN, GRU, CNN-GRU), defined as the architectures attaining the lowest Phase 1 validation RMSE among the spatio-temporal candidates. The per-cell statistical methods (moving average, SES, ARIMA, <inline-formula><mml:math id="M98" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>NN) are univariate and have no lookback-window analogue, so they are confined to Phase 1 (Sect. <xref ref-type="sec" rid="Ch1.S3"/>); ARIMA in particular is not carried forward because it is a per-cell model incompatible with the joint spatio-temporal task – and, on the full-domain evaluation, it is not the best method for <inline-formula><mml:math id="M99" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> in any case (Sect. <xref ref-type="sec" rid="Ch1.S4"/>). The three deep learning models are retrained with lookback windows of <inline-formula><mml:math id="M100" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>, 6, and 12 h. At <inline-formula><mml:math id="M101" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">12</mml:mn></mml:mrow></mml:math></inline-formula> h, the lookback approximately covers one full semidiurnal M2 tidal cycle, providing complete tidal phase information.</p>
</sec>
<sec id="Ch1.S3.SS4">
  <label>3.4</label><title>Phase 3: multi-step nowcasting architectures</title>
      <p id="d2e2309">Phase 3 addresses a different task – six-hour multi-step nowcasting rather than one-step prediction – which requires architectures that emit a full <inline-formula><mml:math id="M102" display="inline"><mml:mi>H</mml:mi></mml:math></inline-formula>-step forecast. We therefore introduce three multi-step architectures rather than reusing the Phase 1 one-step models directly; the design builds on the Phase-1/2 finding that the CNN-GRU hybrid is the strongest one-step spatio-temporal model, which motivates the direct multi-step CNN-GRU-MS evaluated here against two autoregressive ConvLSTM baselines. Three architectures are compared for six-hour nowcasting, all using <inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">12</mml:mn></mml:mrow></mml:math></inline-formula>:</p>
      <p id="d2e2333">The <italic>ConvLSTM encoder–decoder</italic> (ConvLSTM-ED) uses a single ConvLSTM layer <xref ref-type="bibr" rid="bib1.bibx28" id="paren.34"/> with 64 hidden channels and <inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:mn mathvariant="normal">3</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula> kernels to encode the <inline-formula><mml:math id="M105" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula>-step input. The decoder is another ConvLSTM layer that autoregressively generates <inline-formula><mml:math id="M106" display="inline"><mml:mi>H</mml:mi></mml:math></inline-formula> output frames, each conditioned on the previous prediction. A <inline-formula><mml:math id="M107" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> convolution maps hidden states to the two-channel (<inline-formula><mml:math id="M108" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M109" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>) output at each step (305 K parameters). Figure <xref ref-type="fig" rid="F6"/> illustrates the ConvLSTM-ED architecture.</p>

      <fig id="F6" specific-use="star"><label>Figure 6</label><caption><p id="d2e2399">ConvLSTM-ED architecture. The encoder processes <inline-formula><mml:math id="M110" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula> input frames through a ConvLSTM layer (64 hidden channels); the decoder autoregressively generates <inline-formula><mml:math id="M111" display="inline"><mml:mi>H</mml:mi></mml:math></inline-formula> output frames, each conditioned on the previous prediction via feedback (dashed arrows). A <inline-formula><mml:math id="M112" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> convolution maps hidden states to the output at each step (305 K parameters).</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f06.png"/>

        </fig>

      <p id="d2e2435">The <italic>bidirectional encoder–forecaster</italic> (BiEF) extends the encoder with forward and backward ConvLSTM passes <xref ref-type="bibr" rid="bib1.bibx27" id="paren.35"/>. The forward and backward hidden and cell states are concatenated and merged via a shared <inline-formula><mml:math id="M113" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> convolution (<inline-formula><mml:math id="M114" display="inline"><mml:mrow><mml:mn mathvariant="normal">128</mml:mn><mml:mo>→</mml:mo><mml:mn mathvariant="normal">64</mml:mn></mml:mrow></mml:math></inline-formula> channels), and the decoder is identical to ConvLSTM-ED (465 K parameters). Figure <xref ref-type="fig" rid="F7"/> illustrates the BiEF architecture.</p>

      <fig id="F7" specific-use="star"><label>Figure 7</label><caption><p id="d2e2472">BiEF architecture. Bidirectional forward and backward ConvLSTM encoders process the input sequence in both directions; the resulting states are merged via a <inline-formula><mml:math id="M115" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> convolution (<inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:mn mathvariant="normal">128</mml:mn><mml:mo>→</mml:mo><mml:mn mathvariant="normal">64</mml:mn></mml:mrow></mml:math></inline-formula> channels) before the same autoregressive decoder as ConvLSTM-ED (465 K parameters).</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f07.png"/>

        </fig>

      <p id="d2e2505">The <italic>direct multi-step CNN-GRU</italic> (CNN-GRU-MS) extends the Phase 1 CNN-GRU to predict all <inline-formula><mml:math id="M117" display="inline"><mml:mi>H</mml:mi></mml:math></inline-formula> future frames simultaneously in a single forward pass, without autoregressive decoding. The GRU hidden state is mapped to <inline-formula><mml:math id="M118" display="inline"><mml:mrow><mml:mi>H</mml:mi><mml:mo>×</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">21</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">21</mml:mn></mml:mrow></mml:math></inline-formula> via a fully connected layer with sigmoid activation (2.85 M parameters). Two reduced-capacity variants are also trained to assess the influence of model capacity on the comparison with the autoregressive models: <italic>CNN-GRU-MS-Small</italic> (GRU hidden size 90, 1.00 M parameters) and <italic>CNN-GRU-MS-Matched</italic> (GRU hidden size 40, 0.48 M parameters, capacity-matched to BiEF). Figure <xref ref-type="fig" rid="F8"/> illustrates the CNN-GRU-MS architecture.</p>

      <fig id="F8" specific-use="star"><label>Figure 8</label><caption><p id="d2e2549">CNN-GRU-MS architecture. The shared CNN encoder extracts spatial features from each lookback frame independently; the GRU processes the resulting feature sequence; and a fully connected layer maps the final hidden state directly to all <inline-formula><mml:math id="M119" display="inline"><mml:mrow><mml:mi>H</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:math></inline-formula> future frames without autoregressive decoding (2.85 M parameters).</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f08.png"/>

        </fig>

      <p id="d2e2572">Phase 3 models are trained with Adam (lr <inline-formula><mml:math id="M120" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M121" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>), early stopping (patience 5), batch size 32, maximum 20 epochs, and gradient clipping at 1.0 for all three architectures. The reduced batch size (32 vs. 64 in Phase 1) accommodates the larger memory footprint of multi-step sequences. The Phase-1 and Phase-2 models were trained on an Intel Core i7-12700 CPU with Intel Extension for PyTorch (IPEX); the controlled multi-step comparison in Tables <xref ref-type="table" rid="T4"/> and <xref ref-type="table" rid="T5"/> was run separately, retraining all five Phase-3 architectures over five seeds (<inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">42</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">123</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">456</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">789</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1024</mml:mn><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>) under these identical settings on a single GPU (per-seed training and single-CPU inference times are reported in Table <xref ref-type="table" rid="T5"/>).</p>
</sec>
<sec id="Ch1.S3.SS5">
  <label>3.5</label><title>Evaluation metrics</title>
      <p id="d2e2640">The primary metric is root-mean-square error (RMSE) computed over all 291 sea grid points and test timesteps. Following <xref ref-type="bibr" rid="bib1.bibx23" id="text.36"/>, the forecast skill score relative to persistence is defined using the MSE as

            <disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M123" display="block"><mml:mrow><mml:mtext>SS</mml:mtext><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mtext>MSE</mml:mtext><mml:mtext>model</mml:mtext></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mtext>MSE</mml:mtext><mml:mtext>persistence</mml:mtext></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M124" display="inline"><mml:mrow><mml:mtext>SS</mml:mtext><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> indicates improvement and <inline-formula><mml:math id="M125" display="inline"><mml:mrow><mml:mtext>SS</mml:mtext><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> is a perfect forecast. For Phase 3, the step-<inline-formula><mml:math id="M126" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> persistence (<inline-formula><mml:math id="M127" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold">X</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="bold">X</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>) is used as the reference at each lead time.</p>
      <p id="d2e2744">Diurnal error patterns are assessed by stratifying per-timestep squared errors by hour of day. To quantify stochastic training variability, the three Phase 1 deep learning models and the Phase 3 CNN-GRU-MS are retrained with five random seeds (42, 123, 456, 789, 1024) each; mean and standard deviation of the RMSE are reported alongside single-run values. Following the philosophy of <xref ref-type="bibr" rid="bib1.bibx9" id="text.37"/>, we interpret model ranking primarily through effect sizes (RMSE differences) rather than <inline-formula><mml:math id="M128" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula>-values.</p>
</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Results</title>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Phase 1: one-step prediction</title>
      <p id="d2e2773">Table <xref ref-type="table" rid="T2"/> presents the RMSE and skill scores for all twelve methods, evaluated consistently over the same 291 sea cells and the full test period. The three deep learning models achieve the lowest errors for <italic>both</italic> components, with CNN leading for <inline-formula><mml:math id="M129" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> (11.31 cm s<sup>−1</sup>, SS<inline-formula><mml:math id="M131" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>U</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">0.502</mml:mn></mml:mrow></mml:math></inline-formula>) and CNN-GRU for <inline-formula><mml:math id="M132" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> (15.44 cm s<sup>−1</sup>, SS<inline-formula><mml:math id="M134" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>V</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.417</mml:mn></mml:mrow></mml:math></inline-formula>). Multi-seed experiments (five seeds per model) confirm that inter-model differences are comparable to stochastic training variability: CNN <inline-formula><mml:math id="M135" display="inline"><mml:mrow><mml:mi>U</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">11.32</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula>, GRU <inline-formula><mml:math id="M136" display="inline"><mml:mrow><mml:mi>U</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">11.45</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula>, CNN-GRU <inline-formula><mml:math id="M137" display="inline"><mml:mrow><mml:mi>U</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">11.32</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.04</mml:mn></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup>. CNN and CNN-GRU are statistically indistinguishable for <inline-formula><mml:math id="M139" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>, while CNN-GRU shows a marginal advantage for <inline-formula><mml:math id="M140" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> (<inline-formula><mml:math id="M141" display="inline"><mml:mrow><mml:mn mathvariant="normal">15.48</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.07</mml:mn></mml:mrow></mml:math></inline-formula> vs. <inline-formula><mml:math id="M142" display="inline"><mml:mrow><mml:mn mathvariant="normal">15.55</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.06</mml:mn></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup>).</p>

<table-wrap id="T2" specific-use="star"><label>Table 2</label><caption><p id="d2e2972">Phase 1 test-set RMSE (cm s<sup>−1</sup>), MSE-based skill scores (SS) relative to persistence, and Pearson correlation coefficients (<inline-formula><mml:math id="M145" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula>) for one-step prediction (<inline-formula><mml:math id="M146" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>). All methods are evaluated over the same 291 sea cells and the full test period. Bold indicates the best performance. ARIMA(2,0,2) is the order favoured by AIC within the low-order search (<inline-formula><mml:math id="M147" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>); higher-order and seasonal extensions do not materially improve out-of-sample skill (Sect. <xref ref-type="sec" rid="Ch1.S3"/>). ARIMA(1,0,1) is the parsimonious specification. Correlations are computed for deep learning models only (see Sect. <xref ref-type="sec" rid="Ch1.S4.SS7"/>).</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="8">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:colspec colnum="8" colname="col8" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Category</oasis:entry>
         <oasis:entry colname="col2">Method</oasis:entry>
         <oasis:entry colname="col3">RMSE<sub><italic>U</italic></sub></oasis:entry>
         <oasis:entry colname="col4">SS<sub><italic>U</italic></sub></oasis:entry>
         <oasis:entry colname="col5">RMSE<sub><italic>V</italic></sub></oasis:entry>
         <oasis:entry colname="col6">SS<sub><italic>V</italic></sub></oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M152" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi>U</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M153" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi>V</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Baseline</oasis:entry>
         <oasis:entry colname="col2">Persistence</oasis:entry>
         <oasis:entry colname="col3">16.03</oasis:entry>
         <oasis:entry colname="col4">0.000</oasis:entry>
         <oasis:entry colname="col5">20.22</oasis:entry>
         <oasis:entry colname="col6">0.000</oasis:entry>
         <oasis:entry colname="col7">–</oasis:entry>
         <oasis:entry colname="col8">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">UTide</oasis:entry>
         <oasis:entry colname="col3">36.43</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M154" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>4.167</oasis:entry>
         <oasis:entry colname="col5">35.90</oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M155" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>2.151</oasis:entry>
         <oasis:entry colname="col7">–</oasis:entry>
         <oasis:entry colname="col8">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Classical</oasis:entry>
         <oasis:entry colname="col2">Moving average</oasis:entry>
         <oasis:entry colname="col3">24.35</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M156" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>1.309</oasis:entry>
         <oasis:entry colname="col5">27.42</oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M157" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.839</oasis:entry>
         <oasis:entry colname="col7">–</oasis:entry>
         <oasis:entry colname="col8">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Exp. smoothing (SES)</oasis:entry>
         <oasis:entry colname="col3">16.03</oasis:entry>
         <oasis:entry colname="col4">0.000</oasis:entry>
         <oasis:entry colname="col5">20.09</oasis:entry>
         <oasis:entry colname="col6">0.013</oasis:entry>
         <oasis:entry colname="col7">–</oasis:entry>
         <oasis:entry colname="col8">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">ARIMA(1,0,1)</oasis:entry>
         <oasis:entry colname="col3">14.95</oasis:entry>
         <oasis:entry colname="col4">0.129</oasis:entry>
         <oasis:entry colname="col5">18.95</oasis:entry>
         <oasis:entry colname="col6">0.122</oasis:entry>
         <oasis:entry colname="col7">–</oasis:entry>
         <oasis:entry colname="col8">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">ARIMA(2,0,2)</oasis:entry>
         <oasis:entry colname="col3">13.98</oasis:entry>
         <oasis:entry colname="col4">0.239</oasis:entry>
         <oasis:entry colname="col5">18.19</oasis:entry>
         <oasis:entry colname="col6">0.191</oasis:entry>
         <oasis:entry colname="col7">–</oasis:entry>
         <oasis:entry colname="col8">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Spatio-temp. statistical</oasis:entry>
         <oasis:entry colname="col2">EOF-VAR</oasis:entry>
         <oasis:entry colname="col3">12.50</oasis:entry>
         <oasis:entry colname="col4">0.392</oasis:entry>
         <oasis:entry colname="col5">16.27</oasis:entry>
         <oasis:entry colname="col6">0.353</oasis:entry>
         <oasis:entry colname="col7">–</oasis:entry>
         <oasis:entry colname="col8">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Shallow ML</oasis:entry>
         <oasis:entry colname="col2">Temporal <inline-formula><mml:math id="M158" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>NN</oasis:entry>
         <oasis:entry colname="col3">14.00</oasis:entry>
         <oasis:entry colname="col4">0.237</oasis:entry>
         <oasis:entry colname="col5">19.08</oasis:entry>
         <oasis:entry colname="col6">0.110</oasis:entry>
         <oasis:entry colname="col7">–</oasis:entry>
         <oasis:entry colname="col8">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Perceptron</oasis:entry>
         <oasis:entry colname="col3">12.26</oasis:entry>
         <oasis:entry colname="col4">0.415</oasis:entry>
         <oasis:entry colname="col5">16.62</oasis:entry>
         <oasis:entry colname="col6">0.324</oasis:entry>
         <oasis:entry colname="col7">–</oasis:entry>
         <oasis:entry colname="col8">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">MLP</oasis:entry>
         <oasis:entry colname="col3">12.22</oasis:entry>
         <oasis:entry colname="col4">0.418</oasis:entry>
         <oasis:entry colname="col5">16.42</oasis:entry>
         <oasis:entry colname="col6">0.340</oasis:entry>
         <oasis:entry colname="col7">–</oasis:entry>
         <oasis:entry colname="col8">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Deep learning</oasis:entry>
         <oasis:entry colname="col2">CNN</oasis:entry>
         <oasis:entry colname="col3"><bold>11.31</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>0.502</bold></oasis:entry>
         <oasis:entry colname="col5">15.65</oasis:entry>
         <oasis:entry colname="col6">0.401</oasis:entry>
         <oasis:entry colname="col7">0.973</oasis:entry>
         <oasis:entry colname="col8">0.948</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">GRU</oasis:entry>
         <oasis:entry colname="col3">11.56</oasis:entry>
         <oasis:entry colname="col4">0.480</oasis:entry>
         <oasis:entry colname="col5">15.63</oasis:entry>
         <oasis:entry colname="col6">0.403</oasis:entry>
         <oasis:entry colname="col7">0.973</oasis:entry>
         <oasis:entry colname="col8">0.948</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">CNN-GRU</oasis:entry>
         <oasis:entry colname="col3">11.33</oasis:entry>
         <oasis:entry colname="col4">0.500</oasis:entry>
         <oasis:entry colname="col5"><bold>15.44</bold></oasis:entry>
         <oasis:entry colname="col6"><bold>0.417</bold></oasis:entry>
         <oasis:entry colname="col7"><bold>0.973</bold></oasis:entry>
         <oasis:entry colname="col8"><bold>0.949</bold></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e3530">A clear, monotone performance hierarchy emerges (Fig. <xref ref-type="fig" rid="F9"/>): deep learning (SS <inline-formula><mml:math id="M159" display="inline"><mml:mo>≈</mml:mo></mml:math></inline-formula> 0.40–0.50) <inline-formula><mml:math id="M160" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> the reduced-rank spatio-temporal statistical model and shallow machine learning (EOF-VAR, perceptron, MLP; SS <inline-formula><mml:math id="M161" display="inline"><mml:mo>≈</mml:mo></mml:math></inline-formula> 0.32–0.42) <inline-formula><mml:math id="M162" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> pointwise classical time series (ARIMA and <inline-formula><mml:math id="M163" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>NN, SS <inline-formula><mml:math id="M164" display="inline"><mml:mo>≈</mml:mo></mml:math></inline-formula> 0.11–0.24) <inline-formula><mml:math id="M165" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> persistence and exponential smoothing (SS <inline-formula><mml:math id="M166" display="inline"><mml:mo>≈</mml:mo></mml:math></inline-formula> 0) <inline-formula><mml:math id="M167" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> the UTide and moving-average baselines (SS <inline-formula><mml:math id="M168" display="inline"><mml:mo>&lt;</mml:mo></mml:math></inline-formula> 0), and this ordering holds for <italic>both</italic> velocity components. The contrast between EOF-VAR (SS<inline-formula><mml:math id="M169" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>U</mml:mi></mml:msub><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.39</mml:mn></mml:mrow></mml:math></inline-formula>, SS<inline-formula><mml:math id="M170" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>V</mml:mi></mml:msub><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.35</mml:mn></mml:mrow></mml:math></inline-formula>) and the pointwise ARIMA (SS<inline-formula><mml:math id="M171" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>U</mml:mi></mml:msub><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.13</mml:mn></mml:mrow></mml:math></inline-formula>, SS<inline-formula><mml:math id="M172" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>V</mml:mi></mml:msub><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">0.12</mml:mn></mml:mrow></mml:math></inline-formula>) is informative: introducing spatial correlation through a reduced-rank spatio-temporal statistical model recovers most of the gap to the neural networks, confirming that the dominant deficiency of the classical baselines is their neglect of spatial structure rather than of nonlinearity. The neural networks nonetheless retain a consistent edge over EOF-VAR (CNN-GRU SS<inline-formula><mml:math id="M173" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>V</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.42</mml:mn></mml:mrow></mml:math></inline-formula> vs. <inline-formula><mml:math id="M174" display="inline"><mml:mn mathvariant="normal">0.35</mml:mn></mml:math></inline-formula>), attributable to their nonlinear spatio-temporal representation. All three deep learning models achieve overall correlation coefficients <inline-formula><mml:math id="M175" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.97</mml:mn></mml:mrow></mml:math></inline-formula> for <inline-formula><mml:math id="M176" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M177" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.94</mml:mn></mml:mrow></mml:math></inline-formula> for <inline-formula><mml:math id="M178" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> (Table <xref ref-type="table" rid="T2"/>).</p>

      <fig id="F9" specific-use="star"><label>Figure 9</label><caption><p id="d2e3741">Phase 1 RMSE for all twelve methods, grouped by category. Bold outlines highlight the best model for each component.</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f09.png"/>

        </fig>

      <p id="d2e3750">The classical autoregressive models are competitive with persistence but do not approach the deep learning tier. The ARIMA and temporal <inline-formula><mml:math id="M179" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>NN figures in Table <xref ref-type="table" rid="T2"/> are evaluated on the full 291-cell domain over the entire test period: ARIMA(1,0,1) reaches 14.95 and 18.95 cm s<sup>−1</sup> (SS 0.13 and 0.12), worse for the meridional component than every deep learning model (CNN-GRU: 15.44 cm s<sup>−1</sup>) and below the MLP. The order favoured by AIC within the low-order search is in fact (2,0,2) rather than (1,0,1) (54 of 60 cell-series; Sect. <xref ref-type="sec" rid="Ch1.S3"/>); even the AIC-optimal ARIMA(2,0,2) reaches only 13.98 and 18.19 cm s<sup>−1</sup> over the full domain, still well short of deep learning. Extending the order search to <inline-formula><mml:math id="M183" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula> or adding a 12 h seasonal component does not materially improve out-of-sample accuracy (median RMSE changes of <inline-formula><mml:math id="M184" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> % and <inline-formula><mml:math id="M185" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula> % on the diagnostics sample, respectively, with the mean change for the higher-order specifications being a slight degradation; Sect. <xref ref-type="sec" rid="Ch1.S3"/>), so this conclusion is not an artifact of the low-order specification. ARIMA's residuals retain significant autocorrelation at both orders (Ljung–Box <inline-formula><mml:math id="M186" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> in essentially all cells), confirming that a per-cell low-order ARMA cannot capture the full tidal harmonic structure. The strong one-step autocorrelation of the tidal signal nonetheless allows ARIMA to exceed persistence, but temporal modelling alone does not match the spatio-temporal networks for either component.</p>
      <p id="d2e3851">The UTide baseline produces strongly negative skill scores (SS<inline-formula><mml:math id="M187" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>U</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">4.17</mml:mn></mml:mrow></mml:math></inline-formula>). This is expected: for one-step prediction, persistence already contains the current tidal phase plus all non-tidal variability, while UTide relies on long-term harmonic fits that cannot capture the residual variability. We note that UTide predictions are also used for gap-filling (Sect. <xref ref-type="sec" rid="Ch1.S2"/>), so the UTide baseline is partially evaluated on data it helped generate; however, this would bias its RMSE <italic>downward</italic>, making the strongly negative skill scores a conservative estimate of its true limitation. The moving average fails (SS<inline-formula><mml:math id="M188" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>U</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1.31</mml:mn></mml:mrow></mml:math></inline-formula>) because averaging over the lookback window smooths the tidal oscillations.</p>
      <p id="d2e3891">Notably, the perceptron has the most parameters (1.81 M) of any Phase 1 model due to the large flattened input dimensionality (<inline-formula><mml:math id="M189" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mo>×</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">21</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">21</mml:mn><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2646</mml:mn></mml:mrow></mml:math></inline-formula> features <inline-formula><mml:math id="M190" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 512 units), yet performs worse than the CNN (1.34 M) and CNN-GRU (1.72 M). This indicates that, where it matters, architectural inductive bias rather than raw parameter count drives performance: the convolutional structure provides a spatial prior that fully connected layers cannot match. We caution, however, against over-generalising from the deep learning comparison: for one-step <italic>zonal</italic> prediction the CNN, GRU, and CNN-GRU are nearly indistinguishable (Table <xref ref-type="table" rid="T2"/>; five-seed <inline-formula><mml:math id="M191" display="inline"><mml:mrow><mml:mi>U</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">11.32</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M192" display="inline"><mml:mrow><mml:mn mathvariant="normal">11.45</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M193" display="inline"><mml:mrow><mml:mn mathvariant="normal">11.32</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.04</mml:mn></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup>), even though the GRU treats the grid as unstructured while the CNN is explicitly spatial. In this tidally dominated regime the repeating semidiurnal signal is strong enough that the choice of spatial architecture matters little for one-step <inline-formula><mml:math id="M195" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>; spatial structure yields a measurable benefit chiefly for <inline-formula><mml:math id="M196" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>, where along-strait gradients are steeper, and – as Phase 2 shows – when a longer lookback lets the hybrid exploit tidal-cycle phase.</p>
      <p id="d2e3999">The skill score comparison (Fig. <xref ref-type="fig" rid="F10"/>) shows all deep learning models consistently achieving SS <inline-formula><mml:math id="M197" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> 0.40.</p>

      <fig id="F10" specific-use="star"><label>Figure 10</label><caption><p id="d2e4014">MSE-based skill scores relative to persistence for all methods. Positive values indicate improvement.</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f10.png"/>

        </fig>

</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Phase 2: lookback sensitivity</title>
      <p id="d2e4031">Figure <xref ref-type="fig" rid="F11"/> and Table <xref ref-type="table" rid="T3"/> show the effect of lookback window length on the three deep learning architectures. Note that the CNN-GRU <inline-formula><mml:math id="M198" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula> values differ slightly between Phase 1 (Table <xref ref-type="table" rid="T2"/>: 11.33 cm s<sup>−1</sup>) and Phase 2 (Table <xref ref-type="table" rid="T3"/>: 11.35 cm s<sup>−1</sup>) due to independent training runs with different random seeds. The results reveal a clear architectural difference: <list list-type="bullet"><list-item>
      <p id="d2e4083">CNN-GRU shows a monotonic trend: RMSE<sub><italic>U</italic></sub> decreases from 11.35 (<inline-formula><mml:math id="M202" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>) to 11.28 (<inline-formula><mml:math id="M203" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:math></inline-formula>) to 11.21 cm s<sup>−1</sup> (<inline-formula><mml:math id="M205" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">12</mml:mn></mml:mrow></mml:math></inline-formula>), a 1.2 % improvement. RMSE<sub><italic>V</italic></sub> shows a similar, though smaller, trend (15.47 <inline-formula><mml:math id="M207" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula> 15.42 <inline-formula><mml:math id="M208" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula> 15.41 cm s<sup>−1</sup>). These differences are modest and comparable to the variation expected from stochastic training; however, the consistent direction across both components and all three lookback values suggests a real, if small, effect.</p></list-item><list-item>
      <p id="d2e4186">The standalone CNN and GRU show no consistent benefit from longer lookbacks: CNN RMSE<sub><italic>U</italic></sub> is essentially flat (11.30, 11.32, 11.33 cm s<sup>−1</sup>). GRU performs slightly worse at longer lookbacks, possibly because additional input length introduces noise without commensurate spatial context.</p></list-item></list></p>

      <fig id="F11" specific-use="star"><label>Figure 11</label><caption><p id="d2e4212">Phase 2: RMSE as a function of lookback window length for the three deep learning architectures. <bold>(a)</bold> <inline-formula><mml:math id="M212" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> component. <bold>(b)</bold> <inline-formula><mml:math id="M213" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> component.</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f11.png"/>

        </fig>

<table-wrap id="T3" specific-use="star"><label>Table 3</label><caption><p id="d2e4244">Phase 2: RMSE (cm s<sup>−1</sup>) for deep learning models across lookback windows. Bold indicates the best result per column.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right" colsep="1"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right" colsep="1"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Model</oasis:entry>
         <oasis:entry rowsep="1" namest="col2" nameend="col3" align="center" colsep="1"><inline-formula><mml:math id="M215" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry rowsep="1" namest="col4" nameend="col5" align="center" colsep="1"><inline-formula><mml:math id="M216" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry rowsep="1" namest="col6" nameend="col7" align="center"><inline-formula><mml:math id="M217" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">12</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">RMSE<sub><italic>U</italic></sub></oasis:entry>
         <oasis:entry colname="col3">RMSE<sub><italic>V</italic></sub></oasis:entry>
         <oasis:entry colname="col4">RMSE<sub><italic>U</italic></sub></oasis:entry>
         <oasis:entry colname="col5">RMSE<sub><italic>V</italic></sub></oasis:entry>
         <oasis:entry colname="col6">RMSE<sub><italic>U</italic></sub></oasis:entry>
         <oasis:entry colname="col7">RMSE<sub><italic>V</italic></sub></oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">CNN</oasis:entry>
         <oasis:entry colname="col2"><bold>11.30</bold></oasis:entry>
         <oasis:entry colname="col3">15.57</oasis:entry>
         <oasis:entry colname="col4">11.32</oasis:entry>
         <oasis:entry colname="col5">15.53</oasis:entry>
         <oasis:entry colname="col6">11.33</oasis:entry>
         <oasis:entry colname="col7">15.60</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">GRU</oasis:entry>
         <oasis:entry colname="col2">11.32</oasis:entry>
         <oasis:entry colname="col3">15.58</oasis:entry>
         <oasis:entry colname="col4">11.45</oasis:entry>
         <oasis:entry colname="col5">15.70</oasis:entry>
         <oasis:entry colname="col6">11.43</oasis:entry>
         <oasis:entry colname="col7">15.66</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CNN-GRU</oasis:entry>
         <oasis:entry colname="col2">11.35</oasis:entry>
         <oasis:entry colname="col3"><bold>15.47</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>11.28</bold></oasis:entry>
         <oasis:entry colname="col5"><bold>15.42</bold></oasis:entry>
         <oasis:entry colname="col6"><bold>11.21</bold></oasis:entry>
         <oasis:entry colname="col7"><bold>15.41</bold></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e4490">The CNN-GRU's monotonic improvement with lookback length indicates that the GRU component effectively integrates tidal phase information from longer sequences, while the CNN provides spatial context at each timestep. In contrast, the standalone GRU receives flattened frames without spatial structure, and the CNN stacks all frames along the channel dimension, which does not naturally model temporal ordering.</p>
</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>Phase 3: six-hour nowcasting</title>
      <p id="d2e4501">Table <xref ref-type="table" rid="T4"/> reports the average RMSE across the six lead times for every multi-step model. To make the comparison free of platform, batch-size, and single-seed confounds, all five trainable models were retrained under an <italic>identical</italic> protocol – five random seeds, batch size 32, the same data, chronological splits, optimiser, and early-stopping criterion – on a single GPU (Sect. <xref ref-type="sec" rid="Ch1.S3"/>). Accuracy is governed primarily by model <italic>capacity</italic>: the 1.00 M-parameter CNN-GRU-MS-Small attains the lowest error (<inline-formula><mml:math id="M224" display="inline"><mml:mrow><mml:mn mathvariant="normal">18.39</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.15</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M225" display="inline"><mml:mrow><mml:mn mathvariant="normal">22.07</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.19</mml:mn></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup>) and the 2.85 M full model is statistically identical (<inline-formula><mml:math id="M227" display="inline"><mml:mrow><mml:mn mathvariant="normal">18.49</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.11</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M228" display="inline"><mml:mrow><mml:mn mathvariant="normal">22.29</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.09</mml:mn></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup>), so capacity beyond <inline-formula><mml:math id="M230" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 1 M yields no further gain; error rises only at the smallest sizes. Critically, at matched capacity (<inline-formula><mml:math id="M231" display="inline"><mml:mo lspace="0mm">∼</mml:mo></mml:math></inline-formula> 0.48 M) the direct CNN-GRU-MS-Matched (<inline-formula><mml:math id="M232" display="inline"><mml:mrow><mml:mn mathvariant="normal">19.08</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.18</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M233" display="inline"><mml:mrow><mml:mn mathvariant="normal">22.53</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.18</mml:mn></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup>) and the autoregressive BiEF (<inline-formula><mml:math id="M235" display="inline"><mml:mrow><mml:mn mathvariant="normal">19.00</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.19</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M236" display="inline"><mml:mrow><mml:mn mathvariant="normal">22.28</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.24</mml:mn></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup>) are <italic>statistically indistinguishable</italic>: once capacity and training budget are equalised, direct multi-step prediction confers no accuracy advantage over autoregressive decoding. All models reduce the persistence mean square error by 70 %–77 % (skill scores 0.70–0.77).</p>

<table-wrap id="T4" specific-use="star"><label>Table 4</label><caption><p id="d2e4681">Phase 3: average RMSE (cm s<sup>−1</sup>) over six lead times and MSE skill scores for multi-step nowcasting (<inline-formula><mml:math id="M239" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">12</mml:mn></mml:mrow></mml:math></inline-formula>), with model size. All five trainable models were retrained under an identical protocol (five seeds, batch 32, same splits, optimiser, and early stopping) on one GPU; values are five-seed means <inline-formula><mml:math id="M240" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula> standard deviation. Bold indicates the lowest RMSE. The 0.48 M CNN-GRU-MS-Matched is capacity-matched to BiEF (0.47 M).</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="6">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Model</oasis:entry>
         <oasis:entry colname="col2">Params</oasis:entry>
         <oasis:entry colname="col3">Avg RMSE<sub><italic>U</italic></sub></oasis:entry>
         <oasis:entry colname="col4">SS<sub><italic>U</italic></sub></oasis:entry>
         <oasis:entry colname="col5">Avg RMSE<sub><italic>V</italic></sub></oasis:entry>
         <oasis:entry colname="col6">SS<sub><italic>V</italic></sub></oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Persistence</oasis:entry>
         <oasis:entry colname="col2">–</oasis:entry>
         <oasis:entry colname="col3">37.95</oasis:entry>
         <oasis:entry colname="col4">0.000</oasis:entry>
         <oasis:entry colname="col5">41.56</oasis:entry>
         <oasis:entry colname="col6">0.000</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ConvLSTM-ED (autoregressive)</oasis:entry>
         <oasis:entry colname="col2">0.30 M</oasis:entry>
         <oasis:entry colname="col3">19.36 <inline-formula><mml:math id="M245" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula> 0.34</oasis:entry>
         <oasis:entry colname="col4">0.740</oasis:entry>
         <oasis:entry colname="col5">22.76 <inline-formula><mml:math id="M246" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula> 0.24</oasis:entry>
         <oasis:entry colname="col6">0.700</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">BiEF (autoregressive)</oasis:entry>
         <oasis:entry colname="col2">0.47 M</oasis:entry>
         <oasis:entry colname="col3">19.00 <inline-formula><mml:math id="M247" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula> 0.19</oasis:entry>
         <oasis:entry colname="col4">0.749</oasis:entry>
         <oasis:entry colname="col5">22.28 <inline-formula><mml:math id="M248" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula> 0.24</oasis:entry>
         <oasis:entry colname="col6">0.713</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CNN-GRU-MS-Matched (direct)</oasis:entry>
         <oasis:entry colname="col2">0.48 M</oasis:entry>
         <oasis:entry colname="col3">19.08 <inline-formula><mml:math id="M249" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula> 0.18</oasis:entry>
         <oasis:entry colname="col4">0.747</oasis:entry>
         <oasis:entry colname="col5">22.53 <inline-formula><mml:math id="M250" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula> 0.18</oasis:entry>
         <oasis:entry colname="col6">0.706</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><bold>CNN-GRU-MS-Small (direct)</bold></oasis:entry>
         <oasis:entry colname="col2">1.00 M</oasis:entry>
         <oasis:entry colname="col3"><bold>18.39</bold> <inline-formula><mml:math id="M251" display="inline"><mml:mo mathvariant="bold">±</mml:mo></mml:math></inline-formula> <bold>0.15</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>0.765</bold></oasis:entry>
         <oasis:entry colname="col5"><bold>22.07</bold> <inline-formula><mml:math id="M252" display="inline"><mml:mo mathvariant="bold">±</mml:mo></mml:math></inline-formula> <bold>0.19</bold></oasis:entry>
         <oasis:entry colname="col6"><bold>0.718</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CNN-GRU-MS (direct)</oasis:entry>
         <oasis:entry colname="col2">2.85 M</oasis:entry>
         <oasis:entry colname="col3">18.49 <inline-formula><mml:math id="M253" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula> 0.11</oasis:entry>
         <oasis:entry colname="col4">0.763</oasis:entry>
         <oasis:entry colname="col5">22.29 <inline-formula><mml:math id="M254" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula> 0.09</oasis:entry>
         <oasis:entry colname="col6">0.712</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e5009">Although the architectures achieve comparable <italic>accuracy</italic> at matched capacity, they differ sharply in <italic>computational cost</italic> (Table <xref ref-type="table" rid="T5"/>), which is the decisive factor for operational deployment. The direct multi-step models train about 2.5 times faster – they emit all six steps in a single pass and converge in fewer epochs – and, more importantly for an operational nowcast, run roughly 4.5 times faster at inference, because the autoregressive models must decode the horizon step by step. A single direct model also covers all 291 cells, whereas the strongest classical baseline (per-cell ARIMA) requires fitting and maintaining 291 separate models.</p>

<table-wrap id="T5" specific-use="star"><label>Table 5</label><caption><p id="d2e5024">Phase 3 computational cost: per-seed training time (single NVIDIA RTX 5090 GPU) and CPU inference latency for one six-hour forecast (batch 1, single AMD Ryzen 7 7800X3D thread group). Absolute values are hardware-dependent and given for reference; the relative differences – the direct multi-step models are markedly cheaper on both axes – are hardware-independent.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Model</oasis:entry>
         <oasis:entry colname="col2">Params</oasis:entry>
         <oasis:entry colname="col3">Train (s)</oasis:entry>
         <oasis:entry colname="col4">Inference (ms)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">ConvLSTM-ED (autoregressive)</oasis:entry>
         <oasis:entry colname="col2">0.30 M</oasis:entry>
         <oasis:entry colname="col3">80</oasis:entry>
         <oasis:entry colname="col4">8.0</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">BiEF (autoregressive)</oasis:entry>
         <oasis:entry colname="col2">0.47 M</oasis:entry>
         <oasis:entry colname="col3">128</oasis:entry>
         <oasis:entry colname="col4">13.4</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CNN-GRU-MS-Matched (direct)</oasis:entry>
         <oasis:entry colname="col2">0.48 M</oasis:entry>
         <oasis:entry colname="col3">53</oasis:entry>
         <oasis:entry colname="col4">3.0</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CNN-GRU-MS-Small (direct)</oasis:entry>
         <oasis:entry colname="col2">1.00 M</oasis:entry>
         <oasis:entry colname="col3">50</oasis:entry>
         <oasis:entry colname="col4">3.1</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CNN-GRU-MS (direct)</oasis:entry>
         <oasis:entry colname="col2">2.85 M</oasis:entry>
         <oasis:entry colname="col3">56</oasis:entry>
         <oasis:entry colname="col4">3.3</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e5137">Figure <xref ref-type="fig" rid="F12"/> shows the RMSE degradation with lead time (a representative single-run profile; the controlled multi-seed averages are those in Table <xref ref-type="table" rid="T4"/>). Persistence RMSE grows from 16.0 cm s<sup>−1</sup> (<inline-formula><mml:math id="M256" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>) to 56.4 cm s<sup>−1</sup> (<inline-formula><mml:math id="M258" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:math></inline-formula>) for <inline-formula><mml:math id="M259" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> – a factor of 3.5. The full CNN-GRU-MS degrades from 12.4 to 24.4 cm s<sup>−1</sup>, a factor of only 2.0.</p>

      <fig id="F12" specific-use="star"><label>Figure 12</label><caption><p id="d2e5214">Phase 3: RMSE as a function of lead time for all nowcasting models. <bold>(a)</bold> <inline-formula><mml:math id="M261" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> component. <bold>(b)</bold> <inline-formula><mml:math id="M262" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> component.</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f12.png"/>

        </fig>

      <p id="d2e5243">A crossover occurs at <inline-formula><mml:math id="M263" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>, where the autoregressive models (ConvLSTM-ED: 12.0 cm s<sup>−1</sup>; BiEF: 11.8 cm s<sup>−1</sup>) slightly outperform the full CNN-GRU-MS (12.4 cm s<sup>−1</sup>) for <inline-formula><mml:math id="M267" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> before the latter pulls ahead at longer leads. This reflects a genuine trade-off: autoregressive decoding can focus each step on one-step-ahead prediction (advantageous at short leads) but accumulates error over the horizon, whereas direct prediction distributes capacity across all six steps. We caution, however, that the full model's longer-lead advantage over BiEF in Fig. <xref ref-type="fig" rid="F12"/> partly reflects its six-fold larger capacity: at matched capacity the two strategies are statistically comparable (Table <xref ref-type="table" rid="T4"/>), so the operational case for the direct model rests on its much lower computational cost (Table <xref ref-type="table" rid="T5"/>) rather than on a lead-time accuracy advantage.</p>
      <p id="d2e5308">The per-lead-time skill scores (Fig. <xref ref-type="fig" rid="F13"/>) show that all models maintain SS<inline-formula><mml:math id="M268" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>U</mml:mi></mml:msub><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.4</mml:mn></mml:mrow></mml:math></inline-formula> at every lead time and SS<inline-formula><mml:math id="M269" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>V</mml:mi></mml:msub><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.3</mml:mn></mml:mrow></mml:math></inline-formula>. CNN-GRU-MS reaches SS<inline-formula><mml:math id="M270" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>U</mml:mi></mml:msub><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.81</mml:mn></mml:mrow></mml:math></inline-formula> at <inline-formula><mml:math id="M271" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:math></inline-formula>, reducing 81 % of the persistence mean square error even at the longest forecast horizon.</p>

      <fig id="F13" specific-use="star"><label>Figure 13</label><caption><p id="d2e5372">Phase 3: per-lead-time skill score relative to step-<inline-formula><mml:math id="M272" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> persistence. <bold>(a)</bold> <inline-formula><mml:math id="M273" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> component. <bold>(b)</bold> <inline-formula><mml:math id="M274" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> component.</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f13.png"/>

        </fig>

      <p id="d2e5408">BiEF outperforms ConvLSTM-ED at all lead times, with the advantage increasing from 0.2 cm s<sup>−1</sup> at <inline-formula><mml:math id="M276" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> to 0.8 cm s<sup>−1</sup> at <inline-formula><mml:math id="M278" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:math></inline-formula> for <inline-formula><mml:math id="M279" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>, indicating that bidirectional encoding provides a richer initial state for the decoder.</p>
</sec>
<sec id="Ch1.S4.SS4">
  <label>4.4</label><title>Diurnal error analysis</title>
      <p id="d2e5474">Figure <xref ref-type="fig" rid="F14"/> shows the one-step RMSE (Phase 1) stratified by hour of day. The Sunda Strait is at approximately 105° E (UTC<inline-formula><mml:math id="M280" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>7). The elevated-error period (06:00–18:00 UTC) corresponds to 13:00 to 01:00 LT (following day). The difference between high-error and low-error periods is visually evident across all models.</p>

      <fig id="F14" specific-use="star"><label>Figure 14</label><caption><p id="d2e5488">Diurnal RMSE pattern for the five neural network models (Phase 1). Red shading indicates the high-error period (06:00–18:00 UTC, corresponding to 13:00–01:00 LT). <bold>(a)</bold> <inline-formula><mml:math id="M281" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> component. <bold>(b)</bold> <inline-formula><mml:math id="M282" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> component.</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f14.png"/>

        </fig>

      <p id="d2e5517">To investigate the physical driver, we compare the diurnal RMSE pattern with concurrent wind observations from the three AWS stations at Merak, Ciwandan, and Bakauheni (Fig. <xref ref-type="fig" rid="F15"/>). The three-station average wind speed shows a clear sea-breeze signal, peaking at <inline-formula><mml:math id="M283" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">4.8</mml:mn></mml:mrow></mml:math></inline-formula> m s<sup>−1</sup> near 15:00 LT (08:00 UTC) and dropping to <inline-formula><mml:math id="M285" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">3.0</mml:mn></mml:mrow></mml:math></inline-formula> m s<sup>−1</sup> at night. However, the relationship between the diurnal cycles of wind speed and model RMSE – measured as a Spearman rank correlation across the 24 hourly bins, so the associated <inline-formula><mml:math id="M287" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula>-values are indicative rather than strict given the smooth, autocorrelated diurnal shape and small effective sample – differs by component: for <inline-formula><mml:math id="M288" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>, the correlation is positive (<inline-formula><mml:math id="M289" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">0.48</mml:mn></mml:mrow></mml:math></inline-formula>, nominal <inline-formula><mml:math id="M290" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.018</mml:mn></mml:mrow></mml:math></inline-formula>), consistent with sea-breeze-driven zonal surface currents degrading prediction skill. For <inline-formula><mml:math id="M291" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>, it is negative (<inline-formula><mml:math id="M292" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.41</mml:mn></mml:mrow></mml:math></inline-formula>), so the elevated meridional errors are not explained by wind speed and instead follow a separate diurnal cycle.</p>
      <p id="d2e5639">We examine this component split directly by projecting the three-station AWS wind onto the strait axis – oriented NE–SW, the principal axis of the observed current variability (Sect. <xref ref-type="sec" rid="Ch1.S4.SS7"/>) – and regressing the diurnal component errors on the along- and cross-strait wind. The zonal (<inline-formula><mml:math id="M293" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>) error scales with total wind <italic>speed</italic> (<inline-formula><mml:math id="M294" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.48</mml:mn></mml:mrow></mml:math></inline-formula>) more strongly than with either signed projection (<inline-formula><mml:math id="M295" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.04</mml:mn></mml:mrow></mml:math></inline-formula> cross-strait, <inline-formula><mml:math id="M296" display="inline"><mml:mn mathvariant="normal">0.36</mml:mn></mml:math></inline-formula> along-strait), consistent with the afternoon sea breeze injecting unmodelled, current-only-unpredictable momentum that raises <inline-formula><mml:math id="M297" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> error when the wind is strongest. The meridional (<inline-formula><mml:math id="M298" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>) error, in contrast, is <italic>negatively</italic> correlated with wind speed (<inline-formula><mml:math id="M299" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.41</mml:mn></mml:mrow></mml:math></inline-formula>) yet <italic>positively</italic> correlated with <italic>both</italic> the along- and cross-strait wind projections (<inline-formula><mml:math id="M300" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">0.64</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M301" display="inline"><mml:mn mathvariant="normal">0.76</mml:mn></mml:math></inline-formula>). The decisive feature is the <italic>sign</italic> with respect to speed: wind drag should grow as the wind strengthens, yet the <inline-formula><mml:math id="M302" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> error <italic>falls</italic> when the wind is strongest, its diurnal minimum coinciding with the afternoon sea-breeze maximum. The positive correlation with both orthogonal projections is consistent with – though it does not by itself prove – a shared diurnal cycle rather than direct forcing, since both projections also peak with the afternoon sea breeze; the speed-negative relationship and the tidal-phase timing are the cleaner diagnostics. Taken together, and bearing in mind that these are descriptive correlations over 24 diurnal bins, the evidence is most consistent with the apparent negative <inline-formula><mml:math id="M303" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>-wind correlation being a tidal-phase coincidence: the meridional error – systematically the harder component across all methods and phases – appears dominated by tidal dynamics rather than by wind forcing.</p>

      <fig id="F15" specific-use="star"><label>Figure 15</label><caption><p id="d2e5786">Diurnal RMSE (red, left axis) overlaid with three-station average wind speed (blue dashed, right axis). Red shading marks the high-error period (06:00–18:00 UTC, corresponding to 13:00–01:00 LT). <bold>(a)</bold> <inline-formula><mml:math id="M304" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> component. <bold>(b)</bold> <inline-formula><mml:math id="M305" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> component.</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f15.png"/>

        </fig>

      <p id="d2e5815">All five neural network models exhibit the same diurnal modulation, indicating a physical limitation rather than a model-specific artefact.</p>
</sec>
<sec id="Ch1.S4.SS5">
  <label>4.5</label><title>Seasonal error stratification</title>
      <p id="d2e5827">To assess the influence of monsoon regime on prediction skill, Phase 1 test-set errors for persistence, ARIMA(1,0,1), and the three deep learning models are stratified into three seasons: the southeast (SE) monsoon (July–August), the transition period (September–November), and the northwest (NW) monsoon (December–February). Table <xref ref-type="table" rid="T6"/> summarises the results; ARIMA is stratified over the full domain and period, consistently with the other methods.</p>

<table-wrap id="T6" specific-use="star"><label>Table 6</label><caption><p id="d2e5835">Seasonal RMSE (cm s<sup>−1</sup>) for persistence, ARIMA(1,0,1), and the three deep learning models (Phase 1, <inline-formula><mml:math id="M307" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>). Bold indicates the lowest RMSE per season. ARIMA's lowest-error season is the SE monsoon.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right" colsep="1"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right" colsep="1"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Model</oasis:entry>
         <oasis:entry rowsep="1" namest="col2" nameend="col3" align="center" colsep="1">SE monsoon </oasis:entry>
         <oasis:entry rowsep="1" namest="col4" nameend="col5" align="center" colsep="1">Transition </oasis:entry>
         <oasis:entry rowsep="1" namest="col6" nameend="col7" align="center">NW monsoon </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">RMSE<sub><italic>U</italic></sub></oasis:entry>
         <oasis:entry colname="col3">RMSE<sub><italic>V</italic></sub></oasis:entry>
         <oasis:entry colname="col4">RMSE<sub><italic>U</italic></sub></oasis:entry>
         <oasis:entry colname="col5">RMSE<sub><italic>V</italic></sub></oasis:entry>
         <oasis:entry colname="col6">RMSE<sub><italic>U</italic></sub></oasis:entry>
         <oasis:entry colname="col7">RMSE<sub><italic>V</italic></sub></oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Persistence</oasis:entry>
         <oasis:entry colname="col2">14.64</oasis:entry>
         <oasis:entry colname="col3">17.64</oasis:entry>
         <oasis:entry colname="col4">15.68</oasis:entry>
         <oasis:entry colname="col5">20.64</oasis:entry>
         <oasis:entry colname="col6">17.28</oasis:entry>
         <oasis:entry colname="col7">21.46</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ARIMA(1,0,1)</oasis:entry>
         <oasis:entry colname="col2">13.44</oasis:entry>
         <oasis:entry colname="col3">16.29</oasis:entry>
         <oasis:entry colname="col4">14.77</oasis:entry>
         <oasis:entry colname="col5">19.39</oasis:entry>
         <oasis:entry colname="col6">16.13</oasis:entry>
         <oasis:entry colname="col7">20.20</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CNN</oasis:entry>
         <oasis:entry colname="col2">10.21</oasis:entry>
         <oasis:entry colname="col3">13.35</oasis:entry>
         <oasis:entry colname="col4">11.22</oasis:entry>
         <oasis:entry colname="col5">16.06</oasis:entry>
         <oasis:entry colname="col6">12.12</oasis:entry>
         <oasis:entry colname="col7">16.69</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">GRU</oasis:entry>
         <oasis:entry colname="col2">10.19</oasis:entry>
         <oasis:entry colname="col3">13.30</oasis:entry>
         <oasis:entry colname="col4">11.34</oasis:entry>
         <oasis:entry colname="col5">16.11</oasis:entry>
         <oasis:entry colname="col6">12.05</oasis:entry>
         <oasis:entry colname="col7">16.51</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">CNN-GRU</oasis:entry>
         <oasis:entry colname="col2"><bold>10.19</bold></oasis:entry>
         <oasis:entry colname="col3"><bold>13.11</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>11.19</bold></oasis:entry>
         <oasis:entry colname="col5"><bold>15.94</bold></oasis:entry>
         <oasis:entry colname="col6"><bold>12.25</bold></oasis:entry>
         <oasis:entry colname="col7"><bold>16.50</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><inline-formula><mml:math id="M314" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> (hours)</oasis:entry>
         <oasis:entry namest="col2" nameend="col3" align="center" colsep="1">1488 </oasis:entry>
         <oasis:entry namest="col4" nameend="col5" align="center" colsep="1">2184 </oasis:entry>
         <oasis:entry namest="col6" nameend="col7" align="center">2089 </oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e6130">All models show a consistent seasonal dependence: RMSE is lowest during the SE monsoon (<inline-formula><mml:math id="M315" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>: 10.19 cm s<sup>−1</sup>, <inline-formula><mml:math id="M317" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>: 13.11 cm s<sup>−1</sup> for CNN-GRU) and highest during the NW monsoon (<inline-formula><mml:math id="M319" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>: 12.25 cm s<sup>−1</sup>, <inline-formula><mml:math id="M321" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>: 16.50 cm s<sup>−1</sup>), representing a 20 % (<inline-formula><mml:math id="M323" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>) and 26 % (<inline-formula><mml:math id="M324" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>) seasonal increase in error. The NW monsoon brings stronger and more variable winds from the Java Sea <xref ref-type="bibr" rid="bib1.bibx35" id="paren.38"/>, introducing non-tidal surface current variability that is not captured by models trained predominantly on tidal dynamics. Persistence and ARIMA errors follow the same seasonal pattern, confirming that the increased NW monsoon error reflects genuinely harder-to-predict dynamics rather than model degradation. ARIMA's lowest-error season is the SE monsoon (<inline-formula><mml:math id="M325" display="inline"><mml:mrow><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">16.29</mml:mn></mml:mrow></mml:math></inline-formula> vs. 18.95 cm s<sup>−1</sup> over the full period), consistent with the stronger non-tidal forcing of the NW monsoon being the principal source of seasonal error growth.</p>
</sec>
<sec id="Ch1.S4.SS6">
  <label>4.6</label><title>Tidal–residual error decomposition</title>
      <p id="d2e6260">To quantify the contribution of tidal versus non-tidal variability to prediction errors, we decompose the observed current at each grid point into a tidal component (predicted by UTide fitted to the training period) and a residual. Table <xref ref-type="table" rid="T7"/> presents the results. For persistence, the tidal RMSE is negligible (<inline-formula><mml:math id="M327" display="inline"><mml:mrow><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">0.2</mml:mn></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup>) because the tidal signal changes minimally over a one-hour lag. Nearly all persistence error arises from the residual component (35.8 and 39.7 cm s<sup>−1</sup>). CNN and CNN-GRU reduce residual-component errors by 19 %–20 % (<inline-formula><mml:math id="M330" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>) and 16 % (<inline-formula><mml:math id="M331" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>) relative to persistence (GRU is omitted because it lacks the CNN encoder needed to isolate spatial contributions), indicating that they capture some non-tidal variability – likely wind-driven and sub-tidal fluctuations – beyond the deterministic tidal signal. However, the residual RMSE remains substantially larger than the total RMSE (28.9 vs. 11.3 cm s<sup>−1</sup> for CNN), reflecting the well-known challenge of predicting mesoscale and sub-mesoscale ocean variability.</p>

<table-wrap id="T7" specific-use="star"><label>Table 7</label><caption><p id="d2e6329">Tidal–residual error decomposition for one-step prediction (<inline-formula><mml:math id="M333" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>). Tidal RMSE measures the error in the tidal component of each prediction; residual RMSE measures the error in the non-tidal component. CNN and CNN-GRU total RMSE values differ slightly from Table <xref ref-type="table" rid="T2"/> (by <inline-formula><mml:math id="M334" display="inline"><mml:mrow><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">0.03</mml:mn></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup>) because the tidal decomposition uses an independent training run.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right" colsep="1"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Model</oasis:entry>
         <oasis:entry rowsep="1" namest="col2" nameend="col4" align="center" colsep="1"><inline-formula><mml:math id="M336" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> (cm s<sup>−1</sup>) </oasis:entry>
         <oasis:entry rowsep="1" namest="col5" nameend="col7" align="center"><inline-formula><mml:math id="M338" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> (cm s<sup>−1</sup>) </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Total</oasis:entry>
         <oasis:entry colname="col3">Tidal</oasis:entry>
         <oasis:entry colname="col4">Residual</oasis:entry>
         <oasis:entry colname="col5">Total</oasis:entry>
         <oasis:entry colname="col6">Tidal</oasis:entry>
         <oasis:entry colname="col7">Residual</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Persistence</oasis:entry>
         <oasis:entry colname="col2">16.03</oasis:entry>
         <oasis:entry colname="col3">0.15</oasis:entry>
         <oasis:entry colname="col4">35.85</oasis:entry>
         <oasis:entry colname="col5">20.22</oasis:entry>
         <oasis:entry colname="col6">0.20</oasis:entry>
         <oasis:entry colname="col7">39.73</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CNN</oasis:entry>
         <oasis:entry colname="col2">11.31</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
         <oasis:entry colname="col4">28.85</oasis:entry>
         <oasis:entry colname="col5">15.65</oasis:entry>
         <oasis:entry colname="col6">–</oasis:entry>
         <oasis:entry colname="col7">33.44</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CNN-GRU</oasis:entry>
         <oasis:entry colname="col2">11.35</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
         <oasis:entry colname="col4">29.26</oasis:entry>
         <oasis:entry colname="col5">15.47</oasis:entry>
         <oasis:entry colname="col6">–</oasis:entry>
         <oasis:entry colname="col7">33.42</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e6543">This decomposition also lets us test, rather than assert, the nature of ARIMA's skill. We split each cell's series into the UTide tidal component and a non-tidal residual and re-fit the ARIMA(1,0,1) directly to each band. The deterministic tide changes by only <inline-formula><mml:math id="M340" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">0.07</mml:mn></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup> over a one-hour lag in this decomposition (consistent with the under-0.2 cm s<sup>−1</sup> persistence tidal RMSE in Table <xref ref-type="table" rid="T7"/>), so it is trivially predictable at one step; consequently persistence's one-step error on the raw series (16.0 and 20.2 cm s<sup>−1</sup> for <inline-formula><mml:math id="M344" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M345" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>) is essentially identical to its error on the residual alone (16.0 and 20.2 cm s<sup>−1</sup>) – the one-step error is almost entirely sub-tidal. An ARIMA fitted directly to that residual achieves <italic>zero</italic> skill over persistence (skill score <inline-formula><mml:math id="M347" display="inline"><mml:mn mathvariant="normal">0.00</mml:mn></mml:math></inline-formula> for <inline-formula><mml:math id="M348" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M349" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.02</mml:mn></mml:mrow></mml:math></inline-formula> for <inline-formula><mml:math id="M350" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>; RMSE 16.0 and 20.4 versus persistence's 16.0 and 20.2 cm s<sup>−1</sup>). ARIMA's modest edge over persistence on the raw series (full-domain SS <inline-formula><mml:math id="M352" display="inline"><mml:mo>≈</mml:mo></mml:math></inline-formula> 0.13 and <inline-formula><mml:math id="M353" display="inline"><mml:mn mathvariant="normal">0.12</mml:mn></mml:math></inline-formula>) therefore derives <italic>entirely</italic> from the smooth, strongly autocorrelated tidal oscillation that its AR and MA terms track slightly better than naive persistence, and not from any ability to predict the sub-tidal residual that dominates the total error. This is corroborated by the Ljung–Box tests (Sect. <xref ref-type="sec" rid="Ch1.S3"/>), which show ARIMA residuals retain significant autocorrelation at the tidal periods, and by the strongly negative skill of the pure-harmonic UTide baseline (Table <xref ref-type="table" rid="T2"/>): the deterministic tide alone, without the autoregressive persistence term, is insufficient. These results directly confirm and qualify the tidal-autocorrelation interpretation – temporal autocorrelation explains ARIMA's skill, but that skill is small once the comparison is made over the full domain and period. (ARIMA is not listed as a row in Table <xref ref-type="table" rid="T7"/> because that table scores total-field predictions against the deterministic tidal field, whereas the per-cell autoregressive forecast is decomposed by re-fitting to each band as described here.)</p>
</sec>
<sec id="Ch1.S4.SS7">
  <label>4.7</label><title>Spatial error distribution</title>
      <p id="d2e6700">Figure <xref ref-type="fig" rid="F16"/> shows the per-cell RMSE for the three Phase 1 deep learning models. All three models exhibit a consistent spatial pattern: errors are lowest in the central domain (RMSE<inline-formula><mml:math id="M354" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>U</mml:mi></mml:msub><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:math></inline-formula>–<inline-formula><mml:math id="M355" display="inline"><mml:mn mathvariant="normal">8</mml:mn></mml:math></inline-formula> cm s<sup>−1</sup>, RMSE<inline-formula><mml:math id="M357" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>V</mml:mi></mml:msub><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula>–<inline-formula><mml:math id="M358" display="inline"><mml:mn mathvariant="normal">12</mml:mn></mml:math></inline-formula> cm s<sup>−1</sup>) and increase toward the strait boundaries, particularly in the northwest and southeast corners where current speeds and spatial gradients are highest. The <inline-formula><mml:math id="M360" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> component shows elevated errors along the northeast–southwest axis (the along-strait direction), consistent with the stronger tidal amplification of meridional currents.</p>

      <fig id="F16" specific-use="star"><label>Figure 16</label><caption><p id="d2e6781">Spatial distribution of per-cell RMSE for the three Phase 1 deep learning models. Top row: <inline-formula><mml:math id="M361" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> component. Bottom row: <inline-formula><mml:math id="M362" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> component. Higher errors occur near the strait boundaries where current gradients are steepest.</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f16.png"/>

        </fig>

      <p id="d2e6804">The spatial error pattern is remarkably similar across all three architectures, suggesting that the error distribution is governed by the underlying physical dynamics (tidal amplification near boundaries, coastal interactions) rather than by model-specific limitations. Mean per-cell correlation coefficients are high across all models (0.97 for <inline-formula><mml:math id="M363" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>, 0.93 for <inline-formula><mml:math id="M364" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>), with slightly lower values near the domain edges where the signal-to-noise ratio of the HF radar measurements is also lower.</p>
      <p id="d2e6822">To test whether this spatial structure is intrinsic to the physics or specific to the deep learning models, we computed the per-cell RMSE map of the pointwise ARIMA model and correlated it with the deep learning maps (Fig. <xref ref-type="fig" rid="F17"/>). The spatial patterns are nearly identical: the ARIMA error map correlates with the CNN, GRU, and CNN-GRU maps at <inline-formula><mml:math id="M365" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.96</mml:mn></mml:mrow></mml:math></inline-formula>–<inline-formula><mml:math id="M366" display="inline"><mml:mn mathvariant="normal">0.98</mml:mn></mml:math></inline-formula> for <inline-formula><mml:math id="M367" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M368" display="inline"><mml:mn mathvariant="normal">0.98</mml:mn></mml:math></inline-formula> for <inline-formula><mml:math id="M369" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>. A purely temporal, per-cell statistical model and the spatio-temporal networks therefore share the same error geography – low in the interior, high in the NW and SE corners – confirming that the spatial distribution of error is set by the tidal dynamics (amplification and steep gradients near the boundaries) rather than by model class. This is consistent with the expectation (Sect. <xref ref-type="sec" rid="Ch1.S5"/>) that a dedicated spatio-temporal statistical model would inherit the same error geography while improving the overall level of skill toward the neural networks.</p>

      <fig id="F17" specific-use="star"><label>Figure 17</label><caption><p id="d2e6874">Per-cell meridional RMSE for the pointwise ARIMA(1,0,1) model (left) and the spatio-temporal CNN-GRU (right). Despite ARIMA's higher overall error, the two maps share the same spatial structure (<inline-formula><mml:math id="M370" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">0.98</mml:mn></mml:mrow></mml:math></inline-formula>), indicating that the error geography is governed by the tidal dynamics rather than by model class.</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f17.png"/>

        </fig>

</sec>
<sec id="Ch1.S4.SS8">
  <label>4.8</label><title>Gap-filled versus observed performance</title>
      <p id="d2e6905">Because <inline-formula><mml:math id="M371" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">15</mml:mn></mml:mrow></mml:math></inline-formula> % of the full record is gap-filled with UTide tidal predictions (Sect. <xref ref-type="sec" rid="Ch1.S2"/>), and HF-radar gaps cluster in the low-quality boundary regions where errors are also highest, we examined whether the reported skill is contaminated by models reproducing the deterministic gap-fill. Over the test period the sea-cell gap fraction is 32.9 % (higher than the whole-record average because coverage is lower in this window), and it is higher at the 88 boundary cells (38.0 %) than in the interior (30.6 %). Splitting the test targets into genuinely-observed and gap-filled, every method has <italic>lower</italic> error on the gap-filled targets (e.g. CNN <inline-formula><mml:math id="M372" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>: 12.5 cm s<sup>−1</sup> gap-filled vs. 16.8 observed; persistence <inline-formula><mml:math id="M374" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>: 17.8 vs. 21.3; ARIMA <inline-formula><mml:math id="M375" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>: 15.9 vs. 20.3 cm s<sup>−1</sup>), because the UTide gap-fill is smooth and tidally deterministic and therefore easier to predict. Including gap-filled targets thus slightly <italic>deflates</italic> the reported RMSE; the observation-only values – the honest measure of skill against real data – are modestly higher (e.g. CNN <inline-formula><mml:math id="M377" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>: 16.8 vs. the 15.65 cm s<sup>−1</sup> all-target value). The ranking of methods is unchanged under either split. The per-cell gap fraction is mapped in Fig. <xref ref-type="fig" rid="F18"/>; comparison with the error map (Fig. <xref ref-type="fig" rid="F16"/>) confirms that the gap-rich boundary cells are also the high-error cells. We recommend the observation-only metric for operational assessment.</p>

      <fig id="F18"><label>Figure 18</label><caption><p id="d2e6998">Fraction of test-period targets that are gap-filled (UTide), per sea cell. The gap fraction is higher near the strait boundaries (38.0 % at edge cells) than in the interior (30.6 %), coinciding with the high-error regions in Fig. <xref ref-type="fig" rid="F16"/>.</p></caption>
          <graphic xlink:href="https://ascmo.copernicus.org/articles/12/221/2026/ascmo-12-221-2026-f18.png"/>

        </fig>

</sec>
<sec id="Ch1.S4.SS9">
  <label>4.9</label><title>Vector-valued error metrics</title>
      <p id="d2e7017">Because tidal currents rotate, evaluating <inline-formula><mml:math id="M379" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M380" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> separately does not reveal whether a model reproduces the <italic>direction</italic> of the flow – the quantity most relevant to trajectory forecasting and search-and-rescue. The Sunda Strait tidal current is strongly rectilinear along the NE–SW strait axis (Sect. <xref ref-type="sec" rid="Ch1.S4.SS7"/>), so the dominant rotary structure is a semidiurnal reversal along that axis rather than a smoothly rotating ellipse. We therefore complement the component RMSE with three vector-aware metrics over the test set (Table <xref ref-type="table" rid="T8"/>): the speed RMSE <inline-formula><mml:math id="M381" display="inline"><mml:msqrt><mml:mrow><mml:mo>〈</mml:mo><mml:mo>(</mml:mo><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">u</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:mo>-</mml:mo><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">u</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>〉</mml:mo></mml:mrow></mml:msqrt></mml:math></inline-formula>, the vector (complex) RMSE <inline-formula><mml:math id="M382" display="inline"><mml:msqrt><mml:mrow><mml:mo>〈</mml:mo><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">u</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:mi mathvariant="bold-italic">u</mml:mi><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>〉</mml:mo></mml:mrow></mml:msqrt></mml:math></inline-formula>, and the direction mean absolute error (evaluated where the observed speed exceeds 10 cm s<sup>−1</sup>).</p>

<table-wrap id="T8"><label>Table 8</label><caption><p id="d2e7119">Vector-valued error metrics over the test set (all sea cells): speed RMSE, vector (complex) RMSE, and direction mean absolute error (computed where observed speed <inline-formula><mml:math id="M384" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup>). The deep learning models reproduce the flow direction substantially better than persistence or ARIMA. Bold indicates the best value.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Method</oasis:entry>
         <oasis:entry colname="col2">Speed RMSE</oasis:entry>
         <oasis:entry colname="col3">Vector RMSE</oasis:entry>
         <oasis:entry colname="col4">Direction MAE</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">(cm s<sup>−1</sup>)</oasis:entry>
         <oasis:entry colname="col3">(cm s<sup>−1</sup>)</oasis:entry>
         <oasis:entry colname="col4">(°)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Persistence</oasis:entry>
         <oasis:entry colname="col2">19.68</oasis:entry>
         <oasis:entry colname="col3">25.81</oasis:entry>
         <oasis:entry colname="col4">16.1</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ARIMA(1,0,1)</oasis:entry>
         <oasis:entry colname="col2">18.46</oasis:entry>
         <oasis:entry colname="col3">24.14</oasis:entry>
         <oasis:entry colname="col4">14.8</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CNN</oasis:entry>
         <oasis:entry colname="col2">14.62</oasis:entry>
         <oasis:entry colname="col3">19.16</oasis:entry>
         <oasis:entry colname="col4">11.4</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">GRU</oasis:entry>
         <oasis:entry colname="col2">14.79</oasis:entry>
         <oasis:entry colname="col3">19.31</oasis:entry>
         <oasis:entry colname="col4">11.8</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CNN-GRU</oasis:entry>
         <oasis:entry colname="col2"><bold>14.59</bold></oasis:entry>
         <oasis:entry colname="col3"><bold>19.12</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>11.3</bold></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e7295">The deep learning models cut the direction error to <inline-formula><mml:math id="M388" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">11</mml:mn></mml:mrow></mml:math></inline-formula>–12°, against 16.1° for persistence and 14.8° for ARIMA, and similarly reduce the speed and vector RMSE by <inline-formula><mml:math id="M389" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">25</mml:mn></mml:mrow></mml:math></inline-formula> %. ARIMA exhibits a negative speed bias (<inline-formula><mml:math id="M390" display="inline"><mml:mo lspace="0mm">-</mml:mo></mml:math></inline-formula>3.5 cm s<sup>−1</sup>), under-predicting current magnitude as expected for a mean-reverting model, whereas persistence is approximately unbiased in speed. Thus the deep learning advantage seen in the component RMSE carries over to the operationally relevant full-vector error: the networks reproduce both the magnitude and the direction of the surface current more faithfully than the statistical baselines.</p>
</sec>
</sec>
<sec id="Ch1.S5">
  <label>5</label><title>Discussion</title>
<sec id="Ch1.S5.SS1">
  <label>5.1</label><title>Performance in energetic strait environments</title>
      <p id="d2e7355">The absolute RMSE values in this study are higher than those reported in open-sea environments: 4.5–7.4 cm s<sup>−1</sup> in the Gulf of Thailand <xref ref-type="bibr" rid="bib1.bibx32" id="paren.39"/>, 10–15 cm s<sup>−1</sup> in Monterey Bay <xref ref-type="bibr" rid="bib1.bibx12" id="paren.40"/>, and <inline-formula><mml:math id="M394" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">8</mml:mn></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup> along the US mid-Atlantic coast <xref ref-type="bibr" rid="bib1.bibx2" id="paren.41"/>. This reflects the much stronger tidal dynamics in the Sunda Strait (Table <xref ref-type="table" rid="T9"/>). When normalised by the observed speed range (defined here as the maximum minus minimum of the respective velocity component at each site), relative errors of 3 %–7 % are of the same order as those in other environments (Table <xref ref-type="table" rid="T9"/>). We stress that this normalisation is a first-order framing device, not evidence of superior performance: the low percentage is in large part a mathematical consequence of dividing by the Sunda Strait's very large speed range (nearly seven times that of the smallest comparison site), and the metric does not account for differences in spatial resolution, temporal resolution, prediction horizon, or the (study-dependent) definition of “speed range”. Given these methodological differences, the comparison should be read only as indicating that deep learning does not break down in an energetic strait, not as a quantitative ranking against the cited studies.</p>

<table-wrap id="T9" specific-use="star"><label>Table 9</label><caption><p id="d2e7421">Comparison with previous HF radar current prediction studies.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Study</oasis:entry>
         <oasis:entry colname="col2">Location</oasis:entry>
         <oasis:entry colname="col3">Speed range</oasis:entry>
         <oasis:entry colname="col4">RMSE</oasis:entry>
         <oasis:entry colname="col5">Rel. error</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">(cm s<sup>−1</sup>)</oasis:entry>
         <oasis:entry colname="col4">(cm s<sup>−1</sup>)</oasis:entry>
         <oasis:entry colname="col5">(%)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">
                    <xref ref-type="bibr" rid="bib1.bibx32" id="text.42"/>
                  </oasis:entry>
         <oasis:entry colname="col2">Gulf of Thailand</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M398" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">50</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">4.5–7.4</oasis:entry>
         <oasis:entry colname="col5">9–15</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">
                    <xref ref-type="bibr" rid="bib1.bibx12" id="text.43"/>
                  </oasis:entry>
         <oasis:entry colname="col2">Monterey Bay</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M399" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">80</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">10–15</oasis:entry>
         <oasis:entry colname="col5">13–19</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">
                    <xref ref-type="bibr" rid="bib1.bibx2" id="text.44"/>
                  </oasis:entry>
         <oasis:entry colname="col2">US mid-Atlantic</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M400" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">60</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M401" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">8</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M402" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">13</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">This study (1 step)</oasis:entry>
         <oasis:entry colname="col2">Sunda Strait</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M403" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">340</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">11.3–15.4</oasis:entry>
         <oasis:entry colname="col5">3–5</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">This study (6 h)</oasis:entry>
         <oasis:entry colname="col2">Sunda Strait</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M404" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">340</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">18.7–22.5</oasis:entry>
         <oasis:entry colname="col5">5–7</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S5.SS2">
  <label>5.2</label><title>Tidal cycle coverage in the lookback window</title>
      <p id="d2e7667">Phase 2 reveals that only the CNN-GRU benefits from longer lookbacks. This has a clear physical interpretation: the M2 tidal period is <inline-formula><mml:math id="M405" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">12.42</mml:mn></mml:mrow></mml:math></inline-formula> h, so <inline-formula><mml:math id="M406" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">12</mml:mn></mml:mrow></mml:math></inline-formula> provides the GRU with approximately one full tidal cycle. The CNN-GRU exploits this because the GRU naturally processes ordered sequences and can learn the tidal phase–amplitude relationship, while the CNN provides spatial context at each step. In contrast, the standalone CNN stacks all frames along the channel axis, destroying temporal ordering, and the standalone GRU receives flattened spatial fields, losing the spatial structure needed to benefit from longer sequences.</p>
      <p id="d2e7692">The modest improvement from <inline-formula><mml:math id="M407" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M408" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">12</mml:mn></mml:mrow></mml:math></inline-formula> for one-step prediction (1.2 % for <inline-formula><mml:math id="M409" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>) is consistent with the high autocorrelation at one-hour lag. The effect is amplified for multi-step prediction, where knowledge of the full tidal cycle provides phase information essential for forecasts extending several hours ahead.</p>
      <p id="d2e7726">One caveat is that temporal gaps (<inline-formula><mml:math id="M410" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">15</mml:mn></mml:mrow></mml:math></inline-formula> % of hours) are filled using UTide tidal harmonic predictions prior to model training (Sect. <xref ref-type="sec" rid="Ch1.S2"/>). While this ensures gap-free lookback sequences for all models, it could artificially inflate the apparent benefit of longer lookbacks by providing smoother (tidal-only) inputs at gap-filled timesteps. The effect is mitigated by the fact that the same gap-filling is applied across all lookback lengths, so the relative comparison remains valid. Future work could evaluate models robust to missing inputs as an alternative to gap filling.</p>
</sec>
<sec id="Ch1.S5.SS3">
  <label>5.3</label><title>Autoregressive versus direct multi-step prediction</title>
      <p id="d2e7749">A controlled comparison – all five multi-step models retrained over five seeds with an identical batch size, optimiser, and early-stopping criterion on a single GPU (Table <xref ref-type="table" rid="T4"/>) – shows that <italic>model capacity</italic>, not the direct-versus-autoregressive distinction, governs accuracy. Error decreases with capacity up to <inline-formula><mml:math id="M411" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> M parameters and then plateaus: CNN-GRU-MS-Small (1.00 M) is the most accurate (18.39 and 22.07 cm s<sup>−1</sup>) and is statistically tied with the 2.85 M full model. At matched capacity (<inline-formula><mml:math id="M413" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">0.48</mml:mn></mml:mrow></mml:math></inline-formula> M) the direct CNN-GRU-MS-Matched (19.08 and 22.53 cm s<sup>−1</sup>) and the autoregressive BiEF (19.00 and 22.28 cm s<sup>−1</sup>) are statistically indistinguishable – all differences lie within the five-seed standard deviation (<inline-formula><mml:math id="M416" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">0.2</mml:mn></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup>). Error accumulation in the autoregressive decoder is, at this six-hour horizon, offset by the autoregressive models' ability to focus each step on one-step-ahead prediction. The <inline-formula><mml:math id="M418" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> crossover – where autoregressive models slightly outperform CNN-GRU-MS – suggests that autoregressive decoding retains an advantage at the shortest lead. A hybrid strategy (autoregressive for <inline-formula><mml:math id="M419" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>, direct for longer leads) could in principle exploit this, but the operational complexity is unlikely to justify the marginal improvement. BiEF consistently outperforms ConvLSTM-ED, confirming that bidirectional encoding provides a richer initial state for the decoder.</p>
</sec>
<sec id="Ch1.S5.SS4">
  <label>5.4</label><title>The <inline-formula><mml:math id="M420" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> component is systematically harder to predict</title>
      <p id="d2e7876">Across all methods and phases, <inline-formula><mml:math id="M421" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> RMSE exceeds <inline-formula><mml:math id="M422" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> RMSE by 30 %–50 %. In the Sunda Strait (oriented NE–SW), the meridional component is approximately aligned with the along-strait direction where tidal amplification is strongest. The wider dynamic range (394 cm s<sup>−1</sup> for <inline-formula><mml:math id="M424" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> versus 339 cm s<sup>−1</sup> for <inline-formula><mml:math id="M426" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>) and steeper spatial gradients near coastal boundaries make <inline-formula><mml:math id="M427" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> inherently more difficult to predict.</p>
</sec>
<sec id="Ch1.S5.SS5">
  <label>5.5</label><title>Computational efficiency</title>
      <p id="d2e7948">The multi-step architectures achieve comparable accuracy at matched capacity, but they differ markedly in <italic>computational cost</italic> (Table <xref ref-type="table" rid="T5"/>), which is the decisive factor for operational deployment. Producing the full six-hour horizon in a single forward pass, the direct CNN-GRU-MS trains about 2.5 times faster (it converges within the early-stopping budget, whereas BiEF did not early-stop within the 20-epoch cap) and runs about 4.5 times faster at inference – 3.0 vs. 13.4 ms per forecast on a single CPU (Table <xref ref-type="table" rid="T5"/>) – because the autoregressive models must decode the six steps sequentially. For an hourly operational nowcast all models are fast enough in absolute terms (<inline-formula><mml:math id="M428" display="inline"><mml:mrow><mml:mo>≪</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> s), so inference latency is not itself a binding constraint; the direct model's advantages are its lower training cost (relevant for periodic retraining), its lower inference cost (relevant for scaling to larger grids, higher update frequencies, or constrained CPU-only hardware), and the fact that a single model covers all 291 cells – whereas the per-cell classical baselines (e.g. ARIMA) require fitting and maintaining one model per grid point. On this basis we recommend the 1.00 M CNN-GRU-MS-Small, which attains the best accuracy at the smallest size, for operational deployment.</p>
</sec>
<sec id="Ch1.S5.SS6">
  <label>5.6</label><title>Classical baselines and spatio-temporal statistical models</title>
      <p id="d2e7976">The pointwise classical baselines in this study (moving average, exponential smoothing, ARIMA, temporal <inline-formula><mml:math id="M429" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>NN) are deliberately <italic>univariate, per-cell</italic> models that ignore spatial correlation. To test directly whether their deficit relative to the neural networks is due to neglected spatial structure or to the absence of nonlinearity, we benchmarked them against an explicitly spatio-temporal statistical model: the EOF-VAR reduced-rank dynamic spatio-temporal model <xref ref-type="bibr" rid="bib1.bibx8 bib1.bibx34" id="paren.45"/> described in Sect. <xref ref-type="sec" rid="Ch1.S3"/>, which represents spatial covariance through empirical orthogonal functions and temporal dynamics through a vector autoregression on the leading principal components. EOF-VAR (SS<inline-formula><mml:math id="M430" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>U</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">0.39</mml:mn></mml:mrow></mml:math></inline-formula>, SS<inline-formula><mml:math id="M431" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>V</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.35</mml:mn></mml:mrow></mml:math></inline-formula>) substantially outperforms the pointwise ARIMA (SS<inline-formula><mml:math id="M432" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>U</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">0.13</mml:mn></mml:mrow></mml:math></inline-formula>, SS<inline-formula><mml:math id="M433" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>V</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.12</mml:mn></mml:mrow></mml:math></inline-formula>) and reaches the shallow-machine-learning tier, recovering most of the gap to the neural networks. This decisively attributes the classical methods' deficit to their neglect of spatial correlation: once a linear statistical model is allowed to share information across space, it becomes competitive. The neural networks nonetheless retain a consistent advantage (CNN-GRU 11.33/15.44 vs. EOF-VAR 12.50/16.27 cm s<sup>−1</sup>), which we attribute to their nonlinear representation of the spatio-temporal field; and consistently with this, the per-cell error <italic>geography</italic> is shared across the pointwise ARIMA and the spatio-temporal networks (<inline-formula><mml:math id="M435" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">0.96</mml:mn></mml:mrow></mml:math></inline-formula>–<inline-formula><mml:math id="M436" display="inline"><mml:mn mathvariant="normal">0.98</mml:mn></mml:math></inline-formula>; Sect. <xref ref-type="sec" rid="Ch1.S4.SS7"/>). Extending EOF-VAR or a full integro-difference-equation model to the multi-step setting of Phase 3, and incorporating nonlinear basis functions, are natural directions for future work.</p>
</sec>
<sec id="Ch1.S5.SS7">
  <label>5.7</label><title>Limitations</title>
      <p id="d2e8106">Several limitations should be noted. First, although the Phase-1 deep learning models and all five Phase-3 multi-step models are evaluated over five seeds (standard deviations 0.03–0.34 cm s<sup>−1</sup>), the Phase-2 lookback trends (0.06–0.14 cm s<sup>−1</sup>) are comparable to this variability and should be interpreted with caution. Second, the lead-time breakdown (Fig. <xref ref-type="fig" rid="F12"/>) is shown for a representative single run; the controlled multi-seed comparison is summarised by the lead-time-averaged values in Table <xref ref-type="table" rid="T4"/>, and the training-time and inference-latency figures (Table <xref ref-type="table" rid="T5"/>) are reported for the specific hardware listed and should be read as relative rather than absolute. Third, a substantial fraction of timesteps are gap-filled using UTide tidal predictions (<inline-formula><mml:math id="M439" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">15</mml:mn></mml:mrow></mml:math></inline-formula> % of the full record, but 32.9 % of sea-cell test targets; Sect. <xref ref-type="sec" rid="Ch1.S4.SS8"/>); models trained on these partially synthetic inputs may learn smoother-than-observed dynamics. As shown in Sect. <xref ref-type="sec" rid="Ch1.S4.SS8"/>, errors are lower on gap-filled targets for all methods, so all-target RMSE slightly understates the error against genuine observations; we therefore also report the observation-only RMSE and recommend it for operational assessment. Fourth, all models receive only the recent current history; incorporating external forcing (tidal predictions, wind, atmospheric pressure) could improve skill during sea-breeze-affected periods. Fifth, the study uses a single HF radar system covering a <inline-formula><mml:math id="M440" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">900</mml:mn></mml:mrow></mml:math></inline-formula> km<sup>2</sup> area; generalisation to other settings requires validation. Sixth, models predict at the 291 observed grid points only; extending predictions to the 137 blind-zone cells or beyond the radar footprint requires additional spatial reconstruction methods.</p>
</sec>
</sec>
<sec id="Ch1.S6" sec-type="conclusions">
  <label>6</label><title>Conclusions</title>
      <p id="d2e8182">This study presents a three-phase evaluation of forecasting methods for HF radar surface currents in the Sunda Strait, progressing from one-step prediction to six-hour nowcasting. The main findings are as follows.</p>
      <p id="d2e8185"><list list-type="order">
          <list-item>

      <p id="d2e8190">Among twelve methods for one-step prediction, evaluated consistently over all 291 cells and the full test period, the deep learning models achieve the highest skill for <italic>both</italic> components (CNN, SS<inline-formula><mml:math id="M442" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>U</mml:mi></mml:msub><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">0.50</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M443" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.97</mml:mn></mml:mrow></mml:math></inline-formula>; CNN-GRU, SS<inline-formula><mml:math id="M444" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>V</mml:mi></mml:msub><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.42</mml:mn></mml:mrow></mml:math></inline-formula>). Five-seed experiments confirm that CNN and CNN-GRU are statistically indistinguishable for <inline-formula><mml:math id="M445" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> (<inline-formula><mml:math id="M446" display="inline"><mml:mrow><mml:mn mathvariant="normal">11.32</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M447" display="inline"><mml:mrow><mml:mn mathvariant="normal">11.32</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.04</mml:mn></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup>), reflecting that the dominant semidiurnal signal renders the spatial-architecture choice nearly immaterial for one-step zonal prediction. The pointwise classical models (ARIMA, exponential smoothing, <inline-formula><mml:math id="M449" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>NN) exceed persistence but fall well below deep learning; a reduced-rank spatio-temporal statistical model (EOF-VAR) that accounts for spatial correlation recovers most of the gap (SS<inline-formula><mml:math id="M450" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>U</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.39</mml:mn></mml:mrow></mml:math></inline-formula>, SS<inline-formula><mml:math id="M451" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi>V</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.35</mml:mn></mml:mrow></mml:math></inline-formula>), showing that the classical deficit is dominated by the neglect of spatial structure rather than of nonlinearity.</p>
          </list-item>
          <list-item>

      <p id="d2e8326">CNN-GRU is the only architecture that improves with longer lookback (<inline-formula><mml:math id="M452" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula> to 12 h), demonstrating unique exploitation of tidal cycle information. Standalone CNN and GRU show no benefit.</p>
          </list-item>
          <list-item>

      <p id="d2e8346">For six-hour nowcasting, multi-step accuracy is governed by model capacity, plateauing near 1 M parameters (CNN-GRU-MS-Small: <inline-formula><mml:math id="M453" display="inline"><mml:mrow><mml:mn mathvariant="normal">18.39</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.15</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M454" display="inline"><mml:mrow><mml:mn mathvariant="normal">22.07</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.19</mml:mn></mml:mrow></mml:math></inline-formula> cm s<sup>−1</sup>; SS 0.77 and 0.72). Under a controlled five-seed comparison the direct and autoregressive architectures are statistically indistinguishable at matched capacity (<inline-formula><mml:math id="M456" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">0.48</mml:mn></mml:mrow></mml:math></inline-formula> M: direct 19.08/22.53 versus BiEF 19.00/22.28 cm s<sup>−1</sup>). The direct multi-step CNN-GRU-MS is nonetheless preferred for operational nowcasting because it trains <inline-formula><mml:math id="M458" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">2.5</mml:mn></mml:mrow></mml:math></inline-formula> times faster and runs <inline-formula><mml:math id="M459" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">4.5</mml:mn></mml:mrow></mml:math></inline-formula> times faster at inference (a single forward pass versus sequential decoding).</p>
          </list-item>
          <list-item>

      <p id="d2e8431">All models exhibit a diurnal error pattern. Concurrent AWS wind observations show that the zonal (<inline-formula><mml:math id="M460" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>) error increase correlates with afternoon sea-breeze enhancement (<inline-formula><mml:math id="M461" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">0.48</mml:mn></mml:mrow></mml:math></inline-formula> across the 24 diurnal-mean hours), while the meridional (<inline-formula><mml:math id="M462" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>) diurnal pattern is not wind-driven and likely reflects tidal phase effects.</p>
          </list-item>
          <list-item>

      <p id="d2e8468">Relative errors of 3 %–7 % of the observed speed range are of the same order as those in calmer environments, suggesting that deep learning generalises to energetic strait settings, though this normalisation partly reflects the strait's very large speed range and cross-site comparisons are further limited by differences in resolution and prediction horizon.</p>
          </list-item>
        </list></p>
      <p id="d2e8473">Beyond the site-specific results, four methodological insights may generalise to other HF radar forecasting applications: (i) tidal harmonic models, despite their physical basis, do not outperform persistence at one-step-ahead horizons where autocorrelation dominates, cautioning against the assumption that physics-based baselines are always superior; (ii) the lookback window length matters only when the architecture can jointly exploit spatial and temporal structure; (iii) at matched capacity, direct and autoregressive multi-step architectures achieve comparable accuracy – so a controlled, capacity-matched comparison rather than a headline RMSE is needed to compare architectures, and the direct model's advantage is computational (faster training and single-pass inference) rather than one of accuracy; and (iv) baselines must be evaluated on the same spatial domain and time span as the models they are compared against, since a spatially or temporally restricted evaluation can flatter a per-cell baseline and invert the apparent ranking.</p>
      <p id="d2e8476">The direct multi-step CNN-GRU-MS combines the best forecast skill with the lowest computational cost, making it a promising candidate for operational nowcasting in tidally dominated strait environments.</p>
</sec>

      
      </body>
    <back><notes notes-type="codeavailability"><title>Code availability</title>

      <p id="d2e8483">All model training, evaluation, and analysis code – together with the precomputed result files needed to reproduce every table and figure – is archived at Zenodo (<ext-link xlink:href="https://doi.org/10.5281/zenodo.20580835" ext-link-type="DOI">10.5281/zenodo.20580835</ext-link>, <xref ref-type="bibr" rid="bib1.bibx13" id="altparen.46"/>, released under the MIT License). The code is implemented in Python using PyTorch (CPU inference is sufficient; Intel Extension for PyTorch was used optionally for acceleration) and includes implementations of all twelve Phase 1 methods, the Phase 2 lookback sweep, the Phase 3 ConvLSTM-ED, BiEF, and CNN-GRU-MS architectures, and the referee-requested diagnostics (full-domain ARIMA with stationarity and residual tests, the extended ARIMA order search and seasonal ARIMA evaluation, the ARIMA tidal–residual decomposition, exponential smoothing, the EOF-VAR spatio-temporal statistical baseline, gap-fill stratification, vector error metrics, and the along-/cross-strait wind projection).</p>
  </notes><notes notes-type="dataavailability"><title>Data availability</title>

      <p id="d2e8495">The BADA HF radar data and AWS meteorological observations used in this study are owned and archived by the Indonesian Agency for Meteorology, Climatology, and Geophysics (BMKG). In accordance with the national regulation governing access to meteorological, climatological, and geophysical data (Peraturan Badan Meteorologi, Klimatologi, dan Geofisika Nomor 4 Tahun 2022; <uri>https://jdih.bmkg.go.id/dokumen/detail/4193</uri>, last access: 3 March 2026), these specific datasets cannot be made unconditionally public in open, FAIR-aligned repositories. Access to the underlying data for scientific validation and non-commercial research purposes can be granted upon reasonable request to the corresponding author, subject to the formal approval procedures and data-sharing agreements mandated by BMKG regulations. The BATNAS v1.6 bathymetry used to define the domain masks is openly available from the Indonesian Geospatial Information Agency (BIG, <uri>https://tanahair.indonesia.go.id/portal-web/unduh</uri>, last access: 3 March 2026).</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d2e8507">DG: conceptualisation, methodology, software, formal analysis, investigation, writing (original draft), visualisation. AAP: data curation, validation, writing (review and editing).</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d2e8513">The contact author has declared that neither of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d2e8519">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. The authors bear the ultimate responsibility for providing appropriate place names. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.</p>
  </notes><notes notes-type="sistatement"><title>Special issue statement</title>

      <p id="d2e8525">This article is part of the special issue “Artificial intelligence and machine learning in climate and weather science research”. It is not associated with a conference.</p>
  </notes><ack><title>Acknowledgements</title><p id="d2e8531">The authors gratefully acknowledge BMKG for operating the BADA HF radar system and providing the data.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d2e8537">Alifficionaldo A. Putra is supported by the Indonesia Endowment Fund for Education (LPDP), Ministry of Finance, Republic of Indonesia (grant no. 2025061211202566).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d2e8543">This paper was edited by Trevor Harris and reviewed by Gabriel Huerta and one anonymous referee.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Allard et al.(2014)Allard, Rogers, Martin, Jensen, Chu, Campbell, Dykes, Smith, Choi, and Gravois</label><mixed-citation>Allard, R., Rogers, E., Martin, P., Jensen, T., Chu, P., Campbell, T., Dykes, J., Smith, T., Choi, J., and Gravois, U.: The US Navy coupled ocean-wave prediction system, Oceanography, 27, 92–103, <ext-link xlink:href="https://doi.org/10.5670/oceanog.2014.71" ext-link-type="DOI">10.5670/oceanog.2014.71</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Barrick et al.(2012)Barrick, Fernandez, Ferrer, Whelan, and Breivik</label><mixed-citation>Barrick, D., Fernandez, V., Ferrer, M. I., Whelan, C., and Breivik, Ø.: A short-term predictive system for surface currents from a rapidly deployed coastal HF radar network, Ocean Dynam., 62, 725–740, <ext-link xlink:href="https://doi.org/10.1007/s10236-012-0521-0" ext-link-type="DOI">10.1007/s10236-012-0521-0</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Box et al.(2015)Box, Jenkins, Reinsel, and Ljung</label><mixed-citation> Box, G. E. P., Jenkins, G. M., Reinsel, G. C., and Ljung, G. M.: Time Series Analysis: Forecasting and Control, 5th edn., John Wiley &amp; Sons, ISBN 978-1-118-67502-1, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Chen and Chi(2021)</label><mixed-citation>Chen, P. and Chi, M.-Y.: STAGRU: Ocean surface current spatio-temporal prediction based on deep learning, in: Proc. Int. Conf. Computer Information Science and Artificial Intelligence (CISAI), 495–499, <ext-link xlink:href="https://doi.org/10.1109/CISAI54367.2021.00101" ext-link-type="DOI">10.1109/CISAI54367.2021.00101</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Cho et al.(2014)Cho, van Merriënboer, Gulcehre, Bahdanau, Bougares, Schwenk, and Bengio</label><mixed-citation>Cho, K., van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y.: Learning phrase representations using RNN encoder-decoder for statistical machine translation, in: Proc. 2014 Conf. Empirical Methods in Natural Language Processing (EMNLP),  1724–1734, <ext-link xlink:href="https://doi.org/10.3115/v1/D14-1179" ext-link-type="DOI">10.3115/v1/D14-1179</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Ciani et al.(2025)Ciani, Fanelli, and Buongiorno Nardelli</label><mixed-citation>Ciani, D., Fanelli, C., and Buongiorno Nardelli, B.: Estimating ocean currents from the joint reconstruction of absolute dynamic topography and sea surface temperature through deep learning algorithms, Ocean Sci., 21, 199–216, <ext-link xlink:href="https://doi.org/10.5194/os-21-199-2025" ext-link-type="DOI">10.5194/os-21-199-2025</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Codiga(2011)</label><mixed-citation>Codiga, D. L.: Unified tidal analysis and prediction using the UTide Matlab functions, Tech. Rep. 2011-01, Graduate School of Oceanography, University of Rhode Island, <ext-link xlink:href="https://doi.org/10.13140/RG.2.1.3761.2008" ext-link-type="DOI">10.13140/RG.2.1.3761.2008</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Cressie and Wikle(2011)</label><mixed-citation> Cressie, N. and Wikle, C. K.: Statistics for Spatio-Temporal Data, John Wiley &amp; Sons, ISBN 978-1-119-24304-5, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Diebold and Mariano(1995)</label><mixed-citation>Diebold, F. X. and Mariano, R. S.: Comparing predictive accuracy, J. Bus. Econ. Stat., 13, 253–263, <ext-link xlink:href="https://doi.org/10.1080/07350015.1995.10524599" ext-link-type="DOI">10.1080/07350015.1995.10524599</ext-link>, 1995.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Dong et al.(2022)Dong, Xu, Han, Bethel, Xie, and Zhou</label><mixed-citation>Dong, C., Xu, G., Han, G., Bethel, B. J., Xie, W., and Zhou, S.: Recent developments in artificial intelligence in oceanography, Ocean-Land-Atmosphere Research, 2022, 9870950, <ext-link xlink:href="https://doi.org/10.34133/2022/9870950" ext-link-type="DOI">10.34133/2022/9870950</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>El Aouni et al.(2025)El Aouni, Gaudel, Regnier, Van Gennip, Le Galloudec, Drevillon, Drillet, and Lellouche</label><mixed-citation>El Aouni, A., Gaudel, Q., Regnier, C., Van Gennip, S., Le Galloudec, O., Drevillon, M., Drillet, Y., and Lellouche, J.-M.: GLONET: Mercator's end-to-end neural global ocean forecasting system, Journal of Geophysical Research: Machine Learning and Computation, 2, <ext-link xlink:href="https://doi.org/10.1029/2025jh000686" ext-link-type="DOI">10.1029/2025jh000686</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Frolov et al.(2012)Frolov, Paduan, Cook, and Bellingham</label><mixed-citation>Frolov, S., Paduan, J., Cook, M., and Bellingham, J.: Improved statistical prediction of surface currents based on historic HF-radar observations, Ocean Dynam., 62, 1111–1122, <ext-link xlink:href="https://doi.org/10.1007/s10236-012-0553-5" ext-link-type="DOI">10.1007/s10236-012-0553-5</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Gautama and Putra(2026)</label><mixed-citation>Gautama, D. and Putra, A. A.: Analysis code for: Comparative evaluation of statistical and deep learning methods for high-frequency radar surface current forecasting in a narrow tropical strait (Version 1.2.0), Zenodo [computer software], <ext-link xlink:href="https://doi.org/10.5281/zenodo.20580835" ext-link-type="DOI">10.5281/zenodo.20580835</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>He et al.(2023)He, Zhou, Tian, Huang, Yang, Wang, and Huang</label><mixed-citation>He, S., Zhou, H., Tian, Y., Huang, D., Yang, J., Wang, C., and Huang, W.: Quality control for ocean current measurement using high-frequency direction-finding radar, Remote Sensing, 15, 5553, <ext-link xlink:href="https://doi.org/10.3390/rs15235553" ext-link-type="DOI">10.3390/rs15235553</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Hochreiter and Schmidhuber(1997)</label><mixed-citation>Hochreiter, S. and Schmidhuber, J.: Long short-term memory, Neural Comput., 9, 1735–1780, <ext-link xlink:href="https://doi.org/10.1162/neco.1997.9.8.1735" ext-link-type="DOI">10.1162/neco.1997.9.8.1735</ext-link>, 1997.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Jirakittayakorn et al.(2017)Jirakittayakorn, Kormongkolkul, Vateekul, Jitkajornwanich, and Lawawirojwong</label><mixed-citation>Jirakittayakorn, A., Kormongkolkul, T., Vateekul, P., Jitkajornwanich, K., and Lawawirojwong, S.: Temporal kNN for short-term ocean current prediction based on HF radar observations, in: Proc. 14th Int. Joint Conf. Computer Science and Software Engineering (JCSSE), IEEE, <ext-link xlink:href="https://doi.org/10.1109/JCSSE.2017.8025921" ext-link-type="DOI">10.1109/JCSSE.2017.8025921</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Kalinić et al.(2017)Kalinić, Mihanović, Cosoli, and Vilibić</label><mixed-citation>Kalinić, H., Mihanović, H., Cosoli, S., and Vilibić, I.: Predicting ocean surface currents using numerical weather prediction model and Kohonen neural network: A northern Adriatic study, Neural Computing and Applications, 28, 611–620, <ext-link xlink:href="https://doi.org/10.1007/s00521-016-2395-4" ext-link-type="DOI">10.1007/s00521-016-2395-4</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Kingma and Ba(2015)</label><mixed-citation>Kingma, D. P. and Ba, J.: Adam: A method for stochastic optimization, in: Proc. 3rd Int. Conf. Learning Representations (ICLR), arXiv, <ext-link xlink:href="https://doi.org/10.48550/arxiv.1412.6980" ext-link-type="DOI">10.48550/arxiv.1412.6980</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>LeCun et al.(1998)LeCun, Bottou, Bengio, and Haffner</label><mixed-citation>LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P.: Gradient-based learning applied to document recognition, P. IEEE, 86, 2278–2324, <ext-link xlink:href="https://doi.org/10.1109/5.726791" ext-link-type="DOI">10.1109/5.726791</ext-link>, 1998.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Li et al.(2018)Li, Wei, Susanto, Zhu, Setiawan, Xu, Fan, Agustiadi, Trenggono, and Fang</label><mixed-citation>Li, S., Wei, Z., Susanto, R. D., Zhu, Y., Setiawan, A., Xu, T., Fan, B., Agustiadi, T., Trenggono, M., and Fang, G.: Observations of intraseasonal variability in the Sunda Strait throughflow, J. Oceanogr., 74, 541–547, <ext-link xlink:href="https://doi.org/10.1007/s10872-018-0476-y" ext-link-type="DOI">10.1007/s10872-018-0476-y</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Liu et al.(2024)Liu, Zhang, Hao, Zhang, and Huang</label><mixed-citation>Liu, Y., Zhang, L., Hao, W., Zhang, L., and Huang, L.: Predicting temporal and spatial 4-D ocean temperature using satellite data based on a novel deep learning model, Ocean Model., 188, 102333, <ext-link xlink:href="https://doi.org/10.1016/j.ocemod.2024.102333" ext-link-type="DOI">10.1016/j.ocemod.2024.102333</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Mujiasih et al.(2021)Mujiasih, Hartanto, Beckers, and Barth</label><mixed-citation>Mujiasih, S., Hartanto, D., Beckers, J.-M., and Barth, A.: Reducing the error in estimates of the Sunda Strait currents by blending HF radar currents with model results, Cont. Shelf Res., 228, 104512, <ext-link xlink:href="https://doi.org/10.1016/j.csr.2021.104512" ext-link-type="DOI">10.1016/j.csr.2021.104512</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Murphy(1988)</label><mixed-citation>Murphy, A. H.: Skill scores based on the mean square error and their relationships to the correlation coefficient, Mon. Weather Rev., 116, 2417–2424, <ext-link xlink:href="https://doi.org/10.1175/1520-0493(1988)116&lt;2417:SSBOTM&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0493(1988)116&lt;2417:SSBOTM&gt;2.0.CO;2</ext-link>, 1988.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Paduan and Washburn(2013)</label><mixed-citation>Paduan, J. D. and Washburn, L.: High-frequency radar observations of ocean surface currents, Annu. Rev. Mar. Sci., 5, 115–136, <ext-link xlink:href="https://doi.org/10.1146/annurev-marine-121211-172315" ext-link-type="DOI">10.1146/annurev-marine-121211-172315</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Ren et al.(2018)Ren, Hu, and Hartnett</label><mixed-citation>Ren, L., Hu, Z., and Hartnett, M.: Short-term forecasting of coastal surface currents using high frequency radar data and artificial neural networks, Remote Sens., 10, 850, <ext-link xlink:href="https://doi.org/10.3390/rs10060850" ext-link-type="DOI">10.3390/rs10060850</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Roarty et al.(2024)Roarty, Updyke, Nazzaro, Smith, Glenn, and Schofield</label><mixed-citation>Roarty, H., Updyke, T., Nazzaro, L., Smith, M., Glenn, S., and Schofield, O.: Real-time quality assurance and quality control for a high frequency radar network, Frontiers in Marine Science, 11, 1352226, <ext-link xlink:href="https://doi.org/10.3389/fmars.2024.1352226" ext-link-type="DOI">10.3389/fmars.2024.1352226</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Schuster and Paliwal(1997)</label><mixed-citation>Schuster, M. and Paliwal, K. K.: Bidirectional recurrent neural networks, IEEE T. Signal Proces., 45, 2673–2681, <ext-link xlink:href="https://doi.org/10.1109/78.650093" ext-link-type="DOI">10.1109/78.650093</ext-link>, 1997.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Shi et al.(2015)Shi, Chen, Wang, Yeung, Wong, and Woo</label><mixed-citation>Shi, X., Chen, Z., Wang, H., Yeung, D.-Y., Wong, W.-K., and Woo, W.-C.: Convolutional LSTM network: A machine learning approach for precipitation nowcasting, in: Advances in Neural Information Processing Systems, vol. 28, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arxiv.1506.04214" ext-link-type="DOI">10.48550/arxiv.1506.04214</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>Sprintall et al.(2019)Sprintall, Gordon, Wijffels, Feng, Hu, Koch-Larrouy, Phillips, Nugroho, Napitu, Pujiana, Susanto, Sloyan, Peña Molino, Yuan, Riama, Siswanto, Kuswardani, Arifin, Wahyudi, Zhou, Nagai, Ansong, Bourdallé-Badié, Chanut, Lyard, Arbic, Ramdhani, and Setiawan</label><mixed-citation>Sprintall, J., Gordon, A. L., Wijffels, S. E., Feng, M., Hu, S., Koch-Larrouy, A., Phillips, H., Nugroho, D., Napitu, A., Pujiana, K., Susanto, R. D., Sloyan, B., Peña Molino, B., Yuan, D., Riama, N. F., Siswanto, S., Kuswardani, A., Arifin, Z., Wahyudi, A. J., Zhou, H., Nagai, T., Ansong, J. K., Bourdallé-Badié, R., Chanut, J., Lyard, F., Arbic, B. K., Ramdhani, A., and Setiawan, A.: Detecting change in the Indonesian Seas, Frontiers in Marine Science, 6, 257, <ext-link xlink:href="https://doi.org/10.3389/fmars.2019.00257" ext-link-type="DOI">10.3389/fmars.2019.00257</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>Susanto et al.(2016)Susanto, Wei, Adi, Zheng, Fang, Fan, Supangat, Agustiadi, Li, Trenggono, and Setiawan</label><mixed-citation>Susanto, R. D., Wei, Z., Adi, T. R., Zheng, Q., Fang, G., Fan, B., Supangat, A., Agustiadi, T., Li, S., Trenggono, M., and Setiawan, A.: Oceanography surrounding Krakatau Volcano in the Sunda Strait, Indonesia, Oceanography, 29, 264–272, <ext-link xlink:href="https://doi.org/10.5670/oceanog.2016.31" ext-link-type="DOI">10.5670/oceanog.2016.31</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Thongniran et al.(2019a)Thongniran, Jitkajornwanich, Lawawirojwong, Srestasathiern, and Vateekul</label><mixed-citation>Thongniran, N., Jitkajornwanich, K., Lawawirojwong, S., Srestasathiern, P., and Vateekul, P.: Combining attentional CNN and GRU networks for ocean current prediction based on HF radar observations, in: Proc. 8th Int. Conf. Computing and Pattern Recognition (ICCPR),  440–446, <ext-link xlink:href="https://doi.org/10.1145/3373509.3373549" ext-link-type="DOI">10.1145/3373509.3373549</ext-link>, 2019a.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Thongniran et al.(2019b)Thongniran, Vateekul, Jitkajornwanich, Lawawirojwong, and Srestasathiern</label><mixed-citation>Thongniran, N., Vateekul, P., Jitkajornwanich, K., Lawawirojwong, S., and Srestasathiern, P.: Spatio-temporal deep learning for ocean current prediction based on HF radar data, in: Proc. 16th Int. Joint Conf. Computer Science and Software Engineering (JCSSE), IEEE, 254–259, <ext-link xlink:href="https://doi.org/10.1109/JCSSE.2019.8864215" ext-link-type="DOI">10.1109/JCSSE.2019.8864215</ext-link>, 2019b.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Wei and Guan(2022)</label><mixed-citation>Wei, L. and Guan, L.: Seven-day sea surface temperature prediction using a 3DConv-LSTM model, Frontiers in Marine Science, 9, 905848, <ext-link xlink:href="https://doi.org/10.3389/fmars.2022.905848" ext-link-type="DOI">10.3389/fmars.2022.905848</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>Wikle et al.(2019)Wikle, Zammit-Mangion, and Cressie</label><mixed-citation>Wikle, C. K., Zammit-Mangion, A., and Cressie, N.: Spatio-Temporal Statistics with R, Chapman and Hall/CRC, <ext-link xlink:href="https://doi.org/10.1201/9781351769723" ext-link-type="DOI">10.1201/9781351769723</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx35"><label>Wyrtki(1961)</label><mixed-citation>Wyrtki, K.: Physical Oceanography of the Southeast Asian Waters, Scripps Institution of Oceanography, La Jolla, CA, nAGA Report Vol. 2, <uri>https://escholarship.org/uc/item/49n9x3t4</uri> (last access: 12 September 2026), 1961.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>Xiao et al.(2019)Xiao, Chen, Hu, Wang, Xu, Cai, Xu, Chen, and Gong</label><mixed-citation>Xiao, C., Chen, N., Hu, C., Wang, K., Xu, Z., Cai, Y., Xu, L., Chen, Z., and Gong, J.: A spatiotemporal deep learning model for sea surface temperature field prediction using time-series satellite data, Environ. Modell. Softw., 120, 104502, <ext-link xlink:href="https://doi.org/10.1016/j.envsoft.2019.104502" ext-link-type="DOI">10.1016/j.envsoft.2019.104502</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Zhang et al.(2020)Zhang, Geng, and Yan</label><mixed-citation>Zhang, K., Geng, X., and Yan, X.-H.: Prediction of 3-D ocean temperature by multilayer convolutional LSTM, IEEE Geosci. Remote S., 17, 1303–1307, <ext-link xlink:href="https://doi.org/10.1109/LGRS.2019.2947170" ext-link-type="DOI">10.1109/LGRS.2019.2947170</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Zhang et al.(2024)Zhang, Duan, Cui, Liu, and Huang</label><mixed-citation>Zhang, L., Duan, W., Cui, X., Liu, Y., and Huang, L.: Surface current prediction based on a physics-informed deep learning model, Appl. Ocean Res., 148, 104005, <ext-link xlink:href="https://doi.org/10.1016/j.apor.2024.104005" ext-link-type="DOI">10.1016/j.apor.2024.104005</ext-link>, 2024. </mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Zhang and Yin(2024)</label><mixed-citation>Zhang, Z. and Yin, J.: Spatial-temporal offshore current field forecasting using residual-learning based purely CNN methodology with attention mechanism, Appl. Artif. Intell., 38, 2323827, <ext-link xlink:href="https://doi.org/10.1080/08839514.2024.2323827" ext-link-type="DOI">10.1080/08839514.2024.2323827</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx40"><label>Zhao et al.(2024)Zhao, Peng, Wang, Li, Hou, and Zhong</label><mixed-citation>Zhao, Q., Peng, S., Wang, S., Li, Y., Hou, Y., and Zhong, G.: Applications of deep learning in physical oceanography: A comprehensive review, Frontiers in Marine Science, 11, 1396322, <ext-link xlink:href="https://doi.org/10.3389/fmars.2024.1396322" ext-link-type="DOI">10.3389/fmars.2024.1396322</ext-link>, 2024.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Comparative evaluation of statistical and deep learning methods for high-frequency radar surface current forecasting in a narrow tropical strait</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Allard et al.(2014)Allard, Rogers, Martin, Jensen, Chu, Campbell,
Dykes, Smith, Choi, and Gravois</label><mixed-citation>
      
Allard, R., Rogers, E., Martin, P., Jensen, T., Chu, P., Campbell, T., Dykes,
J., Smith, T., Choi, J., and Gravois, U.: The US Navy coupled ocean-wave
prediction system, Oceanography, 27, 92–103, <a href="https://doi.org/10.5670/oceanog.2014.71" target="_blank">https://doi.org/10.5670/oceanog.2014.71</a>,
2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Barrick et al.(2012)Barrick, Fernandez, Ferrer, Whelan, and
Breivik</label><mixed-citation>
      
Barrick, D., Fernandez, V., Ferrer, M. I., Whelan, C., and Breivik, Ø.: A
short-term predictive system for surface currents from a rapidly deployed
coastal HF radar network, Ocean Dynam., 62, 725–740,
<a href="https://doi.org/10.1007/s10236-012-0521-0" target="_blank">https://doi.org/10.1007/s10236-012-0521-0</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Box et al.(2015)Box, Jenkins, Reinsel, and Ljung</label><mixed-citation>
      
Box, G. E. P., Jenkins, G. M., Reinsel, G. C., and Ljung, G. M.: Time Series
Analysis: Forecasting and Control, 5th edn., John Wiley &amp; Sons, ISBN 978-1-118-67502-1, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Chen and Chi(2021)</label><mixed-citation>
      
Chen, P. and Chi, M.-Y.: STAGRU: Ocean surface current spatio-temporal
prediction based on deep learning, in: Proc. Int. Conf. Computer Information
Science and Artificial Intelligence (CISAI), 495–499,
<a href="https://doi.org/10.1109/CISAI54367.2021.00101" target="_blank">https://doi.org/10.1109/CISAI54367.2021.00101</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Cho et al.(2014)Cho, van Merriënboer, Gulcehre, Bahdanau,
Bougares, Schwenk, and Bengio</label><mixed-citation>
      
Cho, K., van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F.,
Schwenk, H., and Bengio, Y.: Learning phrase representations using RNN
encoder-decoder for statistical machine translation, in: Proc. 2014 Conf.
Empirical Methods in Natural Language Processing (EMNLP),  1724–1734,
<a href="https://doi.org/10.3115/v1/D14-1179" target="_blank">https://doi.org/10.3115/v1/D14-1179</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Ciani et al.(2025)Ciani, Fanelli, and
Buongiorno Nardelli</label><mixed-citation>
      
Ciani, D., Fanelli, C., and Buongiorno Nardelli, B.: Estimating ocean currents from the joint reconstruction of absolute dynamic topography and sea surface temperature through deep learning algorithms, Ocean Sci., 21, 199–216, <a href="https://doi.org/10.5194/os-21-199-2025" target="_blank">https://doi.org/10.5194/os-21-199-2025</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Codiga(2011)</label><mixed-citation>
      
Codiga, D. L.: Unified tidal analysis and prediction using the UTide Matlab
functions, Tech. Rep. 2011-01, Graduate School of Oceanography, University of
Rhode Island, <a href="https://doi.org/10.13140/RG.2.1.3761.2008" target="_blank">https://doi.org/10.13140/RG.2.1.3761.2008</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Cressie and Wikle(2011)</label><mixed-citation>
      
Cressie, N. and Wikle, C. K.: Statistics for Spatio-Temporal Data, John Wiley
&amp; Sons, ISBN 978-1-119-24304-5, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Diebold and Mariano(1995)</label><mixed-citation>
      
Diebold, F. X. and Mariano, R. S.: Comparing predictive accuracy, J.
Bus. Econ. Stat., 13, 253–263,
<a href="https://doi.org/10.1080/07350015.1995.10524599" target="_blank">https://doi.org/10.1080/07350015.1995.10524599</a>, 1995.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Dong et al.(2022)Dong, Xu, Han, Bethel, Xie, and Zhou</label><mixed-citation>
      
Dong, C., Xu, G., Han, G., Bethel, B. J., Xie, W., and Zhou, S.: Recent
developments in artificial intelligence in oceanography,
Ocean-Land-Atmosphere Research, 2022, 9870950, <a href="https://doi.org/10.34133/2022/9870950" target="_blank">https://doi.org/10.34133/2022/9870950</a>,
2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>El Aouni et al.(2025)El Aouni, Gaudel, Regnier, Van Gennip,
Le Galloudec, Drevillon, Drillet, and Lellouche</label><mixed-citation>
      
El Aouni, A., Gaudel, Q., Regnier, C., Van Gennip, S., Le Galloudec, O.,
Drevillon, M., Drillet, Y., and Lellouche, J.-M.: GLONET: Mercator's
end-to-end neural global ocean forecasting system, Journal of Geophysical
Research: Machine Learning and Computation, 2, <a href="https://doi.org/10.1029/2025jh000686" target="_blank">https://doi.org/10.1029/2025jh000686</a>,
2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Frolov et al.(2012)Frolov, Paduan, Cook, and Bellingham</label><mixed-citation>
      
Frolov, S., Paduan, J., Cook, M., and Bellingham, J.: Improved statistical
prediction of surface currents based on historic HF-radar observations,
Ocean Dynam., 62, 1111–1122, <a href="https://doi.org/10.1007/s10236-012-0553-5" target="_blank">https://doi.org/10.1007/s10236-012-0553-5</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Gautama and Putra(2026)</label><mixed-citation>
      
Gautama, D. and Putra, A. A.: Analysis code for: Comparative evaluation of statistical and deep learning methods for high-frequency radar surface current forecasting in a narrow tropical strait (Version 1.2.0), Zenodo [computer software], <a href="https://doi.org/10.5281/zenodo.20580835" target="_blank">https://doi.org/10.5281/zenodo.20580835</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>He et al.(2023)He, Zhou, Tian, Huang, Yang, Wang, and
Huang</label><mixed-citation>
      
He, S., Zhou, H., Tian, Y., Huang, D., Yang, J., Wang, C., and Huang, W.:
Quality control for ocean current measurement using high-frequency
direction-finding radar, Remote Sensing, 15, 5553, <a href="https://doi.org/10.3390/rs15235553" target="_blank">https://doi.org/10.3390/rs15235553</a>,
2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Hochreiter and Schmidhuber(1997)</label><mixed-citation>
      
Hochreiter, S. and Schmidhuber, J.: Long short-term memory, Neural Comput.,
9, 1735–1780, <a href="https://doi.org/10.1162/neco.1997.9.8.1735" target="_blank">https://doi.org/10.1162/neco.1997.9.8.1735</a>, 1997.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Jirakittayakorn et al.(2017)Jirakittayakorn, Kormongkolkul, Vateekul,
Jitkajornwanich, and Lawawirojwong</label><mixed-citation>
      
Jirakittayakorn, A., Kormongkolkul, T., Vateekul, P., Jitkajornwanich, K., and
Lawawirojwong, S.: Temporal kNN for short-term ocean current prediction
based on HF radar observations, in: Proc. 14th Int. Joint Conf. Computer
Science and Software Engineering (JCSSE), IEEE,
<a href="https://doi.org/10.1109/JCSSE.2017.8025921" target="_blank">https://doi.org/10.1109/JCSSE.2017.8025921</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Kalinić et al.(2017)Kalinić, Mihanović, Cosoli, and
Vilibić</label><mixed-citation>
      
Kalinić, H., Mihanović, H., Cosoli, S., and Vilibić, I.: Predicting
ocean surface currents using numerical weather prediction model and Kohonen
neural network: A northern Adriatic study, Neural Computing and
Applications, 28, 611–620, <a href="https://doi.org/10.1007/s00521-016-2395-4" target="_blank">https://doi.org/10.1007/s00521-016-2395-4</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Kingma and Ba(2015)</label><mixed-citation>
      
Kingma, D. P. and Ba, J.: Adam: A method for stochastic optimization, in: Proc.
3rd Int. Conf. Learning Representations (ICLR), arXiv,
<a href="https://doi.org/10.48550/arxiv.1412.6980" target="_blank">https://doi.org/10.48550/arxiv.1412.6980</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>LeCun et al.(1998)LeCun, Bottou, Bengio, and Haffner</label><mixed-citation>
      
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P.: Gradient-based learning
applied to document recognition, P. IEEE, 86, 2278–2324,
<a href="https://doi.org/10.1109/5.726791" target="_blank">https://doi.org/10.1109/5.726791</a>, 1998.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Li et al.(2018)Li, Wei, Susanto, Zhu, Setiawan, Xu, Fan, Agustiadi,
Trenggono, and Fang</label><mixed-citation>
      
Li, S., Wei, Z., Susanto, R. D., Zhu, Y., Setiawan, A., Xu, T., Fan, B.,
Agustiadi, T., Trenggono, M., and Fang, G.: Observations of intraseasonal
variability in the Sunda Strait throughflow, J. Oceanogr., 74,
541–547, <a href="https://doi.org/10.1007/s10872-018-0476-y" target="_blank">https://doi.org/10.1007/s10872-018-0476-y</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Liu et al.(2024)Liu, Zhang, Hao, Zhang, and Huang</label><mixed-citation>
      
Liu, Y., Zhang, L., Hao, W., Zhang, L., and Huang, L.: Predicting temporal and
spatial 4-D ocean temperature using satellite data based on a novel deep
learning model, Ocean Model., 188, 102333,
<a href="https://doi.org/10.1016/j.ocemod.2024.102333" target="_blank">https://doi.org/10.1016/j.ocemod.2024.102333</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Mujiasih et al.(2021)Mujiasih, Hartanto, Beckers, and
Barth</label><mixed-citation>
      
Mujiasih, S., Hartanto, D., Beckers, J.-M., and Barth, A.: Reducing the error
in estimates of the Sunda Strait currents by blending HF radar currents
with model results, Cont. Shelf Res., 228, 104512,
<a href="https://doi.org/10.1016/j.csr.2021.104512" target="_blank">https://doi.org/10.1016/j.csr.2021.104512</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Murphy(1988)</label><mixed-citation>
      
Murphy, A. H.: Skill scores based on the mean square error and their
relationships to the correlation coefficient, Mon. Weather Rev., 116,
2417–2424, <a href="https://doi.org/10.1175/1520-0493(1988)116&lt;2417:SSBOTM&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0493(1988)116&lt;2417:SSBOTM&gt;2.0.CO;2</a>, 1988.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Paduan and Washburn(2013)</label><mixed-citation>
      
Paduan, J. D. and Washburn, L.: High-frequency radar observations of ocean
surface currents, Annu. Rev. Mar. Sci., 5, 115–136,
<a href="https://doi.org/10.1146/annurev-marine-121211-172315" target="_blank">https://doi.org/10.1146/annurev-marine-121211-172315</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Ren et al.(2018)Ren, Hu, and Hartnett</label><mixed-citation>
      
Ren, L., Hu, Z., and Hartnett, M.: Short-term forecasting of coastal surface
currents using high frequency radar data and artificial neural networks,
Remote Sens., 10, 850, <a href="https://doi.org/10.3390/rs10060850" target="_blank">https://doi.org/10.3390/rs10060850</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Roarty et al.(2024)Roarty, Updyke, Nazzaro, Smith, Glenn, and
Schofield</label><mixed-citation>
      
Roarty, H., Updyke, T., Nazzaro, L., Smith, M., Glenn, S., and Schofield, O.:
Real-time quality assurance and quality control for a high frequency radar
network, Frontiers in Marine Science, 11, 1352226,
<a href="https://doi.org/10.3389/fmars.2024.1352226" target="_blank">https://doi.org/10.3389/fmars.2024.1352226</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Schuster and Paliwal(1997)</label><mixed-citation>
      
Schuster, M. and Paliwal, K. K.: Bidirectional recurrent neural networks, IEEE
T. Signal Proces., 45, 2673–2681, <a href="https://doi.org/10.1109/78.650093" target="_blank">https://doi.org/10.1109/78.650093</a>,
1997.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Shi et al.(2015)Shi, Chen, Wang, Yeung, Wong, and
Woo</label><mixed-citation>
      
Shi, X., Chen, Z., Wang, H., Yeung, D.-Y., Wong, W.-K., and Woo, W.-C.:
Convolutional LSTM network: A machine learning approach for precipitation
nowcasting, in: Advances in Neural Information Processing Systems, vol. 28, arXiv,
<a href="https://doi.org/10.48550/arxiv.1506.04214" target="_blank">https://doi.org/10.48550/arxiv.1506.04214</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Sprintall et al.(2019)Sprintall, Gordon, Wijffels, Feng, Hu,
Koch-Larrouy, Phillips, Nugroho, Napitu, Pujiana, Susanto, Sloyan, Peña
Molino, Yuan, Riama, Siswanto, Kuswardani, Arifin, Wahyudi, Zhou, Nagai,
Ansong, Bourdallé-Badié, Chanut, Lyard, Arbic, Ramdhani, and
Setiawan</label><mixed-citation>
      
Sprintall, J., Gordon, A. L., Wijffels, S. E., Feng, M., Hu, S., Koch-Larrouy,
A., Phillips, H., Nugroho, D., Napitu, A., Pujiana, K., Susanto, R. D.,
Sloyan, B., Peña Molino, B., Yuan, D., Riama, N. F., Siswanto, S.,
Kuswardani, A., Arifin, Z., Wahyudi, A. J., Zhou, H., Nagai, T., Ansong,
J. K., Bourdallé-Badié, R., Chanut, J., Lyard, F., Arbic, B. K.,
Ramdhani, A., and Setiawan, A.: Detecting change in the Indonesian Seas,
Frontiers in Marine Science, 6, 257, <a href="https://doi.org/10.3389/fmars.2019.00257" target="_blank">https://doi.org/10.3389/fmars.2019.00257</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Susanto et al.(2016)Susanto, Wei, Adi, Zheng, Fang, Fan, Supangat,
Agustiadi, Li, Trenggono, and Setiawan</label><mixed-citation>
      
Susanto, R. D., Wei, Z., Adi, T. R., Zheng, Q., Fang, G., Fan, B., Supangat,
A., Agustiadi, T., Li, S., Trenggono, M., and Setiawan, A.: Oceanography
surrounding Krakatau Volcano in the Sunda Strait, Indonesia,
Oceanography, 29, 264–272, <a href="https://doi.org/10.5670/oceanog.2016.31" target="_blank">https://doi.org/10.5670/oceanog.2016.31</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Thongniran et al.(2019a)Thongniran, Jitkajornwanich,
Lawawirojwong, Srestasathiern, and Vateekul</label><mixed-citation>
      
Thongniran, N., Jitkajornwanich, K., Lawawirojwong, S., Srestasathiern, P., and
Vateekul, P.: Combining attentional CNN and GRU networks for ocean
current prediction based on HF radar observations, in: Proc. 8th Int. Conf.
Computing and Pattern Recognition (ICCPR),  440–446,
<a href="https://doi.org/10.1145/3373509.3373549" target="_blank">https://doi.org/10.1145/3373509.3373549</a>, 2019a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Thongniran et al.(2019b)Thongniran, Vateekul,
Jitkajornwanich, Lawawirojwong, and Srestasathiern</label><mixed-citation>
      
Thongniran, N., Vateekul, P., Jitkajornwanich, K., Lawawirojwong, S., and
Srestasathiern, P.: Spatio-temporal deep learning for ocean current
prediction based on HF radar data, in: Proc. 16th Int. Joint Conf. Computer
Science and Software Engineering (JCSSE), IEEE, 254–259,
<a href="https://doi.org/10.1109/JCSSE.2019.8864215" target="_blank">https://doi.org/10.1109/JCSSE.2019.8864215</a>, 2019b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Wei and Guan(2022)</label><mixed-citation>
      
Wei, L. and Guan, L.: Seven-day sea surface temperature prediction using a
3DConv-LSTM model, Frontiers in Marine Science, 9, 905848,
<a href="https://doi.org/10.3389/fmars.2022.905848" target="_blank">https://doi.org/10.3389/fmars.2022.905848</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Wikle et al.(2019)Wikle, Zammit-Mangion, and
Cressie</label><mixed-citation>
      
Wikle, C. K., Zammit-Mangion, A., and Cressie, N.: Spatio-Temporal Statistics
with R, Chapman and Hall/CRC, <a href="https://doi.org/10.1201/9781351769723" target="_blank">https://doi.org/10.1201/9781351769723</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Wyrtki(1961)</label><mixed-citation>
      
Wyrtki, K.: Physical Oceanography of the Southeast Asian Waters, Scripps
Institution of Oceanography, La Jolla, CA, nAGA Report Vol. 2, <a href="https://escholarship.org/uc/item/49n9x3t4" target="_blank"/> (last access: 12 September 2026), 1961.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Xiao et al.(2019)Xiao, Chen, Hu, Wang, Xu, Cai, Xu, Chen, and
Gong</label><mixed-citation>
      
Xiao, C., Chen, N., Hu, C., Wang, K., Xu, Z., Cai, Y., Xu, L., Chen, Z., and
Gong, J.: A spatiotemporal deep learning model for sea surface temperature
field prediction using time-series satellite data, Environ. Modell.
Softw., 120, 104502, <a href="https://doi.org/10.1016/j.envsoft.2019.104502" target="_blank">https://doi.org/10.1016/j.envsoft.2019.104502</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Zhang et al.(2020)Zhang, Geng, and Yan</label><mixed-citation>
      
Zhang, K., Geng, X., and Yan, X.-H.: Prediction of 3-D ocean temperature by
multilayer convolutional LSTM, IEEE Geosci. Remote S.,
17, 1303–1307, <a href="https://doi.org/10.1109/LGRS.2019.2947170" target="_blank">https://doi.org/10.1109/LGRS.2019.2947170</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Zhang et al.(2024)Zhang, Duan, Cui, Liu, and
Huang</label><mixed-citation>
      
Zhang, L., Duan, W., Cui, X., Liu, Y., and Huang, L.: Surface current
prediction based on a physics-informed deep learning model, Appl. Ocean
Res., 148, 104005, <a href="https://doi.org/10.1016/j.apor.2024.104005" target="_blank">https://doi.org/10.1016/j.apor.2024.104005</a>, 2024.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Zhang and Yin(2024)</label><mixed-citation>
      
Zhang, Z. and Yin, J.: Spatial-temporal offshore current field forecasting
using residual-learning based purely CNN methodology with attention
mechanism, Appl. Artif. Intell., 38, 2323827,
<a href="https://doi.org/10.1080/08839514.2024.2323827" target="_blank">https://doi.org/10.1080/08839514.2024.2323827</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Zhao et al.(2024)Zhao, Peng, Wang, Li, Hou, and
Zhong</label><mixed-citation>
      
Zhao, Q., Peng, S., Wang, S., Li, Y., Hou, Y., and Zhong, G.: Applications of
deep learning in physical oceanography: A comprehensive review, Frontiers in
Marine Science, 11, 1396322, <a href="https://doi.org/10.3389/fmars.2024.1396322" target="_blank">https://doi.org/10.3389/fmars.2024.1396322</a>, 2024.

    </mixed-citation></ref-html>--></article>
