<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" dtd-version="3.0">
  <front>
    <journal-meta>
<journal-id journal-id-type="publisher">ASCMO</journal-id>
<journal-title-group>
<journal-title>Advances in Statistical Climatology, Meteorology and Oceanography</journal-title>
<abbrev-journal-title abbrev-type="publisher">ASCMO</abbrev-journal-title>
<abbrev-journal-title abbrev-type="nlm-ta">Adv. Stat. Clim. Meteorol. Oceanogr.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2364-3587</issn>
<publisher><publisher-name>Copernicus Publications</publisher-name>
<publisher-loc>Göttingen, Germany</publisher-loc>
</publisher>
</journal-meta>

    <article-meta>
      <article-id pub-id-type="doi">10.5194/ascmo-3-1-2017</article-id><title-group><article-title>Reconstruction of spatio-temporal temperature from sparse historical records using robust probabilistic principal component regression</article-title>
      </title-group><?xmltex \runningtitle{Robust principal component regression for sparse historical records}?><?xmltex \runningauthor{J.~Tipton et~al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Tipton</surname><given-names>John</given-names></name>
          <email>jtipton25@gmail.com</email>
        <ext-link>https://orcid.org/0000-0002-6135-8191</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2 aff3 aff1">
          <name><surname>Hooten</surname><given-names>Mevin</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff4">
          <name><surname>Goring</surname><given-names>Simon</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>Department of Statistics, Colorado State
University, Fort Collins, CO 80523, USA</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>U.S. Geological Survey, Colorado Cooperative Fish and
Wildlife Research Unit, Fort Collins, CO 80523, USA</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>Department of Fish, Wildlife, and Conservation Biology,
Colorado State University, Fort Collins,<?xmltex \hack{\newline}?> CO 80523, USA</institution>
        </aff>
        <aff id="aff4"><label>4</label><institution>Department of Geography, University of Wisconsin, Madison, WI 53706, USA</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">John Tipton (jtipton25@gmail.com)</corresp></author-notes><pub-date><day>27</day><month>January</month><year>2017</year></pub-date>
      
      <volume>3</volume>
      <issue>1</issue>
      <fpage>1</fpage><lpage>16</lpage>
      <history>
        <date date-type="received"><day>30</day><month>June</month><year>2016</year></date>
           <date date-type="rev-request"><year/></date>
           <date date-type="rev-recd"><day>21</day><month>October</month><year>2016</year></date>
           <date date-type="accepted"><day>21</day><month>October</month><year>2016</year></date>
      </history>
      <permissions>
<license license-type="open-access">
<license-p>This work is licensed under a Creative Commons Attribution 3.0 Unported License. To view a copy of this license, visit <ext-link ext-link-type="uri" xlink:href="http://creativecommons.org/licenses/by/3.0/">http://creativecommons.org/licenses/by/3.0/</ext-link></license-p>
</license>
</permissions><self-uri xlink:href="https://ascmo.copernicus.org/articles/.html">This article is available from https://ascmo.copernicus.org/articles/.html</self-uri>
<self-uri xlink:href="https://ascmo.copernicus.org/articles/.pdf">The full text article is available as a PDF file from https://ascmo.copernicus.org/articles/.pdf</self-uri>


      <abstract>
    <p>Scientific records of temperature and precipitation have been kept
for several hundred years, but for many areas, only a shorter record exists.
To understand climate change, there is a need for rigorous statistical
reconstructions of the paleoclimate using proxy data. Paleoclimate proxy data
are often sparse, noisy, indirect measurements of the climate process of
interest, making each proxy uniquely challenging to model statistically. We
reconstruct spatially explicit temperature surfaces from sparse and noisy
measurements recorded at historical United States military forts and other
observer stations from 1820 to 1894. One common method for reconstructing the
paleoclimate from proxy data is principal component regression (PCR). With
PCR, one learns a statistical relationship between the paleoclimate proxy
data and a set of climate observations that are used as patterns for
potential reconstruction scenarios. We explore PCR in a Bayesian hierarchical
framework, extending classical PCR in a variety of ways. First, we model the
latent principal components probabilistically, accounting for measurement
error in the observational data. Next, we extend our method to better
accommodate outliers that occur in the proxy data. Finally, we explore
alternatives to the truncation of lower-order principal components using
different regularization techniques. One fundamental challenge in
paleoclimate reconstruction efforts is the lack of out-of-sample data for
predictive validation. Cross-validation is of potential value, but is
computationally expensive and potentially sensitive to outliers in sparse
data scenarios. To overcome the limitations that a lack of out-of-sample
records presents, we test our methods using a simulation study, applying
proper scoring rules including a computationally efficient approximation to
leave-one-out cross-validation using the log score to validate model
performance. The result of our analysis is a spatially explicit
reconstruction of spatio-temporal temperature from a very sparse historical
record.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <title>Introduction</title>
      <p>There is a need for accurate estimates of paleoclimate, especially
temperature and precipitation, to better understand how climate has changed
in the past. Scientific measurements of temperature and precipitation have
been recorded for several hundred years, and in many locations for a much
shorter time. Because of long-standing interest in weather, there are a vast
number of anecdotal, nonscientific records of weather. However, many
reconstructions of paleoclimate using compiled historical records are not
amenable to direct statistical analysis because they consist of imprecise
measurements of weather reported in letters, newspapers, books, and other
documents
<xref ref-type="bibr" rid="bib1.bibx3 bib1.bibx29 bib1.bibx26 bib1.bibx5" id="paren.1"/>.
The large quantity of historical weather records, combined with appropriate
statistical models, has the potential to facilitate the extension of
scientific understanding of climate further back in time. Thus, there is a
need for a statistical framework that can model historical data compiled from
a variety of disparate sources by leveraging climate data from the recent
past.</p>
      <p>Historical observer weather data are often unreliable, sparse both temporally
and spatially, and noisy because these data were recorded before widespread
adoption of scientific measurement standards. As a result, historical
observer weather data have not been widely used for rigorous statistical
reconstructions of climate because these challenges make it difficult to
create generic statistical approaches for analysis. Historical observer
climate data can occur at hourly, daily, or monthly timescales, and the
current-era analog data used to train statistical models can also vary in
temporal resolution. Therefore, there is often a change of temporal support
between the historical observer and current-era analog data that must be
accounted for <xref ref-type="bibr" rid="bib1.bibx19" id="paren.2"/>.</p>
      <p>Another complication is that the true target one wishes to predict (the
historical, unobserved climate) is never available to evaluate model
predictive performance. Moreover, the historical observer data are often of
unknown or of varying reliability and are typically sparse, sometimes
involving only a few locations per year. The consequences of such data
characteristics for evaluating model performance are underexplored; hence, we
explore methods to validate historical observer-era model predictions under
these sparse data scenarios.</p>
      <p>We used spatially and temporally sparse historical observer measurements of
temperature recorded at United States (US) military forts and other
historical observer stations to reconstruct spatially explicit maps of mean
mid-day July temperature by leveraging modern spatially explicit current-era
analog data to impute missing spatial structure. We perform the
reconstruction within a model framework that accounts for uncertainty in
current-era data products and uncertainty in parameter estimation, and
properly evaluates predictive skill. We test eight model specifications using
a simulation study, generate predictions for mid-day July temperature at
approximately 20 000 locations for each year in 1820–1894 with associated
uncertainties, and evaluate model performance using a computationally
efficient approximation to leave-one-out cross-validation.</p>
</sec>
<sec id="Ch1.S2">
  <title>Data</title>
      <p>We used two datasets we refer to as the <italic>historical observer dataset</italic>
and the <italic>current-era analog</italic> dataset. The historical observer dataset
consists of temperature records from 1820 to 1894 at US forts in the Upper
Midwestern US as well as non-military observer stations. These data were
compiled as part of the Climate Database Modernization Program
(<xref ref-type="bibr" rid="bib1.bibx1" id="altparen.3"/>; CDMP 19th Century Forts and Voluntary
Observers Database Build Project:
<uri>http://www.isws.illinois.edu/atmos/clirecord.asp</uri>;
<xref ref-type="bibr" rid="bib1.bibx8" id="altparen.4"/>). At the observer stations, measurements were
recorded with time and date; however, the timing of measurements varied among
and within individual observer stations and was often temporally imprecise
(“daily min”, “daily max”, “mid-day”, etc.).</p>
      <p>Protocols varied across the observer stations through space and time, leading
to many irregularities in the historical observer data. Temperature
measurements were obtained by a variety of methods: some records report daily
minimum and maximum temperatures, others report hourly measurements, and
sometimes there are days or weeks with missing measurements. In addition, the
number and locations of the observer stations change through time, containing
between 1 and 234 locations per year; this variation is due to historical
events, including the Civil War and the westward expansion of the US in the
late 19th century. Most years have only a few observations and, in general,
the number of observer locations per year increases through time. Therefore,
the model must align the temporal and spatial scales of the two data sources
to reconstruct continuous temperature fields across the Upper Midwest. An
example of 4 years of historical data is shown in Fig. <xref ref-type="fig" rid="Ch1.F1"/>a.</p>
      <p>Because the historical observer data are spatially sparse, traditional
spatial statistical methods, such as Kriging, are not applicable, as these
methods require larger sample sizes to produce reasonable predictive
surfaces. Thus, we used the current-era analog data to provide spatial
structure for the reconstruction. For the current-era analog data, we used
the Parameter-elevation Relationships on Independent Slopes Model (PRISM)
monthly mean mid-day temperature surfaces created by interpolation of the US
Historical Climate Network (USHCN) data over the period 1895–2010
<xref ref-type="bibr" rid="bib1.bibx33" id="paren.5"/>. The PRISM data include 115 years of mean mid-day July
temperatures resolved to an 800 m<inline-formula><mml:math id="M1" display="inline"><mml:mrow><mml:mo>×</mml:mo><mml:mn>800</mml:mn></mml:mrow></mml:math></inline-formula> m grid, resulting in almost
20 000 spatial locations of interest in the study region
(Fig. <xref ref-type="fig" rid="Ch1.F1"/>b). Unlike the historical observer data, the PRISM
data are compiled from the USHCN and consist of commonly used model
interpolated temperature records. Other data products are available,
including high-quality data from satellite measurements; however, we used
PRISM for the current-era analog data due to the longer temporal coverage
that provides the model with more examples of the spatial structure of
mid-day July temperature. Because PRISM is a data product and not raw data,
we account for potential measurement errors in the current-era analog data
using our modeling framework.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><caption><p>Four years of the historical observer temperature data <bold>(a)</bold>
and the current-era analog temperature data <bold>(b)</bold>.</p></caption>
        <?xmltex \igopts{width=327.206693pt}?><graphic xlink:href="https://ascmo.copernicus.org/articles/3/1/2017/ascmo-3-1-2017-f01.pdf"/>

      </fig>

<sec id="Ch1.S2.SS1">
  <title>Temporal change of support</title>
      <p>To enable statistical learning about climate in the historical observer
period, we aligned the two data sources to common spatial and temporal
scales. We assigned each historical observer station to the closest grid cell
in the current-era analog data, thus accounting for any potential spatial
misalignment. Because the grid we aligned to is very fine scale
(800 m <inline-formula><mml:math id="M2" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 800 m grid cells) and temperature surfaces are generally
smooth over this spatial resolution, we assume any errors induced by the
spatial alignment are negligible relative to other sources of noise in the
data and ignore potential effects of spatial misalignment. Aligning the data
sources in time was more complicated because the historical observer station
data are highly irregular, whereas the current-era analog data are monthly
mean mid-day temperatures. We modeled the historical observer period mean
mid-day July temperature using cyclic cubic splines that are highly flexible,
able to accommodate the irregular nature of the historical data, and
constrained to reconstruct diurnal patterns <xref ref-type="bibr" rid="bib1.bibx46" id="paren.6"/>. We
focused on the month of July because the annual temperature curve peaks in
July and thus there is little/no seasonal change in temperature that needs to
be accounted for when computing a monthly average. The methodology could be
applied to other months, but the calibration in Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>) would
need to account for seasonal trend.</p>
      <p>We define our models using the following notation. Scalars are denoted by
lowercase letters, vectors are bold lowercase letters, and matrices are bold
uppercase letters. Fixed values, like data, are generally represented by
Latin letters and parameters are written in Greek letters. Using this
notation, the linear mixed model for estimating daily historical observer
mean mid-day July temperature is

                <disp-formula id="Ch1.E1" content-type="numbered"><mml:math id="M3" display="block"><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mfenced open="(" close=")"><mml:mi>s</mml:mi></mml:mfenced></mml:mrow><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mo>′</mml:mo></mml:msup><mml:mi mathvariant="bold-italic">α</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">η</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">η</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">η</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">ε</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mfenced close=")" open="("><mml:mi>s</mml:mi></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M4" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mfenced close=")" open="("><mml:mi>s</mml:mi></mml:mfenced></mml:mrow></mml:math></inline-formula> is the raw historical observer
temperature observation at location <inline-formula><mml:math id="M5" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>, year <inline-formula><mml:math id="M6" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, day <inline-formula><mml:math id="M7" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>, and hour <inline-formula><mml:math id="M8" display="inline"><mml:mi>s</mml:mi></mml:math></inline-formula>. The
covariate <inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the latitude at location <inline-formula><mml:math id="M10" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> and gives rise to a spatially
varying intercept for temperature parameterized by the coefficient <inline-formula><mml:math id="M11" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>.
The vector <inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">b</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is a cyclic cubic spline basis expansion of order
4 over the 24 h daily cycle with coefficients <inline-formula><mml:math id="M13" display="inline"><mml:mi mathvariant="bold-italic">α</mml:mi></mml:math></inline-formula> that
account for the diurnal pattern in temperature. The random effects
<inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">η</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">η</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">η</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> adjust the model fit with varying
intercepts for location <inline-formula><mml:math id="M17" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>, year <inline-formula><mml:math id="M18" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, and the interaction between location
and year. The model is completed by the inclusion of independent,
uncorrelated Gaussian error <inline-formula><mml:math id="M19" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ε</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mfenced open="(" close=")"><mml:mi>s</mml:mi></mml:mfenced></mml:mrow></mml:math></inline-formula>, giving rise to
interpolated daily temperature curves for July at each observer station
location <inline-formula><mml:math id="M20" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> and year <inline-formula><mml:math id="M21" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>. From the daily temperature curves, we estimated
mean mid-day July temperature by first predicting
<inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mfenced open="(" close=")"><mml:mover accent="true"><mml:mi>s</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover></mml:mfenced><mml:mo>=</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mover accent="true"><mml:mi mathvariant="italic">β</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>+</mml:mo><mml:mi mathvariant="bold">B</mml:mi><mml:mo>(</mml:mo><mml:mover accent="true"><mml:mi>s</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:msup><mml:mo>)</mml:mo><mml:mo>′</mml:mo></mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">α</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>+</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="italic">η</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="italic">η</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="italic">η</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> at 1 min intervals (<inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi>s</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>+</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mn>60</mml:mn></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mn>23</mml:mn><mml:mo>+</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn>59</mml:mn><mml:mn>60</mml:mn></mml:mfrac></mml:mstyle><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>) for each fort location and year. We estimated the mean
mid-day temperature using the same formula as the current-era analog data,
<inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfenced close=")" open="("><mml:munder><mml:mo movablelimits="false">min⁡</mml:mo><mml:mover accent="true"><mml:mi>s</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover></mml:munder><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mfenced close=")" open="("><mml:mover accent="true"><mml:mi>s</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover></mml:mfenced><mml:mo>+</mml:mo><mml:munder><mml:mo movablelimits="false">max⁡</mml:mo><mml:mover accent="true"><mml:mi>s</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover></mml:munder><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mfenced open="(" close=")"><mml:mover accent="true"><mml:mi>s</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover></mml:mfenced></mml:mfenced><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>, aligning
the sparse, irregular historical observer data to the monthly timescale of
the current-era analog data.</p>
      <p>To facilitate parameter estimation in the presence of sparse data, the
calibration model borrows strength among days, sites, and years within the
historical observer data for the month of July, reducing the influence of
measurement error and improving prediction of the mid-day diurnal temperature
curve. By borrowing strength, the calibration model produced a mean mid-day
estimate that has less variability than the raw historical observer data. We
fit the calibration model to the historical data using <monospace>R</monospace> package
<monospace>mgcv</monospace> <xref ref-type="bibr" rid="bib1.bibx47" id="paren.7"/> and refer to the pre-processed mid-day
estimates <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> as the historical observer data in what follows. We
justify the loss of information induced by using the calibration model
predictions instead of the raw historical observer data because the linear
mixed model explained approximately 70 % of the variability in the data
(<inline-formula><mml:math id="M26" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn>0.69</mml:mn></mml:mrow></mml:math></inline-formula>) and provided a mechanism for changing temporal support by
integrating uncertainty over the within-month mid-day temperature.</p>
</sec>
<sec id="Ch1.S2.SS2">
  <title>Modeling outline</title>
      <p>After aligning the two data sources to a common temporal scale, we
constructed a modeling framework to perform our
reconstruction. One method commonly used for the reconstruction of
paleoclimate is principal component regression (PCR), often called
empirical orthogonal function (EOF) regression in the paleoclimate
literature <xref ref-type="bibr" rid="bib1.bibx32" id="paren.8"/>. The use of PCR for the statistical
reconstruction of climate has a long tradition, dating back to
<xref ref-type="bibr" rid="bib1.bibx28" id="text.9"/>. In PCR reconstructions,
the climate proxy observations are regressed on a set of patterns
created from direct observations of the climate process. After learning
about the regression parameters, the model is used to predict climate at
the unobserved locations.</p>
      <p>To build our spatio-temporal predictive model, we used traditional principal
component regression (PCR) as well as probabilistic principal component
regression (pPCR) that assumes the empirical principal components are a noisy
measure of the true, latent principal components
<xref ref-type="bibr" rid="bib1.bibx39" id="paren.10"/>. We explore the temporal PCR and pPCR models
in a Bayesian hierarchical framework using regularization methods to select
important principal components for each year's reconstruction. Within this
framework, we assign hierarchical pooling priors to improve parameter
estimation for years with few observations by borrowing strength from years
with many observations <xref ref-type="bibr" rid="bib1.bibx13" id="paren.11"/>. We also develop robust,
Student's <inline-formula><mml:math id="M27" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> specifications of PCR and pPCR models that accommodate
potentially outlying measurements of mid-day July temperature in the
historical observer data that may have arisen from the non-standardized data
collection.</p>
      <p>We introduce traditional PCR in Sect. <xref ref-type="sec" rid="Ch1.S3.SS1"/> within a temporal
framework that allows for flexibility among years while borrowing strength
among years to improve estimation in years with few observations and define
the probabilistic extension (pPCR) of PCR that accounts for measurement error
in Sect. <xref ref-type="sec" rid="Ch1.S3.SS2"/>. In Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>, we introduce the robust
specification of our PCR and pPCR models that better accommodate outlying
observations, and in Sect. <xref ref-type="sec" rid="Ch1.S3.SS4"/>, we show how to improve
computation by integrating out the latent principal components in the pPCR
model. We describe three scoring rules to validate model performance in
Sect. <xref ref-type="sec" rid="Ch1.S4"/>, and describe a simulation study in
Sect. <xref ref-type="sec" rid="Ch1.S5"/> where we evaluate predictive performance in a
synthetic data scenario. In Sect. <xref ref-type="sec" rid="Ch1.S6"/>, we apply our models to
reconstruct historical mean mid-day July temperature in the Upper Midwestern
US, choosing the model that performs best based on scoring rules.</p>
</sec>
</sec>
<sec id="Ch1.S3">
  <title>Model statement</title>
<sec id="Ch1.S3.SS1">
  <title>Principal component regression</title>
      <p>A common statistical approach for reconstruction of the historical climate
using current-era analog data is to regress the partially observed historical
observer data onto the current-era analog observations. For a given
reconstruction year, define the regression model

                <disp-formula id="Ch1.E2" content-type="numbered"><mml:math id="M28" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi><mml:mo>′</mml:mo></mml:msubsup><mml:msub><mml:mi mathvariant="bold-italic">α</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is a historical observer period observation (pre-processed
using calibration model Eq. <xref ref-type="disp-formula" rid="Ch1.E1"/>) of mean mid-day July temperature
at location <inline-formula><mml:math id="M30" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> for year <inline-formula><mml:math id="M31" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>. The vector <inline-formula><mml:math id="M32" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> consists of the
<inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> historical observer period observations of the temperature field for
year <inline-formula><mml:math id="M34" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, where we observe only <inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> out of the <inline-formula><mml:math id="M36" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> locations, with the
number and locations of observations changing through time
(Figs. <xref ref-type="fig" rid="Ch1.F1"/>a and <xref ref-type="fig" rid="Ch1.F7"/>b). The columns of the <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:mi>n</mml:mi><mml:mo>×</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:math></inline-formula> matrix <inline-formula><mml:math id="M38" display="inline"><mml:mi mathvariant="bold">X</mml:mi></mml:math></inline-formula> contain <inline-formula><mml:math id="M39" display="inline"><mml:mi>d</mml:mi></mml:math></inline-formula> replicates of the current-era analog
temperature surfaces at the <inline-formula><mml:math id="M40" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> locations in the domain of interest, forming
a basis set of patterns for the regression, where <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi><mml:mo>′</mml:mo></mml:msubsup></mml:mrow></mml:math></inline-formula> represents the
<inline-formula><mml:math id="M42" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>th row of <inline-formula><mml:math id="M43" display="inline"><mml:mi mathvariant="bold">X</mml:mi></mml:math></inline-formula>. A greater number of replicates of current-era
analog temperature surfaces <inline-formula><mml:math id="M44" display="inline"><mml:mi>d</mml:mi></mml:math></inline-formula> gives a larger set of potential spatial
patterns that can be used to learn about spatial patterns in the historical
observer data. The <inline-formula><mml:math id="M45" display="inline"><mml:mi>d</mml:mi></mml:math></inline-formula>-dimensional vector of regression coefficients
<inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">α</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> link the historical observer data <inline-formula><mml:math id="M47" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> with
the set of climate patterns <inline-formula><mml:math id="M48" display="inline"><mml:mi mathvariant="bold">X</mml:mi></mml:math></inline-formula> in the current-era analog data for
each year <inline-formula><mml:math id="M49" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, allowing for climate fields that are linear combinations of
observed current-era analogs, up to uncorrelated model error. Thus, we can
model temperatures that are warmer or cooler than the current-era analog
period, but patterns that are not linear combinations of the current-era
analogs are not accommodated in the model, necessitating use of a
sufficiently long temporal record of current-era analogs. The uncorrelated
model error <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is assumed to be independent and identically
distributed Gaussian with variance <inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula>, prior <inline-formula><mml:math id="M52" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mtext>log</mml:mtext><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, and vague hyperpriors <inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mtext>N</mml:mtext><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mi>U</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Because the likelihood is
unaffected if <inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is integrated out, we assume that the data <inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>
are centered and assume <inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> (i.e., anomalies).</p>
      <p>In Eq. (<xref ref-type="disp-formula" rid="Ch1.E2"/>), the columns in <inline-formula><mml:math id="M58" display="inline"><mml:mi mathvariant="bold">X</mml:mi></mml:math></inline-formula> are highly
multicollinear. Multicollinearity inflates the coefficient estimate variance
and, in cases of severe multicollinearity, the least squares solution is
nearly singular, causing algorithm instability and unreliable estimation. One
could use this model to estimate the dynamics influencing a given year's
temperature surface by interpreting the estimated regression coefficients,
but because we are interested in prediction of the dependent variable
<inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and less interested in interpretation of the regression
coefficients, we manipulate the form of <inline-formula><mml:math id="M60" display="inline"><mml:mi mathvariant="bold">X</mml:mi></mml:math></inline-formula> to improve statistical
learning. We begin by computing the singular value decomposition (SVD) of
<inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:mi mathvariant="bold">X</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="bold">U</mml:mi><mml:mi mathvariant="bold">Λ</mml:mi><mml:msup><mml:mi mathvariant="bold">V</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula>, where the columns
of <inline-formula><mml:math id="M62" display="inline"><mml:mi mathvariant="bold">U</mml:mi></mml:math></inline-formula> are the left singular vectors of <inline-formula><mml:math id="M63" display="inline"><mml:mi mathvariant="bold">X</mml:mi></mml:math></inline-formula>, the diagonal
matrix <inline-formula><mml:math id="M64" display="inline"><mml:mi mathvariant="bold">Λ</mml:mi></mml:math></inline-formula> has the singular values in descending order on
the diagonal, and the columns of <inline-formula><mml:math id="M65" display="inline"><mml:mi mathvariant="bold">V</mml:mi></mml:math></inline-formula> are the right singular vectors.
The PCR model using the SVD is

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M66" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">u</mml:mi><mml:mi>i</mml:mi><mml:mo>′</mml:mo></mml:msubsup><mml:mi mathvariant="bold">Λ</mml:mi><mml:msup><mml:mi mathvariant="bold">V</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:msub><mml:mi mathvariant="bold-italic">α</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">u</mml:mi><mml:mi>i</mml:mi><mml:mo>′</mml:mo></mml:msubsup><mml:msup><mml:mi mathvariant="bold">Λ</mml:mi><mml:mfrac><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:msup><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E3"><mml:mtd/><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi><mml:mo>′</mml:mo></mml:msubsup><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">u</mml:mi><mml:mi>i</mml:mi><mml:mo>′</mml:mo></mml:msubsup></mml:mrow></mml:math></inline-formula> is the <inline-formula><mml:math id="M68" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>th row of <inline-formula><mml:math id="M69" display="inline"><mml:mi mathvariant="bold">U</mml:mi></mml:math></inline-formula>. If the regression
coefficient is given the prior <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">α</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mtext>N</mml:mtext><mml:mo>(</mml:mo><mml:mn mathvariant="bold">0</mml:mn><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mi mathvariant="bold">I</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, then
<inline-formula><mml:math id="M71" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="bold">Λ</mml:mi><mml:mfrac><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:msup><mml:msup><mml:mi mathvariant="bold">V</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:msub><mml:mi mathvariant="bold-italic">α</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="bold">0</mml:mn><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:mi mathvariant="italic">β</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mi mathvariant="bold">Λ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> where <inline-formula><mml:math id="M72" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:mi mathvariant="italic">β</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula>. In
this model, the columns of <inline-formula><mml:math id="M73" display="inline"><mml:mi mathvariant="bold">U</mml:mi></mml:math></inline-formula> are the eigenvectors of
<inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold">X</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mi mathvariant="bold">X</mml:mi></mml:mrow></mml:math></inline-formula>, the diagonal elements of <inline-formula><mml:math id="M75" display="inline"><mml:mi mathvariant="bold">Λ</mml:mi></mml:math></inline-formula> are
the eigenvalues of <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold">X</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mi mathvariant="bold">X</mml:mi></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi><mml:mo>′</mml:mo></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">u</mml:mi><mml:mi>i</mml:mi><mml:mo>′</mml:mo></mml:msubsup><mml:msup><mml:mi mathvariant="bold">Λ</mml:mi><mml:mfrac><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:msup></mml:mrow></mml:math></inline-formula> is the scaled principal component at
location <inline-formula><mml:math id="M78" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>. In applications of PCR, one often performs dimension reduction
by retaining only the first <inline-formula><mml:math id="M79" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> eigenvectors of <inline-formula><mml:math id="M80" display="inline"><mml:mi mathvariant="bold">U</mml:mi></mml:math></inline-formula> in the <inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:mi>n</mml:mi><mml:mo>×</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:math></inline-formula> matrix <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">U</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and the first <inline-formula><mml:math id="M83" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> eigenvalues in the <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>×</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:math></inline-formula>
matrix <inline-formula><mml:math id="M85" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Λ</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. After truncation, the truncated PCR design
matrix is <inline-formula><mml:math id="M86" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Z</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold">U</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:msubsup><mml:mi mathvariant="bold">Λ</mml:mi><mml:mi>p</mml:mi><mml:mfrac><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:msubsup></mml:mrow></mml:math></inline-formula>.
Typically, <inline-formula><mml:math id="M87" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> is chosen by cross-validation or by choosing the smallest <inline-formula><mml:math id="M88" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula>
so that the proportion of variability explained in the model is a large
value. Because truncation removes the highest-frequency eigenvectors, the
truncation implies a prior that shrinks the regression coefficients and
provides an implicit regularization on the model <xref ref-type="bibr" rid="bib1.bibx21" id="paren.12"/>.
Although truncation of lower-order principal components disregards
small-scale variability and therefore can only reduce the theoretical minimum
prediction error, the truncated model is often more computationally stable
than Eq. (<xref ref-type="disp-formula" rid="Ch1.E2"/>) and can improve prediction in practice.</p>
      <p>Preexisting research suggests that truncation of the trailing principal
components is not always appropriate because the higher-frequency components
are often important predictors <xref ref-type="bibr" rid="bib1.bibx20 bib1.bibx24" id="paren.13"/>. In our
paleoclimate reconstruction method, inclusion of lower-order principal
components is important, especially if there are climate signals that are
slowly varying or show up occasionally (i.e., every decade or century). If
these uncommon processes appear in the lesser eigenvectors of the current-era
analog data (which is likely because such processes are not the primary
contributors to the annual-scale variability in climate), these signals would
be discarded by truncation as high-frequency noise. Allowing the important
principal components in the regression to vary with time, the model is
capable of detecting the changes in temperature. For example, the irregular
but periodic cycles of the Pacific Decadal Oscillation and the Atlantic
Multidecadal Oscillation likely do not explain a large portion of the
variability in the temperature records. Therefore, these and other similar
climate signals could be removed through the truncation of the principal
components. Ideally, one chooses the truncation <inline-formula><mml:math id="M89" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> to be as large as
computationally possible and then performs a variable selection or
regularization method to select important principal components.
<xref ref-type="bibr" rid="bib1.bibx44" id="text.14"/> approached the problem of choosing the important
principal components through the Bayesian model selection technique known as
stochastic search variable selection (SSVS; <xref ref-type="bibr" rid="bib1.bibx15" id="altparen.15"/>; see
<xref ref-type="bibr" rid="bib1.bibx23" id="altparen.16"/>, for a review). The SSVS variable selection assumes
the hierarchical prior on the <inline-formula><mml:math id="M90" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th regression coefficient at time <inline-formula><mml:math id="M91" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>

                <disp-formula id="Ch1.E4" content-type="numbered"><mml:math id="M92" display="block"><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>∼</mml:mo><mml:mfenced open="{" close=""><mml:mtable class="cases" rowspacing="0.2ex" columnspacing="1em" columnalign="left left" framespacing="0em"><mml:mtr><mml:mtd><mml:mrow><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:msub><mml:mi mathvariant="italic">λ</mml:mi><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext> if </mml:mtext><mml:msub><mml:mi mathvariant="italic">ξ</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:msubsup><mml:mi mathvariant="italic">κ</mml:mi><mml:mi>j</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msubsup><mml:msub><mml:mi mathvariant="italic">λ</mml:mi><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext> if </mml:mtext><mml:msub><mml:mi mathvariant="italic">ξ</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mfenced></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M93" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">λ</mml:mi><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the <inline-formula><mml:math id="M94" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th diagonal element of
<inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Λ</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. The SSVS prior for the pPCR model is similar, but
does not include the <inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">λ</mml:mi><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> terms. The variables <inline-formula><mml:math id="M97" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ξ</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> are
indicators of the importance of the <inline-formula><mml:math id="M98" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th latent principal component in the
regression for year <inline-formula><mml:math id="M99" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> and have independent Bernoulli(<inline-formula><mml:math id="M100" display="inline"><mml:mn>0.5</mml:mn></mml:math></inline-formula>) priors. The
regression coefficient variance <inline-formula><mml:math id="M101" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> could be assigned a
prior if desired, but the shrinkage value <inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">κ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> must be fixed. A
large <inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">κ</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> produces a mixture distribution of a broad, relatively
uninformative prior with large variance <inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> (the “slab”)
and a highly informative prior at a small neighborhood around zero (the
“spike”) that provides shrinkage by truncating less important principal
components using probabilistic learning, thus reducing the chance of omitting
important principal components while avoiding the computationally expensive
task of exploring all <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">2</mml:mn><mml:mi>p</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> possible model configurations. We set <inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">κ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>1000</mml:mn></mml:mrow></mml:math></inline-formula> and pool across years by assuming the hierarchical model with prior
<inline-formula><mml:math id="M107" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>∼</mml:mo><mml:mtext>log</mml:mtext><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">β</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">β</mml:mi></mml:msub></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and vague hyperpriors <inline-formula><mml:math id="M108" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">β</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>∼</mml:mo><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">β</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>∼</mml:mo><mml:mtext>U</mml:mtext><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
      <p>An alternative to variable selection methods like SSVS is penalized
regression <xref ref-type="bibr" rid="bib1.bibx21" id="paren.17"/>. Common forms of penalized regression
include ridge regression (Tikhonov or <inline-formula><mml:math id="M110" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> shrinkage;
<xref ref-type="bibr" rid="bib1.bibx22" id="altparen.18"/>), where one minimizes

                <disp-formula id="Ch1.Ex3"><mml:math id="M111" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi><mml:mo>′</mml:mo></mml:msubsup><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">γ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>p</mml:mi></mml:munderover><mml:msubsup><mml:mi mathvariant="italic">β</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></disp-formula>

          with respect to <inline-formula><mml:math id="M112" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and the least angle subset selection
operator (LASSO or <inline-formula><mml:math id="M113" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> shrinkage; <xref ref-type="bibr" rid="bib1.bibx36" id="altparen.19"/>),
which minimizes

                <disp-formula id="Ch1.Ex4"><mml:math id="M114" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi><mml:mo>′</mml:mo></mml:msubsup><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">γ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>p</mml:mi></mml:munderover><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:math></disp-formula>

          with respect to <inline-formula><mml:math id="M115" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> given the penalty term <inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">γ</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. The
<inline-formula><mml:math id="M117" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> penalty shrinks the coefficients non-linearly toward zero and the <inline-formula><mml:math id="M118" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>
penalty shrinks large coefficients linearly, but in a way that the
coefficients can equal zero exactly. When viewed from this perspective, the
LASSO can be viewed as a compromise between regularization and variable
selection methods because, as the coefficients in the LASSO model approach
zero, there is nonzero probability that the LASSO will shrink the covariate
estimates to zero, thereby removing that variable from the model
<xref ref-type="bibr" rid="bib1.bibx10" id="paren.20"/>. We apply both SSVS and LASSO shrinkage methods to
explore the empirical consequences of the choice of regularizer. One drawback
to regularization methods is the need to estimate the penalty parameter
<inline-formula><mml:math id="M119" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">γ</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Often, the optimal <inline-formula><mml:math id="M120" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">γ</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is determined by cross-validation
using predictive skill. In the Bayesian framework, the shrinkage can be
estimated by cross-validation or by assigning a prior distribution and
performing a fully Bayesian inference
<xref ref-type="bibr" rid="bib1.bibx30 bib1.bibx23" id="paren.21"/>.</p>
      <p>The <inline-formula><mml:math id="M121" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> penalty implies the prior <inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="bold">0</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">γ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mi mathvariant="bold">I</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and the <inline-formula><mml:math id="M123" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> LASSO penalty assigns a
Laplace (double exponential) prior <inline-formula><mml:math id="M124" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:msubsup><mml:mo>∏</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>d</mml:mi></mml:msubsup><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi mathvariant="italic">γ</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:msqrt><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mstyle><mml:mi>exp⁡</mml:mi><mml:mo mathvariant="italic">{</mml:mo><mml:mo>-</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi mathvariant="italic">γ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:msqrt><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:msqrt></mml:mfrac></mml:mstyle><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>. The LASSO penalty can also be specified
using the more computationally efficient hierarchical-scale mixture of
Gaussian distributions with exponential mixing distribution by assigning the
hierarchical prior<?xmltex \hack{\newpage}?><?xmltex \hack{\noindent}?>

                <disp-formula specific-use="align"><mml:math id="M125" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>∼</mml:mo><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="bold">0</mml:mn><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:msub><mml:mi mathvariant="bold">D</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="italic">γ</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi mathvariant="italic">γ</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>∼</mml:mo><mml:mtext>Exp</mml:mtext><mml:mfenced close=")" open="("><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mi mathvariant="italic">λ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M126" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">D</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="italic">γ</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtext>diag</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">γ</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">γ</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>
<xref ref-type="bibr" rid="bib1.bibx30" id="paren.22"/>. We hierarchically pool the error standard deviation
by assigning the hyperprior <inline-formula><mml:math id="M127" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mtext>log</mml:mtext><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> with vague hyperparameters <inline-formula><mml:math id="M128" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and
<inline-formula><mml:math id="M129" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mi>U</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where learning across years is achieved by
updating <inline-formula><mml:math id="M130" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M131" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. To perform a fully Bayesian
regularization that properly accounts for parameter uncertainty, we assign
the hyperpriors <inline-formula><mml:math id="M132" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">λ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>∼</mml:mo><mml:mtext>Gamma</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mi mathvariant="italic">λ</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="italic">λ</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> to allow for differential regularization through time. We
assign the hierarchical pooling prior by modeling the parameters
<inline-formula><mml:math id="M133" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mi mathvariant="italic">λ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M134" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="italic">λ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, reparameterizing the Gamma
distribution using its mean <inline-formula><mml:math id="M135" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">λ</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mi mathvariant="italic">λ</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="italic">λ</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></inline-formula> and variance <inline-formula><mml:math id="M136" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">λ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mi mathvariant="italic">λ</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="italic">λ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></inline-formula> and assigning the vague hyperpriors <inline-formula><mml:math id="M137" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">λ</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mtext>logN</mml:mtext><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M138" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">λ</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mtext>U</mml:mtext><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
</sec>
<sec id="Ch1.S3.SS2">
  <title>Probabilistic principal component regression</title>
      <p>PCR assumes the data <inline-formula><mml:math id="M139" display="inline"><mml:mi mathvariant="bold">X</mml:mi></mml:math></inline-formula>, and therefore the principal components
derived from <inline-formula><mml:math id="M140" display="inline"><mml:mi mathvariant="bold">X</mml:mi></mml:math></inline-formula> are observed without measurement error. The
current-era analog data are model interpolated, and, therefore, the principal
components have unaccounted for measurement error that violates the
assumptions of traditional PCR. Hence, the eigenvectors in <inline-formula><mml:math id="M141" display="inline"><mml:mi mathvariant="bold">U</mml:mi></mml:math></inline-formula> can
be thought of as estimates of the true eigenvectors under an appropriate
probabilistic model. As a remedy, probabilistic principal component models
assume the data matrix <inline-formula><mml:math id="M142" display="inline"><mml:mi mathvariant="bold">X</mml:mi></mml:math></inline-formula> is a noisy measurement of the true
process <xref ref-type="bibr" rid="bib1.bibx39" id="paren.23"/>. Letting <inline-formula><mml:math id="M143" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi><mml:mo>′</mml:mo></mml:msubsup></mml:mrow></mml:math></inline-formula> be the <inline-formula><mml:math id="M144" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>th
row of <inline-formula><mml:math id="M145" display="inline"><mml:mi mathvariant="bold">X</mml:mi></mml:math></inline-formula>, the model for the noisy observations is

                <disp-formula id="Ch1.E5" content-type="numbered"><mml:math id="M146" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mi mathvariant="bold-italic">m</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="bold">K</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">η</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M147" display="inline"><mml:mi mathvariant="bold-italic">m</mml:mi></mml:math></inline-formula> is the <inline-formula><mml:math id="M148" display="inline"><mml:mi>d</mml:mi></mml:math></inline-formula>-vector of means, <inline-formula><mml:math id="M149" display="inline"><mml:mi mathvariant="bold">K</mml:mi></mml:math></inline-formula> is a <inline-formula><mml:math id="M150" display="inline"><mml:mrow><mml:mi>d</mml:mi><mml:mo>×</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:math></inline-formula>
rotation matrix, <inline-formula><mml:math id="M151" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is a <inline-formula><mml:math id="M152" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula>-vector that represents the latent
eigenvectors of the process of interest, and <inline-formula><mml:math id="M153" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">η</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is zero
mean, independent Gaussian error with variance <inline-formula><mml:math id="M154" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>. Note that
<inline-formula><mml:math id="M155" display="inline"><mml:mi mathvariant="bold-italic">m</mml:mi></mml:math></inline-formula> can be integrated out of Eq. (<xref ref-type="disp-formula" rid="Ch1.E5"/>) without changing the
likelihood; thus, we assume that the data <inline-formula><mml:math id="M156" display="inline"><mml:mi mathvariant="bold">X</mml:mi></mml:math></inline-formula> have centered rows and
set <inline-formula><mml:math id="M157" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">m</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="bold">0</mml:mn></mml:mrow></mml:math></inline-formula> (i.e., anomalies). Because principal component
vectors are orthonormal, we complete the principal component model
specification by assigning independent priors <inline-formula><mml:math id="M158" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="bold">0</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="bold">I</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></inline-formula> for <inline-formula><mml:math id="M159" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:math></inline-formula>. A more general model is the factor
analysis model, where the error term <inline-formula><mml:math id="M160" display="inline"><mml:mi mathvariant="bold-italic">η</mml:mi></mml:math></inline-formula> has a generic
diagonal covariance matrix <inline-formula><mml:math id="M161" display="inline"><mml:mi mathvariant="bold">Σ</mml:mi></mml:math></inline-formula>
<xref ref-type="bibr" rid="bib1.bibx39" id="paren.24"/>. Thus, the probabilistic principal component
model can be viewed as a special case of factor analysis where the error term
<inline-formula><mml:math id="M162" display="inline"><mml:mi mathvariant="bold-italic">η</mml:mi></mml:math></inline-formula> is constrained to be diagonal with variance <inline-formula><mml:math id="M163" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>.</p>
      <p><xref ref-type="bibr" rid="bib1.bibx39" id="text.25"/> showed the maximum likelihood estimate (MLE)
of the rotation matrix <inline-formula><mml:math id="M164" display="inline"><mml:mi mathvariant="bold">K</mml:mi></mml:math></inline-formula> with <inline-formula><mml:math id="M165" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> components under the pPCR model
is

                <disp-formula id="Ch1.Ex7"><mml:math id="M166" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mover accent="true"><mml:mi mathvariant="bold">K</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold">U</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:msup><mml:mfenced close=")" open="("><mml:msub><mml:mi mathvariant="bold">Λ</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="italic">λ</mml:mi><mml:mo mathvariant="normal">¯</mml:mo></mml:mover><mml:msub><mml:mi mathvariant="bold">I</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mfenced><mml:mfrac><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:msup><mml:mi mathvariant="bold">R</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          <?xmltex \hack{\newpage}?><?xmltex \hack{\noindent}?>where <inline-formula><mml:math id="M167" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">U</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is a <inline-formula><mml:math id="M168" display="inline"><mml:mrow><mml:mi>d</mml:mi><mml:mo>×</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:math></inline-formula> matrix with the first <inline-formula><mml:math id="M169" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> columns
containing the leading eigenvectors, <inline-formula><mml:math id="M170" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Λ</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is a <inline-formula><mml:math id="M171" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>×</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:math></inline-formula>
diagonal matrix with the associated eigenvalues <inline-formula><mml:math id="M172" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">λ</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>≥</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>≥</mml:mo><mml:msub><mml:mi mathvariant="italic">λ</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> of <inline-formula><mml:math id="M173" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold">X</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mi mathvariant="bold">X</mml:mi></mml:mrow></mml:math></inline-formula> on the diagonal, the matrix
<inline-formula><mml:math id="M174" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">I</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the <inline-formula><mml:math id="M175" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>×</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:math></inline-formula> identity matrix, <inline-formula><mml:math id="M176" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="italic">λ</mml:mi><mml:mo mathvariant="normal">¯</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>d</mml:mi></mml:msubsup><mml:msub><mml:mi mathvariant="italic">λ</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>-</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></inline-formula> is the average variance
contribution for the truncated eigenvectors, and <inline-formula><mml:math id="M177" display="inline"><mml:mi mathvariant="bold">R</mml:mi></mml:math></inline-formula> is an arbitrary
orthogonal rotation matrix (which we set to be <inline-formula><mml:math id="M178" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">I</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>). We set
<inline-formula><mml:math id="M179" display="inline"><mml:mi mathvariant="bold">K</mml:mi></mml:math></inline-formula> at the MLE and rewrite Eq. (<xref ref-type="disp-formula" rid="Ch1.E5"/>) as

                <disp-formula id="Ch1.E6" content-type="numbered"><mml:math id="M180" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi><mml:mo>′</mml:mo></mml:msubsup></mml:mrow><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold">K</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">η</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p>After accounting for the measurement uncertainty in our predictor matrix
<inline-formula><mml:math id="M181" display="inline"><mml:mi mathvariant="bold">X</mml:mi></mml:math></inline-formula> by estimating the unknowns <inline-formula><mml:math id="M182" display="inline"><mml:mi mathvariant="bold">Z</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M183" display="inline"><mml:mi mathvariant="bold-italic">η</mml:mi></mml:math></inline-formula>
in Eq. (<xref ref-type="disp-formula" rid="Ch1.E6"/>), we link the historical observer data and
current-era analog data by regressing <inline-formula><mml:math id="M184" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> onto the latent
eigenvectors <inline-formula><mml:math id="M185" display="inline"><mml:mi mathvariant="bold">Z</mml:mi></mml:math></inline-formula>

                <disp-formula id="Ch1.E7" content-type="numbered"><mml:math id="M186" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi><mml:mo>′</mml:mo></mml:msubsup><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></disp-formula>

          and estimate the unknown regression coefficients <inline-formula><mml:math id="M187" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> in
Eq. (<xref ref-type="disp-formula" rid="Ch1.E7"/>).</p>
</sec>
<sec id="Ch1.S3.SS3">
  <title>Robust regression</title>
      <p>The historical observer data were collected using non-standard methods; thus,
there is likely more variability in the data than can be explained by
assuming a Gaussian error distribution. We propose extending
Eq. (<xref ref-type="disp-formula" rid="Ch1.E7"/>) to a model that is robust to outliers. The robust
pPCR data model using the Student's <inline-formula><mml:math id="M188" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> distribution

                <disp-formula id="Ch1.Ex8"><mml:math id="M189" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>∼</mml:mo><mml:mtext>t</mml:mtext><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi><mml:mo>′</mml:mo></mml:msubsup><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>

          is a model that better accommodates outliers in the data. The parameter
<inline-formula><mml:math id="M190" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the degrees of freedom of the Student's <inline-formula><mml:math id="M191" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> distribution. A common
choice of prior for the degrees of freedom <inline-formula><mml:math id="M192" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is to model the inverse
degrees of freedom with a <inline-formula><mml:math id="M193" display="inline"><mml:mrow><mml:mi>U</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn>0.5</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> distribution. We generalize this prior
to pool across years, assigning the inverse degrees of freedom the prior
<inline-formula><mml:math id="M194" display="inline"><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>∼</mml:mo><mml:mtext>Beta</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn>0.5</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> where the
four-parameter Beta<inline-formula><mml:math id="M195" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">α</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">β</mml:mi><mml:mo>,</mml:mo><mml:mi>L</mml:mi><mml:mo>,</mml:mo><mml:mi>U</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> prior is a Beta<inline-formula><mml:math id="M196" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">α</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">β</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>
prior scaled to the interval <inline-formula><mml:math id="M197" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mi>L</mml:mi><mml:mo>,</mml:mo><mml:mi>U</mml:mi><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>. To hierarchically pool the prior
model, we reparameterize <inline-formula><mml:math id="M198" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub><mml:msub><mml:mi mathvariant="italic">η</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and
<inline-formula><mml:math id="M199" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:msub><mml:mi mathvariant="italic">η</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> for <inline-formula><mml:math id="M200" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> and
<inline-formula><mml:math id="M201" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">η</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">∞</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. We complete the model statement by assigning
the hyperpriors <inline-formula><mml:math id="M202" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mtext>Beta</mml:mtext><mml:mo>(</mml:mo><mml:mn mathvariant="normal">5</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">5</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M203" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">η</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mtext>Gamma</mml:mtext><mml:mo>(</mml:mo><mml:mn>10</mml:mn><mml:mo>,</mml:mo><mml:mn>0.1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Although these priors appear informative, when
reparameterized, the prior specification is similar to the commonly used
vague <inline-formula><mml:math id="M204" display="inline"><mml:mrow><mml:mtext>Gamma</mml:mtext><mml:mo>(</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>,</mml:mo><mml:mn>0.1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> prior on the degrees of freedom <inline-formula><mml:math id="M205" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>
<xref ref-type="bibr" rid="bib1.bibx25" id="paren.26"/>. To regularize the robust data model, we modify the
LASSO prior for the regression coefficients using the variance of the
Student's <inline-formula><mml:math id="M206" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> distribution, resulting in the prior

                <disp-formula id="Ch1.Ex10"><mml:math id="M207" display="block"><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>∼</mml:mo><mml:mi>N</mml:mi><mml:mfenced close=")" open="("><mml:mn mathvariant="bold">0</mml:mn><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:mfrac></mml:mstyle><mml:msub><mml:mi mathvariant="bold">D</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="italic">γ</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where the parameters have the same priors as introduced previously.</p>
</sec>
<sec id="Ch1.S3.SS4">
  <title>Posterior distribution</title>
      <p>The latent principal components <inline-formula><mml:math id="M208" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are high dimensional
(approximately 20 000-dimensional for <inline-formula><mml:math id="M209" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:math></inline-formula>). We aim to avoid the
computational burden of sampling this parameter. Therefore, we wish to
integrate out the latent principal components

                <disp-formula id="Ch1.E8" content-type="numbered"><mml:math id="M210" display="block"><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo movablelimits="false">∫</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>]</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi>d</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          but this integral is not analytically tractable. We could attempt to
numerically integrate out <inline-formula><mml:math id="M211" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, but at great computational cost.
Instead, we write our Student's <inline-formula><mml:math id="M212" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> data model as a scale mixture where
<inline-formula><mml:math id="M213" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>∼</mml:mo><mml:mtext>N</mml:mtext><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi><mml:mo>′</mml:mo></mml:msubsup><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M214" display="inline"><mml:mrow><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>∼</mml:mo><mml:mtext>inv-</mml:mtext><mml:msup><mml:mi mathvariant="italic">χ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M215" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mo>=</mml:mo><mml:mo>∫</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>]</mml:mo><mml:mo>[</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>|</mml:mo><mml:msubsup><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi>d</mml:mi><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>.</mml:mo></mml:mrow></mml:math></inline-formula>
Then, we write the integral Eq. (<xref ref-type="disp-formula" rid="Ch1.E8"/>) as

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M216" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo movablelimits="false">∫</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><?xmltex \hack{\hbox\bgroup\fontsize{9.5}{9.5}\selectfont$\displaystyle}?><mml:mo>[</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>]</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi>d</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><?xmltex \hack{\hbox\bgroup\fontsize{9.5}{9.5}\selectfont$\displaystyle}?><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mfenced close=")" open="("><mml:mo movablelimits="false">∫</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>]</mml:mo><mml:mo>[</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>|</mml:mo><mml:msubsup><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi>d</mml:mi><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mfenced><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>]</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi>d</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><?xmltex \hack{\hbox\bgroup\fontsize{9.5}{9.5}\selectfont$\displaystyle}?><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mfenced open="(" close=")"><mml:mo movablelimits="false">∫</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>]</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>]</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi>d</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mfenced><mml:mo>[</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>|</mml:mo><mml:msubsup><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi>d</mml:mi><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E9"><mml:mtd/><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><?xmltex \hack{\hbox\bgroup\fontsize{9.5}{9.5}\selectfont$\displaystyle}?><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>]</mml:mo><mml:mo>[</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>|</mml:mo><mml:msubsup><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi>d</mml:mi><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>,</mml:mo><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            where the integral in Eq. (<xref ref-type="disp-formula" rid="Ch1.E9"/>) is evaluated by Markov chain
Monte Carlo (MCMC), first sampling <inline-formula><mml:math id="M217" display="inline"><mml:mrow><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>∼</mml:mo><mml:mtext>inv-</mml:mtext><mml:msup><mml:mi mathvariant="italic">χ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, then evaluating the density <inline-formula><mml:math id="M218" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>. These modifications result in the integrated data model
(see the Supplement for details)

                <disp-formula id="Ch1.E10" content-type="numbered"><mml:math id="M219" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>∼</mml:mo><mml:mtext>N</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M220" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi><mml:mo>′</mml:mo></mml:msubsup><mml:mover accent="true"><mml:mi mathvariant="bold">K</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:msubsup><mml:mi mathvariant="bold">M</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msubsup><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M221" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi><mml:mo>′</mml:mo></mml:msubsup><mml:msubsup><mml:mi mathvariant="bold">M</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msubsup><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> with
<inline-formula><mml:math id="M222" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Λ</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="bold">I</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, a
diagonal matrix that can be inverted efficiently. Integration results in
significant computational savings because we avoid sampling <inline-formula><mml:math id="M223" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> vectors of
length <inline-formula><mml:math id="M224" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> (approximately 20 000 each). The cost of not sampling the latent
principal components <inline-formula><mml:math id="M225" display="inline"><mml:mi mathvariant="bold">Z</mml:mi></mml:math></inline-formula> is loss of conjugacy for the regression
coefficients <inline-formula><mml:math id="M226" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> in the MCMC algorithm. The posterior
distribution (for the robust pPCA model with LASSO regularization) from which
we sample using MCMC is

                <disp-formula specific-use="align"><mml:math id="M227" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><?xmltex \hack{\hbox\bgroup\fontsize{8.5}{8.5}\selectfont$\displaystyle}?><mml:munderover><mml:mo movablelimits="false">∏</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:munderover><mml:munder><mml:mo movablelimits="false">∏</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>∈</mml:mo><mml:msub><mml:mi mathvariant="script">H</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mfenced open="[" close="]"><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">η</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">γ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">λ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">λ</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">λ</mml:mi></mml:msub><mml:mo mathsize="1.1em">|</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="bold">X</mml:mi></mml:mfenced><mml:mo>∝</mml:mo><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mspace width="1em" linebreak="nobreak"/><?xmltex \hack{\hbox\bgroup\fontsize{8.5}{8.5}\selectfont$\displaystyle}?><mml:munderover><mml:mo movablelimits="false">∏</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:munderover><mml:mfenced open="(" close=")"><mml:munder><mml:mo movablelimits="false">∏</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>∈</mml:mo><mml:msub><mml:mi mathvariant="script">H</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mfenced close="]" open="["><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo mathsize="1.1em">|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mfenced><mml:mfenced open="[" close="]"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo mathsize="1.1em">|</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mfenced></mml:mfenced><mml:mfenced open="[" close="]"><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo mathsize="1.1em">|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">γ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mfenced><mml:mfenced open="[" close="]"><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo mathsize="1.1em">|</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub></mml:mfenced><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mspace width="1em" linebreak="nobreak"/><mml:mspace width="1em" linebreak="nobreak"/><?xmltex \hack{\hbox\bgroup\fontsize{8.5}{8.5}\selectfont$\displaystyle}?><mml:mo>×</mml:mo><mml:mfenced open="[" close="]"><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub></mml:mfenced><mml:mfenced open="[" close="]"><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub></mml:mfenced><mml:mfenced close="]" open="["><mml:mi mathvariant="italic">σ</mml:mi></mml:mfenced><mml:mfenced open="[" close="]"><mml:msub><mml:mi mathvariant="bold-italic">γ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo mathsize="1.1em">|</mml:mo><mml:msubsup><mml:mi mathvariant="italic">λ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mfenced><mml:mfenced open="[" close="]"><mml:msubsup><mml:mi mathvariant="italic">λ</mml:mi><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo mathsize="1.1em">|</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">λ</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">η</mml:mi><mml:mi mathvariant="italic">λ</mml:mi></mml:msub></mml:mfenced><mml:mfenced open="[" close="]"><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">λ</mml:mi></mml:msub></mml:mfenced><mml:mfenced close="]" open="["><mml:msub><mml:mi mathvariant="italic">η</mml:mi><mml:mi mathvariant="italic">λ</mml:mi></mml:msub></mml:mfenced><mml:mo>,</mml:mo><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M228" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">H</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the set of locations where there are observations
for year <inline-formula><mml:math id="M229" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>. We fit our models using <monospace>JAGS</monospace> <xref ref-type="bibr" rid="bib1.bibx31" id="paren.27"/>
within the <monospace>R</monospace> computing environment <xref ref-type="bibr" rid="bib1.bibx34" id="paren.28"/>.</p>
      <p>For each of the eight candidate models, we fit four parallel chains with
random initial conditions, running 20 000 iterations per chain and
discarding the first 10 000 iterations as burn-in. Fitting all eight models
and the associated post-processing took approximately 18 h on a 2014
dual-core 2.6 GHz MacBook Pro with 8 GB RAM. We thinned our chains every 10
iterations to reduce post-processing time, resulting in a total of 4000
samples and evaluated model convergence using the <inline-formula><mml:math id="M230" display="inline"><mml:mover accent="true"><mml:mi>R</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> statistic
<xref ref-type="bibr" rid="bib1.bibx14" id="paren.29"/>. We chose vague hyperpriors throughout; the
ability to estimate temperature surfaces from these priors implies the
results are not highly sensitive to the prior values. A choice of stronger
hyperprior values could improve inference, but very strong hyperpriors could
also dominate the influence of the data in the posterior estimates.
Preliminary analyses not shown in this paper indicated little sensitivity to
reasonable prior choices.</p>
</sec>
</sec>
<sec id="Ch1.S4">
  <title>Scoring rules</title>
      <p>To evaluate model performance, we apply scoring rules to the estimated
posterior predictive distributions. A highly desirable property of a scoring
rule is propriety <xref ref-type="bibr" rid="bib1.bibx16" id="paren.30"/>. A scoring rule is proper if the
expected score of the optimal prediction is less than or equal to the
expected score of any other prediction <xref ref-type="bibr" rid="bib1.bibx4" id="paren.31"/>. Hence, a
proper scoring rule, on average, chooses the best prediction from a set of
candidate predictions <xref ref-type="bibr" rid="bib1.bibx18" id="paren.32"/>. Often, paleoclimate
reconstructions evaluate predictive performance by holding out some of the
training set data for use in cross-validation, using skill scores like the
coefficient of efficiency (CE) and relative efficiency (RE)
<xref ref-type="bibr" rid="bib1.bibx9 bib1.bibx35 bib1.bibx37 bib1.bibx38" id="paren.33"/>.
Although these scoring rules are common in the paleoclimate reconstruction
community, <xref ref-type="bibr" rid="bib1.bibx17" id="text.34"/> suggest that scoring rules like CE
and RE are improper in general. Because CE and RE are improper, it is
possible that the optimal prediction can, on average, have a worse score than
a sub-optimal prediction, leading to incorrect inference. Therefore, we focus
on three proper scoring rules: mean square prediction error (MSPE), the
continuous ranked probability score (CRPS), and a computationally efficient
approximation to leave-one-out cross-validation (LOO) using the log score. In
general, MSPE is not proper, but because our data models are Gaussian and
Student's <inline-formula><mml:math id="M231" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, MSPE is proper for predictions of the posterior mean in this
case.</p>
      <p>The use of MSPE as a scoring rule implies an <inline-formula><mml:math id="M232" display="inline"><mml:mrow><mml:msup><mml:mi>L</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> loss function on the
posterior distribution; therefore, our predictions are the posterior
predictive means

              <disp-formula id="Ch1.Ex17"><mml:math id="M233" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi mathvariant="normal">E</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub><mml:mo>[</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi>d</mml:mi><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

        where <inline-formula><mml:math id="M234" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mo>=</mml:mo><mml:mo>∫</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi>d</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the posterior predictive distribution for model
parameters <inline-formula><mml:math id="M235" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Given out-of-sample observations
<inline-formula><mml:math id="M236" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mtext>oos</mml:mtext><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, MSPE is

              <disp-formula id="Ch1.Ex19"><mml:math id="M237" display="block"><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>T</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:munderover><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>∉</mml:mo><mml:msub><mml:mi mathvariant="script">H</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:munder><mml:msup><mml:mfenced open="(" close=")"><mml:mi mathvariant="normal">E</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mtext>oos</mml:mtext><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

        where <inline-formula><mml:math id="M238" display="inline"><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the number of out-of-sample locations for year <inline-formula><mml:math id="M239" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> and
<inline-formula><mml:math id="M240" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">H</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the set of observed locations in the historical observer
data. Because MSPE uses the posterior predictive mean (a point prediction)
instead of the full posterior distribution, MSPE ignores much of the
information in the posterior distribution gained by performing Bayesian
inference. Therefore, MSPE is not an ideal scoring rule for a probabilistic
prediction, such as a posterior predictive distribution, even when MSPE is
proper. For example, consider two models that give rise to posterior
predictive distributions with the same posterior predictive mean but
different posterior predictive variances. In this case, it is obvious that
the predictive distribution that has better predictive coverage should be
preferred, but MSPE would score the two models identically, demonstrating how
MSPE loses information by collapsing the posterior distribution into a point
estimate.</p>
      <p>An alternative to MSPE is the CRPS scoring rule. CRPS is proper, utilizes the
full posterior predictive distribution, and allows for a direct comparison of
point predictions and probabilistic predictions <xref ref-type="bibr" rid="bib1.bibx17" id="paren.35"/>.
CRPS resolves the issue presented in the previously described scenario by
including the width of the predictive distribution in the evaluation of the
score. Several recent papers presenting climate reconstructions have made use
of the CRPS for these reasons
<xref ref-type="bibr" rid="bib1.bibx2 bib1.bibx45 bib1.bibx40" id="paren.36"/>.
Given a prediction with the cumulative distribution function, <inline-formula><mml:math id="M241" display="inline"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, at
location <inline-formula><mml:math id="M242" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> and time <inline-formula><mml:math id="M243" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, and out-of-sample observations
<inline-formula><mml:math id="M244" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mtext>oos</mml:mtext><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, the CRPS is defined as

              <disp-formula specific-use="align" content-type="numbered"><mml:math id="M245" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtext>CRPS</mml:mtext><mml:mo>(</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mtext>oos</mml:mtext><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E11"><mml:mtd/><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>-</mml:mo><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:munderover><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>∉</mml:mo><mml:msub><mml:mi mathvariant="script">H</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:munder><mml:munderover><mml:mo movablelimits="false">∫</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow><mml:mi mathvariant="normal">∞</mml:mi></mml:munderover><mml:msup><mml:mfenced close=")" open="("><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mfenced close="}" open="{"><mml:mi>y</mml:mi><mml:mo>≥</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mtext>oos</mml:mtext><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mfenced></mml:mrow></mml:msub></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi>y</mml:mi><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

          <xref ref-type="bibr" rid="bib1.bibx17" id="text.37"/> show that Eq. (<xref ref-type="disp-formula" rid="Ch1.E11"/>) can be written
alternatively as

              <disp-formula specific-use="align" content-type="numbered"><mml:math id="M246" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E12"><mml:mtd/><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtext>CRPS</mml:mtext></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>(</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mtext>oos</mml:mtext></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mtr><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><?xmltex \hack{\hbox\bgroup\fontsize{8.5}{8.5}\selectfont$\displaystyle}?><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:munderover><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>∉</mml:mo><mml:msub><mml:mi mathvariant="script">H</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mfenced open="(" close=")"><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mfenced close="|" open="|"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mtext>oos</mml:mtext><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mfenced><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mfenced open="|" close="|"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mo>*</mml:mo></mml:msubsup></mml:mfenced></mml:mfenced><mml:mo>,</mml:mo><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>

          where <inline-formula><mml:math id="M247" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M248" display="inline"><mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mo>*</mml:mo></mml:msubsup></mml:mrow></mml:math></inline-formula> are independent copies of a random variable
with distribution function <inline-formula><mml:math id="M249" display="inline"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> and the expectation <inline-formula><mml:math id="M250" display="inline"><mml:mi>E</mml:mi></mml:math></inline-formula> is with respect
to the probability density induced by <inline-formula><mml:math id="M251" display="inline"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>. The first expectation in
Eq. (<xref ref-type="disp-formula" rid="Ch1.E12"/>) measures calibration (the absolute error of the
prediction relative to the out-of-sample value) and the second expectation
rewards predictions that are precise (i.e., narrow prediction intervals).</p>
      <p>We can estimate the CRPS after obtaining posterior samples
<inline-formula><mml:math id="M252" display="inline"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>t</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> from the posterior predictive distribution <inline-formula><mml:math id="M253" display="inline"><mml:mrow><mml:mfenced close="]" open="["><mml:msubsup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>t</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo mathsize="1.1em">|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mfenced></mml:mrow></mml:math></inline-formula> at each post burn-in
iteration <inline-formula><mml:math id="M254" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>. Then, Eq. (<xref ref-type="disp-formula" rid="Ch1.E12"/>) is approximated by

              <disp-formula specific-use="align" content-type="numbered"><mml:math id="M255" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mover accent="true"><mml:mtext>CRPS</mml:mtext><mml:mo mathvariant="normal" stretchy="true">^</mml:mo></mml:mover></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>(</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>F</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mtext>oos</mml:mtext></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:munderover><mml:mfenced open="(" close=""><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>∉</mml:mo><mml:msub><mml:mi mathvariant="script">H</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mfenced open="(" close=""><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>K</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>K</mml:mi></mml:munderover><mml:mfenced close="|" open="|"><mml:msubsup><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mtext>oos</mml:mtext><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mfenced></mml:mfenced></mml:mfenced><mml:mo>-</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E13"><mml:mtd/><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mfenced close=")" open="."><mml:mfenced close=")" open="."><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:msup><mml:mi>K</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>K</mml:mi></mml:munderover><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>K</mml:mi></mml:munderover><mml:mfenced close="|" open="|"><mml:msubsup><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mfenced></mml:mfenced></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

          A major disadvantage of both MSPE and CRPS is the need for out-of-sample
validation data. For our simulation study, MSPE and CRPS are
straightforward to calculate because we simulated the out-of-sample
validation data; in practical paleoclimate reconstructions, there are no
out-of-sample data. Therefore, MSPE and CRPS must be approximated using
cross-validation methods, although these methods are computationally
costly and time consuming to implement.</p>
      <p>An alternative is to use the approximate leave-one-out cross-validation
method (LOO; <xref ref-type="bibr" rid="bib1.bibx43" id="altparen.38"/>). LOO uses a proper scoring rule,
the log score, to evaluate predictive skill
<xref ref-type="bibr" rid="bib1.bibx12 bib1.bibx17 bib1.bibx23" id="paren.39"/>. We
estimate the leave-one-out log pointwise predictive density

              <disp-formula specific-use="align" content-type="numbered"><mml:math id="M256" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mtext>lpd</mml:mtext><mml:mtext>loo</mml:mtext></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:munderover><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>∈</mml:mo><mml:msub><mml:mi mathvariant="script">H</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mi>log⁡</mml:mi><mml:mo>[</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo><mml:mo>=</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E14"><mml:mtd/><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:munderover><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>∈</mml:mo><mml:msub><mml:mi mathvariant="script">H</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mi>log⁡</mml:mi><mml:mo movablelimits="false">∫</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi>d</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

          where <inline-formula><mml:math id="M257" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> are the data <inline-formula><mml:math id="M258" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> at time <inline-formula><mml:math id="M259" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> without the <inline-formula><mml:math id="M260" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>th
location. One can calculate Eq. (<xref ref-type="disp-formula" rid="Ch1.E14"/>) directly by
cross-validation at a high computational cost, or one can approximate
Eq. (<xref ref-type="disp-formula" rid="Ch1.E14"/>) using importance sampling from post burn-in posterior
samples using the full data as described in <xref ref-type="bibr" rid="bib1.bibx43" id="text.40"/>.
Importance ratios with high variance can cause the estimate in
Eq. (<xref ref-type="disp-formula" rid="Ch1.E14"/>) to be highly unstable and unreliable, and are therefore
of practical concern. To test for the presence of large variance of the
importance ratios, <xref ref-type="bibr" rid="bib1.bibx27" id="text.41"/> proposed fitting the
generalized Pareto distribution to the upper tail of importance ratios and
examining the empirical estimates of the tail shape parameter <inline-formula><mml:math id="M261" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>. If the
estimated tail parameter <inline-formula><mml:math id="M262" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="italic">ξ</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is less than <inline-formula><mml:math id="M263" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>, the variance of
the importance ratios is finite and the importance ratios approximating the
log posterior score holding out <inline-formula><mml:math id="M264" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> can be used directly to approximate
LOO. If the estimated tail parameter is <inline-formula><mml:math id="M265" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>&lt;</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="italic">ξ</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>, the
variance of the importance ratios is infinite but the mean of the importance
ratios exists. Hence, <xref ref-type="bibr" rid="bib1.bibx41" id="text.42"/> propose using smoothed
importance ratios. If the estimated tail parameter <inline-formula><mml:math id="M266" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="italic">ξ</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>, this
suggests that the mean and variance of the importance ratios do not exist but
that the variance of the smoothed importance ratios is finite, but large, and
the use of LOO is sensitive to the held-out observation. Using the smoothed
importance weights <inline-formula><mml:math id="M267" display="inline"><mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula>, we obtain the Pareto-smoothed importance
sampling approximation

              <disp-formula id="Ch1.E15" content-type="numbered"><mml:math id="M268" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mover accent="true"><mml:mtext>elpd</mml:mtext><mml:mo mathvariant="normal" stretchy="true">^</mml:mo></mml:mover><mml:mtext>PSIS</mml:mtext></mml:msub></mml:mrow><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:munderover><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>∈</mml:mo><mml:msub><mml:mi mathvariant="script">H</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mi>log⁡</mml:mi><mml:mfenced close=")" open="("><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>K</mml:mi></mml:msubsup><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>[</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>K</mml:mi></mml:msubsup><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p>We use the deviance scale and set <inline-formula><mml:math id="M269" display="inline"><mml:mrow><mml:mover accent="true"><mml:mtext>LOO</mml:mtext><mml:mo mathvariant="normal" stretchy="true">^</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:msub><mml:mover accent="true"><mml:mtext>elpd</mml:mtext><mml:mo mathvariant="normal" stretchy="true">^</mml:mo></mml:mover><mml:mtext>PSIS</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> to make LOO a negatively oriented score
(the best model is the one with the lowest score), implementing our score
using <monospace>R</monospace> package loo <xref ref-type="bibr" rid="bib1.bibx42" id="paren.43"/>.</p>
</sec>
<sec id="Ch1.S5">
  <title>Simulation</title>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><caption><p>Simulation study showing observed noisy historical observer data
<bold>(a)</bold>, the simulated true latent climate process we aim to predict
<bold>(b)</bold>, the first four noisy principal components (PCs) estimated from
the historical observer data <bold>(c)</bold>, and the first four simulated true
latent PCs <bold>(d)</bold> that show the effect of measurement error when
compared to <bold>(c)</bold>. Each year of the historical observer period in the
simulation study in <bold>(a)</bold> and <bold>(b)</bold> is assigned a different
color and the historical observer period observations <bold>(a)</bold> are
clustered in space, changing in sample size through time, and noisier than
the latent temperature in <bold>(b)</bold>. Both the noisy <bold>(c)</bold> and
latent <bold>(d)</bold> PCs increase in variability as the number of the
component increases, but the latent PCs are smoother. The first PC is in
black, the second PC is in blue, the third PC is in green, and the fourth PC
is in red.</p></caption>
        <?xmltex \igopts{width=384.112205pt}?><graphic xlink:href="https://ascmo.copernicus.org/articles/3/1/2017/ascmo-3-1-2017-f02.pdf"/>

      </fig>

      <p>With paleoclimate data, it is difficult to verify the predictive ability of
models using cross-validation. With only a handful of observations in the
historical observer data available for each year, cross-validation techniques
could be highly biased due to the effects of unusual observations in small
sample sizes. This is important because we expect noisy and potentially
outlying observations in the historical observer data due to the data
collection procedures. Additionally, the high dimensionality of the field we
aim to reconstruct and the use of computationally intensive MCMC estimation
make cross-validation costly. Instead, we conducted a simulation study to
explore the different models for the historical observer station data and
evaluate model performance using the scoring rules above. Although we do not
simulate from the model that is used for estimation, the simulated data
represent a reasonable approximation to mid-day July temperature, providing
an environment for model testing and exploration of empirical performance.</p>
      <p>We simulate mid-day July temperature in one spatial dimension (we extend to
two dimensions using the real data), allowing for faster computation and
easier graphical exploration of the spatio-temporal process. We simulate <inline-formula><mml:math id="M270" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mo>=</mml:mo><mml:mn>50</mml:mn></mml:mrow></mml:math></inline-formula> realizations of a latent surface from the model

              <disp-formula id="Ch1.Ex25"><mml:math id="M271" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">β</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">η</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

        where the matrix <inline-formula><mml:math id="M272" display="inline"><mml:mi mathvariant="bold">W</mml:mi></mml:math></inline-formula> represents fixed influences on climate, such as
latitude, elevation, and other covariates that explain much of the
temperature surface as well as time-varying components that represent slowly
varying global-scale climate processes. To construct patterns that might be
seen in climate observations, we simulate temporally varying regression
coefficients at different periodicities to represent global-scale climate
processes like the Pacific Decadal Oscillation or the Atlantic Multidecadal
Oscillation. We do not claim our simulation behaves like any climatological
process, only that this example facilitates exploration of complicated
patterns potentially seen in climatological data.</p>
      <p>We include a spatially correlated random effect <inline-formula><mml:math id="M273" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">η</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mtext>N</mml:mtext><mml:mfenced close=")" open="("><mml:mn mathvariant="bold">0</mml:mn><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="italic">η</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mi mathvariant="bold">R</mml:mi><mml:mfenced close=")" open="("><mml:mi mathvariant="italic">ϕ</mml:mi></mml:mfenced></mml:mfenced></mml:mrow></mml:math></inline-formula> that smooths the patterns, generating
realizations of a one-dimensional climate field (Fig. <xref ref-type="fig" rid="Ch1.F2"/>b). A
common choice for the form of <inline-formula><mml:math id="M274" display="inline"><mml:mrow><mml:mi mathvariant="bold">R</mml:mi><mml:mfenced open="(" close=")"><mml:mi mathvariant="italic">ϕ</mml:mi></mml:mfenced></mml:mrow></mml:math></inline-formula> is the Matérn
class of correlation functions. For our simulation, we use the exponential
correlation function, a member of the Matérn family. In the exponential
correlation function, the <inline-formula><mml:math id="M275" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:math></inline-formula>th element <inline-formula><mml:math id="M276" display="inline"><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>exp⁡</mml:mi><mml:mfenced open="(" close=")"><mml:mo>-</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi></mml:mfenced></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M277" display="inline"><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> represents the Euclidean distance
between the <inline-formula><mml:math id="M278" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>th and <inline-formula><mml:math id="M279" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th spatial locations and <inline-formula><mml:math id="M280" display="inline"><mml:mi mathvariant="italic">ϕ</mml:mi></mml:math></inline-formula> is the spatial range
parameter.</p>
      <p>To create observations that match the temporal irregularities and spatial
clustering behavior in the historical observer data, we sample the
one-dimensional spatial field using weighted probabilities that generate
clustered observations in space, storing the simulated temperature
observations at the <inline-formula><mml:math id="M281" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> locations in the vector <inline-formula><mml:math id="M282" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Using this
sampling design, we generate noisy realizations for the simulated historical
observer data for simulated years <inline-formula><mml:math id="M283" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mn>25</mml:mn></mml:mrow></mml:math></inline-formula>
(Fig. <xref ref-type="fig" rid="Ch1.F2"/>a) using

              <disp-formula id="Ch1.E16" content-type="numbered"><mml:math id="M284" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

        where <inline-formula><mml:math id="M285" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is independent Student's <inline-formula><mml:math id="M286" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> error with
variance <inline-formula><mml:math id="M287" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="italic">σ</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn>1.5</mml:mn></mml:mrow></mml:math></inline-formula> and degrees of freedom <inline-formula><mml:math id="M288" display="inline"><mml:mrow><mml:mi mathvariant="italic">ν</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn></mml:mrow></mml:math></inline-formula> that
represent uncertainty in historical temperature measurements. To generate the
set of simulated current-era analog patterns <inline-formula><mml:math id="M289" display="inline"><mml:mi mathvariant="bold">X</mml:mi></mml:math></inline-formula> that will be used
in our regression model, we define for the simulated current-era analog years
<inline-formula><mml:math id="M290" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>26</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mn>50</mml:mn></mml:mrow></mml:math></inline-formula>

              <disp-formula id="Ch1.E17" content-type="numbered"><mml:math id="M291" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mo mathvariant="normal">¨</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

        where, by adding uncorrelated, independent Gaussian noise
<inline-formula><mml:math id="M292" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mo mathvariant="normal">¨</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> with variance <inline-formula><mml:math id="M293" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="italic">σ</mml:mi><mml:mo mathvariant="normal">¨</mml:mo></mml:mover><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>.75, we
account for the measurement error of the temperature process during the
current-era analog period. The measurements used in the model for the
current-era analog period are from PRISM model interpolated data and have
measurement error, where this measurement error should be less than that of
the historical observer data; hence, <inline-formula><mml:math id="M294" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="italic">σ</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>&gt;</mml:mo><mml:msup><mml:mover accent="true"><mml:mi mathvariant="italic">σ</mml:mi><mml:mo mathvariant="normal">¨</mml:mo></mml:mover><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>. We
then combine the simulated values into the noisy pattern matrix <inline-formula><mml:math id="M295" display="inline"><mml:mrow><mml:mi mathvariant="bold">X</mml:mi><mml:mo>≡</mml:mo><mml:mfenced open="(" close=")"><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn>26</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn>50</mml:mn></mml:msub></mml:mfenced></mml:mrow></mml:math></inline-formula> that is used in our
model framework. The latent temperature patterns <inline-formula><mml:math id="M296" display="inline"><mml:mrow><mml:mi mathvariant="bold">S</mml:mi><mml:mo>≡</mml:mo><mml:mfenced open="(" close=")"><mml:msub><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mn>25</mml:mn></mml:msub></mml:mfenced></mml:mrow></mml:math></inline-formula> are the unobserved target
for our reconstruction. Note that, in the real data, <inline-formula><mml:math id="M297" display="inline"><mml:mi mathvariant="bold">S</mml:mi></mml:math></inline-formula> are
unavailable; therefore, our scoring rules can use only the noisy observations
<inline-formula><mml:math id="M298" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, limiting the ability to improve reconstruction skill.
Figure <xref ref-type="fig" rid="Ch1.F2"/>a shows the simulated historical observer period
noisy temperature realizations <inline-formula><mml:math id="M299" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and Fig. <xref ref-type="fig" rid="Ch1.F2"/>b shows
the simulated historical observer period true, latent temperature field that
is the target of our reconstruction, with the <inline-formula><mml:math id="M300" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis representing spatial
location. The noisy principal components derived from the simulated
current-era analog data are plotted in Fig. <xref ref-type="fig" rid="Ch1.F2"/>c with the
simulated latent principal components in Fig. <xref ref-type="fig" rid="Ch1.F2"/>d. Comparing
Fig. <xref ref-type="fig" rid="Ch1.F2"/>c and d shows that using the noisy principal components
is analogous to the errors-in-covariates framework, where noisy observations
of the current-era analog covariates (in our case the principal components)
can lead to bias in the regression coefficients, inflated residual variance,
and a reduction in prediction skill
<xref ref-type="bibr" rid="bib1.bibx7 bib1.bibx11 bib1.bibx6" id="paren.44"/>.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1" specific-use="star"><caption><p>Simulation experiment scores. Smaller values indicate better model performance.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="9">
     <oasis:colspec colnum="1" colname="col1" align="right"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="left"/>
     <oasis:colspec colnum="8" colname="col8" align="right"/>
     <oasis:colspec colnum="9" colname="col9" align="right"/>
     <oasis:thead>
       <oasis:row>  
         <oasis:entry colname="col1"/>  
         <oasis:entry rowsep="1" namest="col2" nameend="col3" align="center">MSPE </oasis:entry>  
         <oasis:entry colname="col4"/>  
         <oasis:entry rowsep="1" namest="col5" nameend="col6" align="center">CRPS </oasis:entry>  
         <oasis:entry colname="col7"/>  
         <oasis:entry rowsep="1" namest="col8" nameend="col9" align="center">LOO </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">Model</oasis:entry>  
         <oasis:entry colname="col2">Gaussian</oasis:entry>  
         <oasis:entry colname="col3">Robust</oasis:entry>  
         <oasis:entry colname="col4"/>  
         <oasis:entry colname="col5">Gaussian</oasis:entry>  
         <oasis:entry colname="col6">Robust</oasis:entry>  
         <oasis:entry colname="col7"/>  
         <oasis:entry colname="col8">Gaussian</oasis:entry>  
         <oasis:entry colname="col9">Robust</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>  
         <oasis:entry colname="col1">SSVS PCR</oasis:entry>  
         <oasis:entry colname="col2">0.379</oasis:entry>  
         <oasis:entry colname="col3">0.380</oasis:entry>  
         <oasis:entry colname="col4"/>  
         <oasis:entry colname="col5">199</oasis:entry>  
         <oasis:entry colname="col6">194</oasis:entry>  
         <oasis:entry colname="col7"/>  
         <oasis:entry colname="col8">2917</oasis:entry>  
         <oasis:entry colname="col9">2901</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">SSVS pPCR</oasis:entry>  
         <oasis:entry colname="col2">0.403</oasis:entry>  
         <oasis:entry colname="col3">0.404</oasis:entry>  
         <oasis:entry colname="col4"/>  
         <oasis:entry colname="col5">206</oasis:entry>  
         <oasis:entry colname="col6">205</oasis:entry>  
         <oasis:entry colname="col7"/>  
         <oasis:entry colname="col8">2920</oasis:entry>  
         <oasis:entry colname="col9">2919</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">LASSO PCR</oasis:entry>  
         <oasis:entry colname="col2">0.508</oasis:entry>  
         <oasis:entry colname="col3">0.456</oasis:entry>  
         <oasis:entry colname="col4"/>  
         <oasis:entry colname="col5">217</oasis:entry>  
         <oasis:entry colname="col6">205</oasis:entry>  
         <oasis:entry colname="col7"/>  
         <oasis:entry colname="col8">2970</oasis:entry>  
         <oasis:entry colname="col9">2934</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">LASSO pPCR</oasis:entry>  
         <oasis:entry colname="col2">0.398</oasis:entry>  
         <oasis:entry colname="col3">0.397</oasis:entry>  
         <oasis:entry colname="col4"/>  
         <oasis:entry colname="col5">206</oasis:entry>  
         <oasis:entry colname="col6">205</oasis:entry>  
         <oasis:entry colname="col7"/>  
         <oasis:entry colname="col8">2918</oasis:entry>  
         <oasis:entry colname="col9">2918</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p>We compare the performance of each model specification using MSPE, CRPS, and
LOO scoring rules in our simulation where the best model is the one with the
smallest score. We fit the PCR and pPCR models using SSVS and LASSO
regularization with both the Gaussian and robust Student's <inline-formula><mml:math id="M301" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> data models,
comparing eight models with the results displayed in Table
<xref ref-type="table" rid="Ch1.T1"/>. Across scores, the models perform similarly, with no
consistently best model, although the robust models generally have lower
scores than the Gaussian models. The LOO Pareto tail parameter estimates for
the robust models show less evidence of misspecification than the traditional
Gaussian data models (figure not shown) and the pPCR models have consistently
smaller tail parameter estimates than the PCR models, suggesting that the
least misspecified models are robust pPCR.</p>
      <p>Predicted versus simulated temperatures for the robust PCR and robust pPCR
models using LASSO regularization are shown in Fig. <xref ref-type="fig" rid="Ch1.F3"/> with
years represented using different colors. Because the predictions for each
year cluster around the 45<inline-formula><mml:math id="M302" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> line, this shows the models reconstruct
the annual differences accurately, which are the main features of the
simulated data and of primary interest in understanding climate change.
However, the models fail to reconstruct much of the within-year spatial
variability (the humps and valleys within a year in Fig. <xref ref-type="fig" rid="Ch1.F2"/>b),
which is unsurprising given the small sample sizes. Despite having similar
predictive scores, there are visible differences in predictions between the
two models. The robust PCR predictions predict some of the spatial structure
of the mean (the point clouds are generally centered on the 45<inline-formula><mml:math id="M303" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>
line), but tend to have unreasonably small predictive standard deviations
(not shown). The robust pPCR predictions estimate a spatially averaged annual
mean in years with small sample sizes, but predict little spatial structure.
Instead, robust pPCR produces predictive standard deviation estimates that
account for uncertainty in years with very little data.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3" specific-use="star"><caption><p>Simulation truth plotted against predicted temperature for the
robust PCR LASSO model on the left and the robust pPCR LASSO model on the
right. Predictions for each simulated year are given different colors and the
clustering of colors represents annual-scale changes in the mean temperature
surface. Results are shown for the LASSO model and the SSVS model performs
similarly.</p></caption>
        <?xmltex \igopts{width=384.112205pt}?><graphic xlink:href="https://ascmo.copernicus.org/articles/3/1/2017/ascmo-3-1-2017-f03.png"/>

      </fig>

</sec>
<sec id="Ch1.S6">
  <title>Observer station data reconstruction</title>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T2"><caption><p>Historical observer reconstruction scores. Smaller values indicate better model performance. </p></caption><oasis:table frame="topbot"><?xmltex \begin{scaleboxenv}{.97}[.97]?><oasis:tgroup cols="6">
     <oasis:colspec colnum="1" colname="col1" align="right"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:thead>
       <oasis:row>  
         <oasis:entry colname="col1"/>  
         <oasis:entry rowsep="1" namest="col2" nameend="col3" align="center">Full data </oasis:entry>  
         <oasis:entry colname="col4"/>  
         <oasis:entry rowsep="1" namest="col5" nameend="col6" align="center">Outlier removed </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">Model</oasis:entry>  
         <oasis:entry colname="col2">Gaussian</oasis:entry>  
         <oasis:entry colname="col3">Robust</oasis:entry>  
         <oasis:entry colname="col4"/>  
         <oasis:entry colname="col5">Gaussian</oasis:entry>  
         <oasis:entry colname="col6">Robust</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>  
         <oasis:entry colname="col1">SSVS PCR</oasis:entry>  
         <oasis:entry colname="col2">7499</oasis:entry>  
         <oasis:entry colname="col3">7183</oasis:entry>  
         <oasis:entry colname="col4"/>  
         <oasis:entry colname="col5">7936</oasis:entry>  
         <oasis:entry colname="col6">7168</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">SSVS pPCR</oasis:entry>  
         <oasis:entry colname="col2">7609</oasis:entry>  
         <oasis:entry colname="col3">7378</oasis:entry>  
         <oasis:entry colname="col4"/>  
         <oasis:entry colname="col5">8033</oasis:entry>  
         <oasis:entry colname="col6">7367</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">LASSO PCR</oasis:entry>  
         <oasis:entry colname="col2">7445</oasis:entry>  
         <oasis:entry colname="col3">7065</oasis:entry>  
         <oasis:entry colname="col4"/>  
         <oasis:entry colname="col5">7935</oasis:entry>  
         <oasis:entry colname="col6">7067</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">LASSO pPCR</oasis:entry>  
         <oasis:entry colname="col2">7616</oasis:entry>  
         <oasis:entry colname="col3">7370</oasis:entry>  
         <oasis:entry colname="col4"/>  
         <oasis:entry colname="col5">8053</oasis:entry>  
         <oasis:entry colname="col6">7352</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup><?xmltex \end{scaleboxenv}?></oasis:table></table-wrap>

      <p>After exploring the model framework using a simulation study, we applied our
models to the historical observer data. We fit the eight models to the data
and present the results from LOO in Table <xref ref-type="table" rid="Ch1.T2"/>. Examination of the
LOO Pareto tail parameter plots in Fig. <xref ref-type="fig" rid="Ch1.F4"/>a and b
identified an outlier occurring at data point 1452, corresponding to an
unrealistic mean mid-day July temperature measurement of
42 F (7.8 <inline-formula><mml:math id="M304" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>C).
After removing the outlier, we fit the models again. Interestingly, the
Gaussian model fits without the outlier showed less predictive skill (larger
LOO values in Table <xref ref-type="table" rid="Ch1.T2"/>) and more model misspecification (a larger
tail parameter estimate in Fig. <xref ref-type="fig" rid="Ch1.F4"/>e and f). The decreased
performance when removing the outlier is explained by the influence of the
outlier on the pooled variance estimate; removing the outlier shrinks the
pooled variance estimate and the Gaussian data model is less able to
accommodate slightly outlying points with the smaller variance. For the
robust models, quality of model fit was slightly improved when fit with the
outlier removed (see Table <xref ref-type="table" rid="Ch1.T2"/>), but there is little change in
model misspecification as defined by large Pareto tail parameter estimates
(Fig. <xref ref-type="fig" rid="Ch1.F4"/>c, d, g, and h).</p>
      <p>With the outlier removed, the best predictive models are robust PCR and
robust pPCR due to having the smallest LOO scores (see Table <xref ref-type="table" rid="Ch1.T2"/>).
Based purely on LOO, it appears the best overall predictive model is robust
PCR, although the increasing sample sizes through time weight the LOO score
to predictions of the most recent years. All models still show evidence of
misspecification because some Pareto tail parameter estimates are greater
than 0.5 (Fig. <xref ref-type="fig" rid="Ch1.F4"/>e, f, g, and h), but the removal of the
outlier improved model fit in general. Figure <xref ref-type="fig" rid="Ch1.F4"/>g and h
suggest that the robust pPCR model might be preferable to the robust PCR
model farther back in time because, for that time period, a much greater
proportion of Pareto tail parameter estimates are less than 0.5 in the robust
pPCR model than the robust PCR model. Hence, there is some evidence that the
choice of model is dependent on the desired inference. If inference over the
entire period is desired, both robust PCR and robust pPCR predict with skill.
If inference is desired on the years furthest back in time,
Fig. <xref ref-type="fig" rid="Ch1.F4"/> suggests that robust pPCR predictions are
preferable.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4" specific-use="star"><caption><p>Historical observer station data LOO Pareto shape estimates with
<bold>(a, b, c, d)</bold> and without <bold>(e, f, g, h)</bold> outlying observation.
Values less than 0.5 show good model performance and values over 1.0 show
poor model performance. Note the presence of the outlier in the upper right
of <bold>(a)</bold> and <bold>(b)</bold> for observation 1452. Also evident is less
model misspecification for the pPCR models <bold>(b, d, f, h)</bold> for
observations furthest back in time. </p></caption>
        <?xmltex \igopts{width=384.112205pt}?><graphic xlink:href="https://ascmo.copernicus.org/articles/3/1/2017/ascmo-3-1-2017-f04.pdf"/>

      </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><caption><p>Reconstruction of mean mid-day July temperature using the robust PCR
model for 4 years. Figures show posterior predictive mean <bold>(a)</bold> and
standard deviation <bold>(b)</bold>.</p></caption>
        <?xmltex \igopts{width=355.659449pt}?><graphic xlink:href="https://ascmo.copernicus.org/articles/3/1/2017/ascmo-3-1-2017-f05.pdf"/>

      </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6" specific-use="star"><caption><p>Reconstruction of the mean mid-day July temperature using the robust
probabilistic PCR model for 4 years. Figures show posterior predictive mean
<bold>(a)</bold> and standard deviation <bold>(b)</bold>.</p></caption>
        <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://ascmo.copernicus.org/articles/3/1/2017/ascmo-3-1-2017-f06.pdf"/>

      </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7" specific-use="star"><caption><p>Posterior predictions of a time series of mid-day July temperature
with associated 95 % credible intervals at Champaign, Illinois, Detroit,
Michigan, Madison, Wisconsin, and Minneapolis, Minnesota <bold>(a)</bold>. The
temporal change in the number of observations is shown in <bold>(b)</bold>. Note
that the uncertainties in the point reconstructions are smallest in years
with larger samples and largest in years with few samples. In general, the
robust pPCR credible intervals are larger than the robust PCR credible
intervals. The differences in reconstruction in years 1847 and 1848 at
Champaign, Illinois, and Detroit, Michigan, demonstrate the bias/variance
trade-offs in predictions between the two models. The robust PCR predictions
tend toward lower variance and the robust pPCR predictions tend to lower bias
with respect to the spatially averaged annual mean mid-day July temperature.</p></caption>
        <?xmltex \igopts{width=412.564961pt}?><graphic xlink:href="https://ascmo.copernicus.org/articles/3/1/2017/ascmo-3-1-2017-f07.pdf"/>

      </fig>

      <p>To visualize our results, we plot reconstructions of 4 years of the
historical temperature surfaces using the robust PCR model
(Fig. <xref ref-type="fig" rid="Ch1.F5"/>) and the robust pPCR model
(Fig. <xref ref-type="fig" rid="Ch1.F6"/>). Visual comparison of the reconstructions
illustrates differences in the two models. The robust PCR model assumes the
climate patterns in the observational data are without error; the stronger
influence of these patterns is seen in the posterior predictive mean surface
(Fig. <xref ref-type="fig" rid="Ch1.F5"/>a), particularly for year 1847. In comparison, the
robust pPCR model shows less influence of these climate patterns in the
posterior predictive mean (Fig. <xref ref-type="fig" rid="Ch1.F6"/>a), particularly for year
1847. Because the robust pPCR model includes the assumption that some of the
pattern is noise, the reconstructions have less spatial structure in years
with little data. Differences in posterior predictive standard deviations are
also evident (Figs. <xref ref-type="fig" rid="Ch1.F5"/>b and <xref ref-type="fig" rid="Ch1.F6"/>b), where the
robust PCR model shows reduction of uncertainty in the spatial locations near
observations and higher uncertainty away from the observations, while the
robust pPCR model has posterior predictive standard deviations that are more
spatially diffuse and perhaps more realistic given the lack of spatial
predictive ability seen in the simulation study. The reconstructed
temperature surfaces for both the robust PCR and robust pPCR models for every
year are included in the Supplement.</p>
      <p>By using the spatial structure in the current-era analog data, we generated
temperature predictions at unobserved locations with corresponding
uncertainties. We chose four locations, Champaign, Illinois, Detroit,
Michigan, Madison, Wisconsin, and Minneapolis, Minnesota, and show the time
series of temperature predictions in Fig. <xref ref-type="fig" rid="Ch1.F7"/>a. From these time
series, we see smaller standard deviations for the years with more historical
observer period observations (sample sizes are shown in
Fig. <xref ref-type="fig" rid="Ch1.F7"/>b), with greater uncertainty the further we go back in
time. We can also see a general trend for the standard deviations of the
robust PCR model to be smaller than the robust pPCR model (the red intervals
are more often inside the blue intervals). The robust PCR model is also more
likely to have structure in the mean that may not be warranted when the
sample size is low (i.e., the spike at Champaign and Detroit that is not
present in Madison or Minneapolis in 1847 and 1848 with sample sizes of 3).
The two models show evidence of a bias/variance trade-off, with the robust
PCR tending toward spatially structured predictions with smaller variance and
the robust pPCR model providing less spatially structured predictions with
larger variance.</p>
</sec>
<sec id="Ch1.S7" sec-type="conclusions">
  <title>Conclusions</title>
      <p>There are many challenges inherent in modeling paleoclimate data. Due to the
lack of direct measurements of climate, paleoclimate reconstructions must
rely on sparse, noisy proxies of climate. The nuances of paleoclimate data
often require specialized modeling techniques and careful investigation into
modeling assumptions and performance. In addition, care is needed to properly
validate paleoclimate reconstruction skill. In summary, we extended principal
component regression methods, applied regularization techniques to choose
important principal components, developed robust models to account for the
presence of outliers, and explored the use of a probabilistic principal
component model to account for measurement uncertainty in the spatially rich
current-era analog data. By rigorously evaluating the predictive skill of our
models, we were able to explore our extensions of PCR for climate
reconstruction, laying the groundwork for future developments with more
complex climate data than PRISM temperature surfaces. The models presented in
this paper would be good candidates for modeling climate variables that are
strongly non-stationary and non-Gaussian (e.g., wind speed or precipitation),
but these extensions are the subject of ongoing research.</p>
      <p>Within our modeling framework, we presented a simulation study for evaluating
paleoclimate reconstructions using proper scoring rules. By using proper
scoring rules and exploring model performance in a simulation framework, we
have stronger support for the quality of the reconstruction. We presented
three statistical scoring rules and explored their strengths and weaknesses.
MSPE is a commonly used and easy to understand scoring rule, but is not
proper in general and only uses a point prediction, ignoring the
probabilistic inference that is gained by using Bayesian techniques. The CRPS
is proper and allows for direct comparison of point predictions and
probabilistic predictions, but requires out-of-sample validation data or
computationally expensive cross-validation. The use of MSPE and CRPS scoring
rules allowed for exploration of the empirical properties of the
computationally efficient LOO approximation to leave-one-out
cross-validation. Our use of LOO to score the historical observer period
model predictions not only enabled us to perform model selection, but also
aided in diagnosing an outlying observation and refining model fit.</p>
      <p>The methods presented in this paper could be applied to other historical
datasets at different locations around the world, further extending the
spatially explicit empirical record of climate further back in time while
rigorously accounting for uncertainties. The methods we presented could also
be extended to model temperature and precipitation for each month of the year
by including a seasonal component in the calibration model and by modeling
dynamics at appropriate timescales. There are many datasets that could be
used within this framework as the current-era analog, including modern
satellite data. The different options of current-era analog datasets present
a trade-off between the number of records available as analogs and the
quality of the data. If there are occasionally rare climate processes that
occur, it seems that a longer record of climate analogs would be preferred.
If the climate processes are relatively stable in time but vary highly in
space, a shorter and more precise modern dataset that is not model
interpolated might be preferred. In addition, use of highly precise
current-era analog data could reduce or eliminate the need to account for
measurement error in the current-era analog data.</p>
      <p>Ultimately, our temperature reconstructions extend the climatological record
in the Upper Midwestern US further into the past. These temperature
reconstructions, with their associated uncertainties, can be used to gain
better understanding of the influences of climate on the biological and
ecological processes observed in the region. By backcasting mean mid-day July
temperature with our models, we gain the potential to better understand how
climate has changed, and this knowledge could be used to improve future
climate reconstructions. Many of the techniques and methods we used –
modeling principal components with a probabilistic model, hierarchical
pooling to borrow strength among years with sparse and dense data, model
selection and regularization, and proper model evaluation – can be adapted
and used in future climate reconstruction problems.</p>
</sec>
<sec id="Ch1.S8">
  <title>Data availability</title>
      <p>The historical observer data are available from <uri>http://www.isws.illinois.edu/atmos/clirecord.asp</uri> and the current-era analog data are
available from <uri>http://www.prism.oregonstate.edu/</uri>. Compiled data and code can be accessed on gitHub at <uri>https://github.com/jtipton25/observer</uri>
(<ext-link xlink:href="http://dx.doi.org/10.5281/zenodo.242996" ext-link-type="DOI">10.5281/zenodo.242996</ext-link>).</p>
</sec>

      
      </body>
    <back><app-group>
        <supplementary-material position="anchor"><p><bold>The Supplement related to this article is available online at <inline-supplementary-material xlink:href="http://dx.doi.org/10.5194/ascmo-3-1-2017-supplement" xlink:title="pdf">doi:10.5194/ascmo-3-1-2017-supplement</inline-supplementary-material>.</bold></p></supplementary-material>
        </app-group><notes notes-type="competinginterests">

      <p>The authors declare that they have no conflict of
interest.</p>
  </notes><ack><title>Acknowledgements</title><p>This paper was greatly improved based on the comments from the two anonymous
reviewers and the associate editor. The authors also thank Ben Bird, Kristin
Broms, Brian Brost, Franny Buderman, Trevor Hefley, and Henry Scharf for
their conversations and input on this work. This research is based upon work
carried out by the PalEON Project (paleonproject.org) with support from the
National Science Foundation MacroSystems Biology program under grant no.
DEB-1241856. Any use of trade, firm, or product names is for descriptive
purposes only and does not imply endorsement by the U.S. Government. Code and
data found in this paper can be accessed at
<uri>https://github.com/jtipton25/observer</uri> (<ext-link xlink:href="http://dx.doi.org/10.5281/zenodo.242996" ext-link-type="DOI">10.5281/zenodo.242996</ext-link>).<?xmltex \hack{\newline}?><?xmltex \hack{\newline}?>Edited
by: A. M. Schmidt <?xmltex \hack{\newline}?> Reviewed by: two anonymous referees</p></ack><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Andsager et al.(2004)Andsager, M.C., and
Spinar</label><mixed-citation>
Andsager, K., Ross, T., Kruk, M.C., and Spinar, M. L.: Climate database modernization program:
pre-20th century task – key climate observations recorded since the founding
of America, 1700s–1800s, in: Combined preprints: 84th AMS annual meeting :
20th Conference on Weather Analysis and Forecasting/16th Conference on
Numerical Weather Prediction, Seattle Washington, Boston, MA, American
Meteorological Society, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Barboza et al.(2014)Barboza, Li, Tingley, and
Viens</label><mixed-citation>
Barboza, L., Li, B., Tingley, M., and Viens, F.: Reconstructing past
temperatures from natural proxies and estimated climate forcings using
short-and long-memory models, Ann. Appl. Stat., 8,
1966–2001, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Bell and Ogilvie(1978)</label><mixed-citation>
Bell, W. and Ogilvie, A.: Weather compilations as a source of data for the
reconstruction of European climate during the medieval period, Climatic
Change, 1, 331–348, 1978.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Bernardo and Smith(2009)</label><mixed-citation>
Bernardo, J. M. and Smith, A.: Bayesian Theory, vol. 405, John Wiley &amp; Sons,
2009.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Brázdil et al.(2006)Brázdil, Kundzewicz, and
Benito</label><mixed-citation>
Brázdil, R., Kundzewicz, Z., and Benito, G.: Historical hydrology for
studying flood risk in Europe, Hydrolog. Sci. J., 51, 739–764,
2006.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Buonaccorsi(2010)</label><mixed-citation>
Buonaccorsi, J. P.: Measurement Error: Models, Methods, and Applications, CRC
Press, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Carroll et al.(2006)Carroll, Ruppert, Stefanski, and
Crainiceanu</label><mixed-citation>
Carroll, R. J., Ruppert, D., Stefanski, L. A., and Crainiceanu, C. M.:
Measurement Error in Nonlinear Models: A Modern Perspective, CRC press, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>CDMP(2016)</label><mixed-citation>CDMP: 19th Century Forts and Voluntary Observers Database Build
Project, available at:
<uri>http://www.isws.illinois.edu/atmos/clirecord.asp</uri>, last access:
21 October 2016.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Cook et al.(1994)Cook, Briffa, and Jones</label><mixed-citation>
Cook, E. R., Briffa, K., and Jones, P.: Spatial regression methods in
dendroclimatology: A review and comparison of two techniques, Int.
J.  Climatol., 14, 379–402, 1994.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Efron et al.(2004)Efron, Hastie, Johnstone, Tibshirani
et al.</label><mixed-citation>
Efron, B., Hastie, T., Johnstone, I., Tibshirani, R., Efron, B., Hastie, T., Johnstone, I., and Tibshirani, R.: Least angle
regression,   Ann. Stat., 32, 407–499, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Fuller(2009)</label><mixed-citation>
Fuller, W. A.: Measurement Error Models, vol. 305, John Wiley &amp; Sons, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Geisser and Eddy(1979)</label><mixed-citation>
Geisser, S. and Eddy, W.: A predictive approach to model selection, J.
Am. Stat. Assoc., 74, 153–160, 1979.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Gelman and Hill(2006)</label><mixed-citation>
Gelman, A. and Hill, J.: Data Analysis Using Regression and
Multilevel/Hierarchical Models, Cambridge University Press, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Gelman and Rubin(1992)</label><mixed-citation>
Gelman, A. and Rubin, D. B.: Inference from iterative simulation using
multiple
sequences, Stat. Sci., 7, 457–472, 1992.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>George and McCulloch(1993)</label><mixed-citation>
George, E. I. and McCulloch, R. E.: Variable selection via Gibbs sampling,
J.
Am. Stat. Assoc., 88, 881–889, 1993.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Gneiting(2011)</label><mixed-citation>
Gneiting, T.: Making and evaluating point forecasts, J.
Am. Stat. Assoc., 106, 746–762, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Gneiting and Raftery(2007)</label><mixed-citation>
Gneiting, T. and Raftery, A.: Strictly proper scoring rules, prediction, and
estimation, J.
Am. Stat. Assoc., 102, 359–378,
2007.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Gneiting et al.(2007)Gneiting, Balabdaoui, and
Raftery</label><mixed-citation>
Gneiting, T., Balabdaoui, F., and Raftery, A.: Probabilistic forecasts,
calibration and sharpness, J. Roy. Stat. Soc. B, 69, 243–268, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Gotway and Young(2002)</label><mixed-citation>
Gotway, C. A. and Young, L.: Combining incompatible spatial data, J.
Am. Stat. Assoc., 97, 632–648, 2002.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Hadi and Ling(1998)</label><mixed-citation>
Hadi, A. S. and Ling, R.: Some cautionary notes on the use of principal
components regression,  Am. Stat., 52, 15–19, 1998.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Hastie et al.(2005)Hastie, Tibshirani, Friedman, and
Franklin</label><mixed-citation>
Hastie, T., Tibshirani, R., Friedman, J., and Franklin, J.: The elements of
statistical learning: data mining, inference and prediction,   Math.
Intell., 27, 83–85, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Hoerl and Kennard(1970)</label><mixed-citation>
Hoerl, A. E. and Kennard, R. W.: Ridge regression: Biased estimation for
nonorthogonal problems, Technometrics, 12, 55–67, 1970.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Hooten and Hobbs(2015)</label><mixed-citation>
Hooten, M. B. and Hobbs, N.: A guide to Bayesian model selection for
ecologists, Ecol. Monogr., 85, 3–28, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Jolliffe(1982)</label><mixed-citation>
Jolliffe, I. T.: A note on the use of principal components in regression,
Appl. Statist., 31, 300–303, 1982.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Juárez and Steel(2010)</label><mixed-citation>
Juárez, M. A. and Steel, M. F.: Model-based clustering of non-Gaussian
panel data based on skew-t distributions, J. Bus. Econ.
Stat., 28, 52–66, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Kastellet et al.(1998)Kastellet, Nesje, and
Pedersen</label><mixed-citation>
Kastellet, E., Nesje, A., and Pedersen, E.: Reconstructing the palaeoclimate
of
Jæren, Southwestern Norway, for the period 1821–1850, from
historical documentary records, Geogr. Ann. A,   80, 51–65, 1998.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Koopman et al.(2009)Koopman, Shephard, and
Creal</label><mixed-citation>
Koopman, S. J., Shephard, N., and Creal, D.: Testing the assumptions behind
importance sampling, Journal of Econometrics, 149, 2–11, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Lorenz(1956)</label><mixed-citation>
Lorenz, E. N.: Empirical orthogonal functions and statistical weather
prediction, Scientific report no. 1: Statistical forecasting project,
Massachusetts Institute of Technology, Department of Meteorology, 1956.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>Ogilvie(1984)</label><mixed-citation>
Ogilvie, A. E.: The past climate and sea-ice record from Iceland, Part 1:
Data to AD 1780, Climatic Change, 6, 131–152, 1984.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>Park and Casella(2008)</label><mixed-citation>
Park, T. and Casella, G.: The Bayesian lasso, J. Am.
Stat. Assoc., 103, 681–686, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Plummer(2003)</label><mixed-citation>
Plummer, M.: JAGS: A program for analysis of Bayesian graphical
models using Gibbs sampling, in: Proceedings of the 3rd international
workshop on distributed statistical computing, vol. 124,  125 pp., Technische
Universität Wien, Wien, Austria, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Preisendorfer(1988)</label><mixed-citation>
Preisendorfer, R.: Principal Component Analysis in Meteorology and
Oceanography, Developments in Atmospheric Science, 17, Elsevier,
1988.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>PRISM Climate Group, Oregon State University(2016)</label><mixed-citation>PRISM Climate Group, Oregon State University: available at:
<uri>http://prism.oregonstate.edu</uri>, last access: 21 October 2016.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>R Core Team(2016)</label><mixed-citation>
R Core Team: R: A Language and Environment for Statistical Computing, R
Foundation for Statistical Computing, Vienna, Austria,
2016.</mixed-citation></ref>
      <ref id="bib1.bibx35"><label>Rutherford et al.(2005)Rutherford, Mann, Osborn, Briffa, Jones,
Bradley, and Hughes</label><mixed-citation>
Rutherford, S., Mann, M., Osborn, T., Briffa, K., Jones, P., Bradley, R., and
Hughes, M.: Proxy-based Northern Hemisphere surface temperature
reconstructions: Sensitivity to method, predictor network, target season,
and target domain, J. Climate, 18, 2308–2329, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>Tibshirani(1996)</label><mixed-citation>
Tibshirani, R.: Regression shrinkage and selection via the lasso, J.
Roy. Stat. Soc.  B, 58, 267–288, 1996.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Tingley and Huybers(2010a)</label><mixed-citation>
Tingley, M. P. and Huybers, P.: A Bayesian algorithm for reconstructing
climate anomalies in space and time. Part I: Development and
applications to paleoclimate reconstruction problems, J. Climate, 23,
2759–2781, 2010a.</mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Tingley and Huybers(2010b)</label><mixed-citation>
Tingley, M. P. and Huybers, P.: A Bayesian algorithm for reconstructing
climate anomalies in space and time. Part II: Comparison with the
regularized expectation-maximization algorithm, J. Climate, 23,
2782–2800, 2010b.</mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Tipping and Bishop(1999)</label><mixed-citation>
Tipping, M. E. and Bishop, C.: Probabilistic principal component analysis,
J. Roy. Stat. Soc. B,
61, 611–622, 1999.</mixed-citation></ref>
      <ref id="bib1.bibx40"><label>Tipton et al.(2016)Tipton, Hooten, Pederson, Tingley, and
Bishop</label><mixed-citation>
Tipton, J., Hooten, M., Pederson, N., Tingley, M., and Bishop, D.:
Reconstruction of late Holocene climate based on tree growth and
mechanistic hierarchical models, Environmetrics, 27, 42–54, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx41"><label>Vehtari and Gelman(2015)</label><mixed-citation>
Vehtari, A. and Gelman, A.: Pareto Smoothed Importance Sampling, arXiv
preprint
arXiv:1507.02646v2, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx42"><label>Vehtari et al.(2016a)Vehtari, Gelman, and Gabry</label><mixed-citation>Vehtari, A., Gelman, A., and Gabry, J.: loo: Efficient leave-one-out
cross-validation and WAIC for Bayesian models, R package version 0.1.6,
available at: <uri>https://github.com/jgabry/loo</uri> (last access: 21 October 2016),
2016a.</mixed-citation></ref>
      <ref id="bib1.bibx43"><label>Vehtari et al.(2016b)Vehtari, Gelman, and
Gabry</label><mixed-citation>
Vehtari, A., Gelman, A., and Gabry, J.: Practical Bayesian model evaluation
using leave-one-out cross-validation and WAIC, arXiv preprint
arXiv:1507.04544, 2016b.</mixed-citation></ref>
      <ref id="bib1.bibx44"><label>Wang(2012)</label><mixed-citation>
Wang, L.: Bayesian principal component regression with data-driven component
selection, J. Appl. Stat., 39, 1177–1189, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx45"><label>Werner and Tingley(2015)</label><mixed-citation>Werner, J. P. and Tingley, M. P.: Technical Note: Probabilistically
constraining proxy age–depth models within a Bayesian hierarchical
reconstruction model, Clim. Past, 11, 533–545, <ext-link xlink:href="http://dx.doi.org/10.5194/cp-11-533-2015" ext-link-type="DOI">10.5194/cp-11-533-2015</ext-link>,
2015.</mixed-citation></ref>
      <ref id="bib1.bibx46"><label>Wood(2006)</label><mixed-citation>Wood, S.: Generalized Additive Models: An Introduction with R, CRC press,
2006.
 </mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx47"><label>Wood(2011)</label><mixed-citation>
Wood, S. N.: Fast stable restricted maximum likelihood and marginal
likelihood
estimation of semiparametric generalized linear models, J. Roy.
Stat. Soc. B, 73, 3–36, 2011.</mixed-citation></ref>

  </ref-list><app-group content-type="float"><app><title/>

    </app></app-group></back>
    <!--<article-title-html>Reconstruction of spatio-temporal temperature from sparse historical records using robust probabilistic principal component regression</article-title-html>
<abstract-html><p class="p">Scientific records of temperature and precipitation have been kept
for several hundred years, but for many areas, only a shorter record exists.
To understand climate change, there is a need for rigorous statistical
reconstructions of the paleoclimate using proxy data. Paleoclimate proxy data
are often sparse, noisy, indirect measurements of the climate process of
interest, making each proxy uniquely challenging to model statistically. We
reconstruct spatially explicit temperature surfaces from sparse and noisy
measurements recorded at historical United States military forts and other
observer stations from 1820 to 1894. One common method for reconstructing the
paleoclimate from proxy data is principal component regression (PCR). With
PCR, one learns a statistical relationship between the paleoclimate proxy
data and a set of climate observations that are used as patterns for
potential reconstruction scenarios. We explore PCR in a Bayesian hierarchical
framework, extending classical PCR in a variety of ways. First, we model the
latent principal components probabilistically, accounting for measurement
error in the observational data. Next, we extend our method to better
accommodate outliers that occur in the proxy data. Finally, we explore
alternatives to the truncation of lower-order principal components using
different regularization techniques. One fundamental challenge in
paleoclimate reconstruction efforts is the lack of out-of-sample data for
predictive validation. Cross-validation is of potential value, but is
computationally expensive and potentially sensitive to outliers in sparse
data scenarios. To overcome the limitations that a lack of out-of-sample
records presents, we test our methods using a simulation study, applying
proper scoring rules including a computationally efficient approximation to
leave-one-out cross-validation using the log score to validate model
performance. The result of our analysis is a spatially explicit
reconstruction of spatio-temporal temperature from a very sparse historical
record.</p></abstract-html>
<ref-html id="bib1.bib1"><label>Andsager et al.(2004)Andsager, M.C., and
Spinar</label><mixed-citation>
Andsager, K., Ross, T., Kruk, M.C., and Spinar, M. L.: Climate database modernization program:
pre-20th century task – key climate observations recorded since the founding
of America, 1700s–1800s, in: Combined preprints: 84th AMS annual meeting :
20th Conference on Weather Analysis and Forecasting/16th Conference on
Numerical Weather Prediction, Seattle Washington, Boston, MA, American
Meteorological Society, 2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Barboza et al.(2014)Barboza, Li, Tingley, and
Viens</label><mixed-citation>
Barboza, L., Li, B., Tingley, M., and Viens, F.: Reconstructing past
temperatures from natural proxies and estimated climate forcings using
short-and long-memory models, Ann. Appl. Stat., 8,
1966–2001, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Bell and Ogilvie(1978)</label><mixed-citation>
Bell, W. and Ogilvie, A.: Weather compilations as a source of data for the
reconstruction of European climate during the medieval period, Climatic
Change, 1, 331–348, 1978.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Bernardo and Smith(2009)</label><mixed-citation>
Bernardo, J. M. and Smith, A.: Bayesian Theory, vol. 405, John Wiley &amp; Sons,
2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Brázdil et al.(2006)Brázdil, Kundzewicz, and
Benito</label><mixed-citation>
Brázdil, R., Kundzewicz, Z., and Benito, G.: Historical hydrology for
studying flood risk in Europe, Hydrolog. Sci. J., 51, 739–764,
2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Buonaccorsi(2010)</label><mixed-citation>
Buonaccorsi, J. P.: Measurement Error: Models, Methods, and Applications, CRC
Press, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Carroll et al.(2006)Carroll, Ruppert, Stefanski, and
Crainiceanu</label><mixed-citation>
Carroll, R. J., Ruppert, D., Stefanski, L. A., and Crainiceanu, C. M.:
Measurement Error in Nonlinear Models: A Modern Perspective, CRC press, 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>CDMP(2016)</label><mixed-citation>
CDMP: 19th Century Forts and Voluntary Observers Database Build
Project, available at:
<a href="http://www.isws.illinois.edu/atmos/clirecord.asp" target="_blank">http://www.isws.illinois.edu/atmos/clirecord.asp</a>, last access:
21 October 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Cook et al.(1994)Cook, Briffa, and Jones</label><mixed-citation>
Cook, E. R., Briffa, K., and Jones, P.: Spatial regression methods in
dendroclimatology: A review and comparison of two techniques, Int.
J.  Climatol., 14, 379–402, 1994.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Efron et al.(2004)Efron, Hastie, Johnstone, Tibshirani
et al.</label><mixed-citation>
Efron, B., Hastie, T., Johnstone, I., Tibshirani, R., Efron, B., Hastie, T., Johnstone, I., and Tibshirani, R.: Least angle
regression,   Ann. Stat., 32, 407–499, 2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Fuller(2009)</label><mixed-citation>
Fuller, W. A.: Measurement Error Models, vol. 305, John Wiley &amp; Sons, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Geisser and Eddy(1979)</label><mixed-citation>
Geisser, S. and Eddy, W.: A predictive approach to model selection, J.
Am. Stat. Assoc., 74, 153–160, 1979.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Gelman and Hill(2006)</label><mixed-citation>
Gelman, A. and Hill, J.: Data Analysis Using Regression and
Multilevel/Hierarchical Models, Cambridge University Press, 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Gelman and Rubin(1992)</label><mixed-citation>
Gelman, A. and Rubin, D. B.: Inference from iterative simulation using
multiple
sequences, Stat. Sci., 7, 457–472, 1992.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>George and McCulloch(1993)</label><mixed-citation>
George, E. I. and McCulloch, R. E.: Variable selection via Gibbs sampling,
J.
Am. Stat. Assoc., 88, 881–889, 1993.
</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Gneiting(2011)</label><mixed-citation>
Gneiting, T.: Making and evaluating point forecasts, J.
Am. Stat. Assoc., 106, 746–762, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Gneiting and Raftery(2007)</label><mixed-citation>
Gneiting, T. and Raftery, A.: Strictly proper scoring rules, prediction, and
estimation, J.
Am. Stat. Assoc., 102, 359–378,
2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Gneiting et al.(2007)Gneiting, Balabdaoui, and
Raftery</label><mixed-citation>
Gneiting, T., Balabdaoui, F., and Raftery, A.: Probabilistic forecasts,
calibration and sharpness, J. Roy. Stat. Soc. B, 69, 243–268, 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Gotway and Young(2002)</label><mixed-citation>
Gotway, C. A. and Young, L.: Combining incompatible spatial data, J.
Am. Stat. Assoc., 97, 632–648, 2002.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Hadi and Ling(1998)</label><mixed-citation>
Hadi, A. S. and Ling, R.: Some cautionary notes on the use of principal
components regression,  Am. Stat., 52, 15–19, 1998.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Hastie et al.(2005)Hastie, Tibshirani, Friedman, and
Franklin</label><mixed-citation>
Hastie, T., Tibshirani, R., Friedman, J., and Franklin, J.: The elements of
statistical learning: data mining, inference and prediction,   Math.
Intell., 27, 83–85, 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Hoerl and Kennard(1970)</label><mixed-citation>
Hoerl, A. E. and Kennard, R. W.: Ridge regression: Biased estimation for
nonorthogonal problems, Technometrics, 12, 55–67, 1970.
</mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Hooten and Hobbs(2015)</label><mixed-citation>
Hooten, M. B. and Hobbs, N.: A guide to Bayesian model selection for
ecologists, Ecol. Monogr., 85, 3–28, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Jolliffe(1982)</label><mixed-citation>
Jolliffe, I. T.: A note on the use of principal components in regression,
Appl. Statist., 31, 300–303, 1982.
</mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Juárez and Steel(2010)</label><mixed-citation>
Juárez, M. A. and Steel, M. F.: Model-based clustering of non-Gaussian
panel data based on skew-t distributions, J. Bus. Econ.
Stat., 28, 52–66, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Kastellet et al.(1998)Kastellet, Nesje, and
Pedersen</label><mixed-citation>
Kastellet, E., Nesje, A., and Pedersen, E.: Reconstructing the palaeoclimate
of
Jæren, Southwestern Norway, for the period 1821–1850, from
historical documentary records, Geogr. Ann. A,   80, 51–65, 1998.
</mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Koopman et al.(2009)Koopman, Shephard, and
Creal</label><mixed-citation>
Koopman, S. J., Shephard, N., and Creal, D.: Testing the assumptions behind
importance sampling, Journal of Econometrics, 149, 2–11, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Lorenz(1956)</label><mixed-citation>
Lorenz, E. N.: Empirical orthogonal functions and statistical weather
prediction, Scientific report no. 1: Statistical forecasting project,
Massachusetts Institute of Technology, Department of Meteorology, 1956.
</mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Ogilvie(1984)</label><mixed-citation>
Ogilvie, A. E.: The past climate and sea-ice record from Iceland, Part 1:
Data to AD 1780, Climatic Change, 6, 131–152, 1984.
</mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Park and Casella(2008)</label><mixed-citation>
Park, T. and Casella, G.: The Bayesian lasso, J. Am.
Stat. Assoc., 103, 681–686, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Plummer(2003)</label><mixed-citation>
Plummer, M.: JAGS: A program for analysis of Bayesian graphical
models using Gibbs sampling, in: Proceedings of the 3rd international
workshop on distributed statistical computing, vol. 124,  125 pp., Technische
Universität Wien, Wien, Austria, 2003.
</mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Preisendorfer(1988)</label><mixed-citation>
Preisendorfer, R.: Principal Component Analysis in Meteorology and
Oceanography, Developments in Atmospheric Science, 17, Elsevier,
1988.
</mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>PRISM Climate Group, Oregon State University(2016)</label><mixed-citation>
PRISM Climate Group, Oregon State University: available at:
<a href="http://prism.oregonstate.edu" target="_blank">http://prism.oregonstate.edu</a>, last access: 21 October 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>R Core Team(2016)</label><mixed-citation>
R Core Team: R: A Language and Environment for Statistical Computing, R
Foundation for Statistical Computing, Vienna, Austria,
2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Rutherford et al.(2005)Rutherford, Mann, Osborn, Briffa, Jones,
Bradley, and Hughes</label><mixed-citation>
Rutherford, S., Mann, M., Osborn, T., Briffa, K., Jones, P., Bradley, R., and
Hughes, M.: Proxy-based Northern Hemisphere surface temperature
reconstructions: Sensitivity to method, predictor network, target season,
and target domain, J. Climate, 18, 2308–2329, 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Tibshirani(1996)</label><mixed-citation>
Tibshirani, R.: Regression shrinkage and selection via the lasso, J.
Roy. Stat. Soc.  B, 58, 267–288, 1996.
</mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Tingley and Huybers(2010a)</label><mixed-citation>
Tingley, M. P. and Huybers, P.: A Bayesian algorithm for reconstructing
climate anomalies in space and time. Part I: Development and
applications to paleoclimate reconstruction problems, J. Climate, 23,
2759–2781, 2010a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Tingley and Huybers(2010b)</label><mixed-citation>
Tingley, M. P. and Huybers, P.: A Bayesian algorithm for reconstructing
climate anomalies in space and time. Part II: Comparison with the
regularized expectation-maximization algorithm, J. Climate, 23,
2782–2800, 2010b.
</mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Tipping and Bishop(1999)</label><mixed-citation>
Tipping, M. E. and Bishop, C.: Probabilistic principal component analysis,
J. Roy. Stat. Soc. B,
61, 611–622, 1999.
</mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Tipton et al.(2016)Tipton, Hooten, Pederson, Tingley, and
Bishop</label><mixed-citation>
Tipton, J., Hooten, M., Pederson, N., Tingley, M., and Bishop, D.:
Reconstruction of late Holocene climate based on tree growth and
mechanistic hierarchical models, Environmetrics, 27, 42–54, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Vehtari and Gelman(2015)</label><mixed-citation>
Vehtari, A. and Gelman, A.: Pareto Smoothed Importance Sampling, arXiv
preprint
arXiv:1507.02646v2, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Vehtari et al.(2016a)Vehtari, Gelman, and Gabry</label><mixed-citation>
Vehtari, A., Gelman, A., and Gabry, J.: loo: Efficient leave-one-out
cross-validation and WAIC for Bayesian models, R package version 0.1.6,
available at: <a href="https://github.com/jgabry/loo" target="_blank">https://github.com/jgabry/loo</a> (last access: 21 October 2016),
2016a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>Vehtari et al.(2016b)Vehtari, Gelman, and
Gabry</label><mixed-citation>
Vehtari, A., Gelman, A., and Gabry, J.: Practical Bayesian model evaluation
using leave-one-out cross-validation and WAIC, arXiv preprint
arXiv:1507.04544, 2016b.
</mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>Wang(2012)</label><mixed-citation>
Wang, L.: Bayesian principal component regression with data-driven component
selection, J. Appl. Stat., 39, 1177–1189, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Werner and Tingley(2015)</label><mixed-citation>
Werner, J. P. and Tingley, M. P.: Technical Note: Probabilistically
constraining proxy age–depth models within a Bayesian hierarchical
reconstruction model, Clim. Past, 11, 533–545, <a href="http://dx.doi.org/10.5194/cp-11-533-2015" target="_blank">doi:10.5194/cp-11-533-2015</a>,
2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>Wood(2006)</label><mixed-citation>
Wood, S.: Generalized Additive Models: An Introduction with R, CRC press,
2006.

</mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>Wood(2011)</label><mixed-citation>
Wood, S. N.: Fast stable restricted maximum likelihood and marginal
likelihood
estimation of semiparametric generalized linear models, J. Roy.
Stat. Soc. B, 73, 3–36, 2011.
</mixed-citation></ref-html>--></article>
