One Sediment Grain Size Standard Changed 8 of 14 Paleotemperature Reconstructions
May 30, 2026 By Jonas Eriksen

For decades, paleoclimatologists have relied on a standard sieve size to isolate foraminifera shells from ocean sediments. That choice, it turns out, may have been quietly skewing our picture of past climates. A new study published in Nature Geoscience finds that switching to a finer mesh alters the reconstructed sea-surface temperature in 8 of 14 published records, with changes of 1–3°C that are large enough to matter for model comparisons.

A single methodological tweak rewrites climate history

The study, led by Dr. Samantha Gibbs at the University of Southampton, began as a routine check. Gibbs and her team were re-running Mg/Ca analyses on archived sediment cores and decided to test whether the standard sieve size—250–300 micrometres—was introducing a bias. They compared results from that fraction with a finer one, 150–250 µm, using published raw data from the NOAA paleoclimate archive.

Of the 14 records they re-evaluated, 8 showed statistically significant shifts in Mg/Ca ratios. In 6 cases, the implied warming trend became stronger; in 2 cases, a previously inferred cooling trend reversed to warming. The remaining 6 records were unchanged within uncertainty. The study estimates that the temperature adjustments range from roughly 1°C to as much as 3°C, depending on the core location and time interval.

“That could be a game changer for how we interpret past interglacials,” said Gibbs in a university press release, though she cautioned that the work is a reanalysis, not a new climate reconstruction. The study is one of several recent efforts to examine how routine laboratory decisions affect large-scale conclusions—a theme also explored in work on incubator humidity settings and catalyst batch lots.

Why grain size matters for paleothermometry

The Mg/Ca ratio in foraminifera shells is a well-established proxy for sea-surface temperature. As water temperature increases, more magnesium is incorporated into the calcium carbonate shell. But the relationship is not purely thermodynamic; it also depends on the growth rate and physiology of the organism. Smaller shells, which come from younger or slower-growing individuals, tend to incorporate less magnesium relative to calcium at a given temperature.

The standard sieve size of 250–300 µm was adopted decades ago, partly for convenience and partly because it yields enough shell material for analysis. But it systematically excludes the smaller fraction, which can be abundant in some sediments. By including the 150–250 µm fraction, Gibbs and colleagues found that the Mg/Ca ratio shifted—in some cases by a margin that changes the inferred temperature by several degrees.

“The effect is not uniform across all species or all ocean basins,” Gibbs noted. The team tested several foraminifera species and found that the magnitude of the shift varied. For some species, the size effect was negligible; for others, it was large enough to flip a cooling trend into a warming one. The study underscores that paleothermometry is not a simple one-to-one translation from chemistry to temperature.

The hidden cost of methodological inertia

Lab protocols in paleoclimatology, like many fields, can persist for decades without systematic revalidation. The standard sieve size is a case in point. It was established in the 1980s and 1990s, when sample sizes were smaller and analytical precision was lower. Since then, instruments have improved, and the amount of shell material needed for a reliable Mg/Ca measurement has decreased. But the protocol has not kept pace.

Peer reviewers rarely question sieve sizes, Gibbs pointed out. “Reviewers ask about age models and calibration, but almost never about the grain size fraction used,” she said. The field’s funding structure may also play a role: agencies tend to prioritize new drilling expeditions and novel proxies over audits of existing methods. Re-running 14 cores for this study cost an estimated $2–4 million in laboratory time and personnel—a modest sum compared to a single ocean drilling campaign, but difficult to justify under current grant metrics.

The economics of paleoscience often incentivize novelty over reproducibility checks. A 2023 survey by the Paleontological Society found that fewer than 15% of paleoclimate studies include any form of method intercomparison. The Gibbs study is part of a growing push to change that, similar to efforts in other fields to re-evaluate long-standing assumptions—for example, how a grant agency's diet rule can shift results in microbiome research.

How the revised temperatures stack up

The team focused on records from the mid-Pliocene (roughly 3 million years ago) and the last interglacial (about 125,000 years ago), two periods often used as analogues for future warming. In the mid-Pliocene, some records had previously suggested that high-latitude sea-surface temperatures were only modestly warmer than today—a finding that climate models struggled to reproduce. After the grain-size adjustment, several of those records now show temperatures 1–2°C higher, bringing them closer to model predictions.

For the last interglacial, the picture is more mixed. Some tropical records shifted toward warmer conditions, while others remained stable. The net effect is to reduce the spread among different reconstructions for the same time interval, which may help resolve long-standing disagreements about how warm the last interglacial really was.

“We are not claiming that our reanalysis is the final word,” Gibbs emphasized. “We are showing that one methodological choice can change the story. That means the community needs to revisit these records with consistent protocols before drawing firm conclusions.” The study used only published raw data, which are freely available through the NOAA archives. That open-data policy made the reanalysis possible at a fraction of the cost of collecting new cores.

Implications for global climate models

Paleoclimate reconstructions are routinely used to evaluate the performance of global climate models. If a model cannot reproduce past warm periods, it may be missing important feedbacks—or the reconstructions themselves may be biased. The Gibbs study suggests that at least some model-data mismatches may stem from methodological artifacts rather than model deficiencies.

For the mid-Pliocene, the revised temperatures bring several high-latitude records into better agreement with model simulations that include enhanced ocean heat transport. The magnitude of the adjustment is comparable to the difference between different model versions, which means that model evaluation studies may need to account for grain-size effects. Some climate sensitivity estimates, which rely on past warm periods to constrain future warming, could shift upward by roughly 0.5°C if the adjusted records are confirmed.

“This is not a crisis for paleoclimatology, but it is a reminder that the devil is in the details,” said Dr. Michaela Tabor, a paleoclimate modeller at the University of Connecticut who was not involved in the study. “We often treat proxy data as truth, but they are measurements with their own uncertainties. Studies like this help us quantify those uncertainties.”

Lessons for the economics of paleoscience

The study’s modest cost relative to its interpretive impact raises questions about how research funding is allocated. A single ocean drilling expedition can cost $10–20 million and yield cores that take years to analyse. Reanalysing those cores with a different sieve size costs perhaps $200,000–300,000 per core, a fraction of the original investment. Yet such reanalyses are rarely funded because they are seen as “not novel.”

Gibbs and her colleagues argue that funding agencies should set aside a portion of their budgets for methodological audits and reproducibility studies. “We need to treat methods as a living part of science, not a fixed recipe,” Gibbs said. The open-data policies of repositories like NOAA and PANGAEA make such audits feasible, but only if the community values them.

The same logic applies to other fields. A recent study on sediment core date shifts similarly showed how small changes in chronology can alter interpretations of ocean circulation. Together, these findings suggest that paleoscience may be sitting on a wealth of under-exploited data that could yield new insights with relatively modest investment—if the incentive structure shifts.

The Gibbs study is not the final word on grain-size effects. Other factors, such as cleaning protocols and instrumental drift, may also introduce biases. But it serves as a concrete example of how a seemingly minor methodological choice can propagate through the scientific literature. As Gibbs put it, “We are not saying the old records are wrong. We are saying they are incomplete. And that is a very different thing.”

Additional case study: The North Atlantic records

To illustrate the magnitude of the effect, consider three specific cores from the North Atlantic region that were part of the reanalysis. Core ODP 982, located near the Rockall Plateau, had previously indicated that sea-surface temperatures during the mid-Pliocene were only about 2°C warmer than modern. After switching to the 150–250 µm fraction, the reconstructed temperature rose to nearly 4°C above modern—a shift of 2°C that brings it in line with model simulations that incorporate stronger meridional overturning circulation. In contrast, core ODP 925 from the equatorial Atlantic showed a modest 0.5°C increase, well within the original uncertainty range. This variability underscores that the grain-size effect is not a simple uniform correction; it depends on local sedimentation rates, foraminifera species composition, and preservation quality.

Another notable example is core MD97-2120 from the southwest Pacific, which previously exhibited a gradual cooling trend across the last interglacial. After reanalysis, the trend reversed to a slight warming, a change of about 1.5°C over the interval. This reversal implies that earlier interpretations of a cool last interglacial in that region may need revision, potentially affecting our understanding of Southern Hemisphere climate dynamics during that period.

These examples highlight the importance of sediment composition: cores with a high abundance of small foraminifera are more susceptible to size-related biases. In cores where the 150–250 µm fraction is sparse, the effect is minimal. Gibbs and her colleagues recommend that future studies routinely report the grain-size distribution and, ideally, analyze multiple size fractions to assess robustness.

Trade-offs and counter-arguments

Some researchers have expressed caution about adopting the finer sieve size universally. One concern is that the 150–250 µm fraction may include a higher proportion of juvenile shells or shells from species with different habitat preferences, potentially introducing a different bias. For example, some species of foraminifera that thrive in colder waters tend to produce smaller shells, so including them might amplify a cold bias. Gibbs acknowledges this possibility but notes that their species-specific analyses showed the opposite effect in most cases—the finer fraction typically indicated warmer temperatures, consistent with a growth-rate effect rather than a habitat effect.

Another trade-off is the increased analytical effort required. The finer fraction yields less carbonate material per sieve, so more sediment must be processed to obtain enough shells for a reliable Mg/Ca measurement. This can increase lab time and cost, especially for cores with low foraminifera abundance. However, Gibbs argues that the additional cost is modest compared to the potential gain in accuracy. “If you invest an extra 10% in sample preparation, you might avoid a 2°C systematic error,” she said.

Critics also point out that the study reanalyzed only 14 records, which is a small fraction of the hundreds of published Mg/Ca records. The generalizability of the findings remains to be tested across a wider range of locations and time periods. Gibbs agrees, and her team has made their raw data and code publicly available to encourage further reanalysis. “We see this as a starting point, not a conclusion,” she said.

Broader implications for reproducibility in paleoscience

The Gibbs study adds to a growing body of literature that examines hidden methodological biases in paleoclimate research. For example, a 2021 study on cleaning protocols for foraminifera shells found that different acid leaching methods can shift Mg/Ca ratios by up to 5%, equivalent to a temperature change of about 1°C. Similarly, a 2022 study on mass spectrometry calibration showed that inter-laboratory differences can introduce systematic offsets of 0.5–1°C. Together, these findings paint a picture of a field where multiple small biases can accumulate, potentially leading to overconfidence in reconstructions.

One way to address this is through community-wide intercomparison exercises, similar to those used in climate modeling. The Paleoclimate Modelling Intercomparison Project (PMIP) has been successful in coordinating model runs, but no equivalent effort exists for proxy data. Gibbs advocates for a “proxy intercomparison project” that would systematically test the impact of different laboratory protocols on a common set of sediment samples. Such an initiative would require funding and coordination, but the payoff could be substantial: a more robust paleoclimate record that can be used with confidence for model evaluation.

In the meantime, the study serves as a practical reminder for researchers. When interpreting published paleotemperature records, it is worth checking the sieve size used and considering whether it might affect the conclusions. For new studies, analyzing multiple size fractions—or at least documenting the size distribution—should become standard practice. As Gibbs concluded, “Science is built on trust, but trust is not a substitute for verification.”

Related Articles