One Lab's Fear Conditioning Protocol Reversed 11 of 20 Mouse Memory Results
May 29, 2026 By Karim Osman

In 2008, a widely cited study reported that mice showed robust retention of fear memories when a tone was paired with a foot shock. The result seemed straightforward: rodents learned to associate the tone with danger and froze when they heard it again. But a decade later, when a replication team tried to reproduce the finding, they discovered something unsettling. By changing the interval between the tone and the shock from zero to roughly two seconds, 11 of the 20 original memory effects reversed direction. Some effects that had been positive became negative; others simply vanished. The story of those two seconds is a case study in how fragile experimental results can be.

A Simple Change, A Reversed Result

The original 2008 experiment used a protocol common in many rodent labs: a tone (the conditioned stimulus) co-terminated with a foot shock (the unconditioned stimulus). That is, the shock began at the same moment the tone ended—what researchers call “immediate shock.” Mice trained this way showed strong freezing when re-exposed to the tone 24 hours later, and even stronger freezing when placed back in the same chamber (contextual fear). The effect sizes were large, typically in the range of Cohen's d around 0.6 to 0.8.

In 2018, a multi-lab replication project attempted to repeat the study exactly. But one participating lab, led by a behavioral neuroscientist at a midwestern university, noticed that their own standard protocol differed from the original: they typically inserted a brief delay—around 2 seconds—between the tone offset and the shock onset. Curious whether this mattered, they ran both versions side by side.

The results were dramatic. With immediate shock, they replicated the original effects reasonably well. With the 2-second delay, 11 of the 20 comparisons flipped: freezing to the tone decreased, contextual freezing increased in some cases, and several effects became non-significant. The pattern was not random; it followed a clear logic tied to which memory system each protocol engaged.

The Two Protocols That Clashed

Fear conditioning protocols vary across labs in ways that are rarely reported in detail. The two that clashed here are known as “delay conditioning” and “trace conditioning,” though the terms are sometimes used loosely. In delay conditioning, the shock overlaps with or immediately follows the tone. In trace conditioning, a gap—the “trace interval”—separates the tone from the shock. The 2-second gap used by the replication team falls into the trace category.

These two protocols are thought to recruit different neural circuits. Delay conditioning depends heavily on the amygdala, a brain region that processes rapid threat signals. Trace conditioning additionally requires the hippocampus, which binds events across time. When the trace interval is present, the mouse must hold a mental representation of the tone during the gap, a process that engages working memory and temporal coding.

The original 2008 study used delay conditioning and found strong amygdala-dependent freezing. The replication team’s trace protocol shifted the balance toward hippocampal involvement. In some mouse strains, this shift enhanced contextual fear (which also relies on the hippocampus) but weakened cued fear. The net result: a reversal of the relative strength of contextual versus cued memory.

Why Timing Matters for Memory

The timing of the shock relative to the tone is not a trivial detail. Trace conditioning requires the animal to sustain attention across the gap, engaging prefrontal cortex and hippocampus in a way that delay conditioning does not. This difference is well established in the literature, but its practical impact on effect sizes is often underestimated.

In delay conditioning, the association between tone and shock is nearly simultaneous, so the amygdala can form a direct link. In trace conditioning, the hippocampus must create a temporal bridge. If the trace interval is too long (more than a few seconds), many rodents fail to learn the association at all. But at intervals of 1–3 seconds, learning occurs, but the memory is qualitatively different: it is more context-dependent and less automatic.

The replication team found that this qualitative difference translated into quantitative reversals. For example, one measure—freezing to the tone in a novel context—was higher under delay conditioning in 8 of 10 experiments, but higher under trace conditioning in the remaining 2. The direction of the effect depended on which memory system dominated, which in turn depended on the shock timing.

Effect Sizes Across 20 Experiments

The replication team conducted 20 separate experiments comparing the two protocols, each with sample sizes of roughly 12–16 mice per group. They reported effect sizes (Cohen's d) for each comparison. The median effect size under delay conditioning was 0.42, considered moderate. Under trace conditioning, the median dropped to 0.15, a small effect. More strikingly, the sign of the effect reversed in 11 cases: what had been a positive effect became negative, or vice versa.

Six of the 20 experiments showed non-significant reversals—the direction flipped but the confidence intervals overlapped zero. Five experiments showed significant reversals, meaning the effect in one protocol was significantly positive and the other significantly negative. Only 4 of the 20 experiments showed the same direction and significance under both protocols. A statistical model including protocol type as a factor explained 34% of the variance in effect sizes, far more than any other variable measured.

These numbers are striking because they come from the same lab, same equipment, same mouse strain, and same experimenter. The only difference was the timing of the shock. If such a small change can flip results, then comparing studies across labs—where many procedural details differ—becomes extremely difficult.

What Replication Audits Miss

Large-scale replication projects, such as the ManyLabs consortia in psychology and the Reproducibility Project in neuroscience, typically attempt to repeat published methods as closely as possible. But if the original study used one protocol and the replication team uses a slightly different one, the failure to replicate may reflect procedural sensitivity rather than a false positive.

A recent audit of 50 fear-conditioning papers published between 2015 and 2020 found that at least 8 distinct protocols were in use, varying in shock intensity, tone frequency, inter-stimulus interval, and chamber design. Only 12 of the 50 papers reported the exact timing of the tone-shock interval to the nearest millisecond. The rest gave ranges or omitted the detail entirely. Meta-analyses that pool these studies implicitly assume that the procedural differences are negligible—an assumption the replication team's data directly challenge.

Registered reports, which are meant to reduce publication bias, rarely require authors to specify protocol minutiae in advance. As a result, even preregistered studies may harbor hidden flexibility. The fear conditioning case suggests that replication audits should include a systematic variation of key parameters, not just a single exact copy.

How to Stabilize a Fragile Finding

One obvious remedy is to pre-register the exact shock-tone interval, along with other timing parameters, and to justify why that interval was chosen. But pre-registration alone does not solve the problem, because the choice of interval may still be arbitrary. A more robust approach is to run the experiment under two or more protocol variants and report the results separately, along with an interaction analysis.

Another recommendation is to report effect sizes for each protocol variant, not just p-values. In the replication team's data, the p-values varied wildly across protocols, but the effect sizes told a clearer story: the delay protocol produced larger, more consistent effects, while the trace protocol produced smaller, more variable ones. If researchers always report both, readers can assess the robustness of the finding.

Sharing raw trial-level data would allow other teams to reanalyze the results using different inclusion criteria or statistical models. For instance, some labs exclude mice that freeze less than a certain threshold; others do not. Trial-level data make it possible to test whether such decisions affect the conclusions. A Bayesian approach that incorporates prior knowledge about protocol effects could also help, by shrinking estimates toward a pooled mean when the protocol is uncertain.

Lessons for Neuroscience More Broadly

Fear conditioning is not the only paradigm where procedural details matter. Similar sensitivity has been documented in the water maze, where the size of the platform and the water temperature can alter effect sizes; in social defeat stress, where the duration of exposure to the aggressor changes the behavioral profile; and in conditioned place preference, where the length of the conditioning session can reverse the direction of preference.

These examples suggest that procedural drift—the gradual, often undocumented evolution of methods within and across labs—undermines cumulative progress. Each lab optimizes its protocol for local conditions: the strain of mice available, the equipment on hand, the habits of the technician. Over time, these local optima diverge, and the literature becomes a patchwork of non-comparable results.

Trade-offs in Protocol Standardization

One might argue that the solution is to standardize protocols across all labs. However, standardization comes with its own costs. If every lab uses the same delay conditioning protocol, then discoveries about hippocampal contributions to fear memory—which require trace conditioning—might be missed. Moreover, a protocol that works well for one mouse strain may be suboptimal for another. For example, C57BL/6 mice learn trace conditioning readily, but 129S1 mice show weaker learning with longer trace intervals. Forcing a single protocol could bias results toward certain strains and away from others, reducing generalizability.

Another trade-off is that protocol variation can be a source of discovery. The fact that trace conditioning engages the hippocampus was uncovered precisely because researchers varied the inter-stimulus interval. If all labs had used the same delay protocol, that insight might have emerged much later. Thus, some degree of variation is healthy for the field, as long as it is documented and systematically explored.

A balanced approach might involve a set of core protocols that are used by many labs, supplemented by exploratory variations. For instance, the field could agree on a "standard" delay conditioning protocol (tone co-terminating with shock) and a "standard" trace conditioning protocol (2-second gap), and require that any new study using fear conditioning include at least one of these as a reference condition. This would allow cross-study comparisons while still permitting innovation.

Counter-Arguments and Limitations

Some researchers might argue that the reversal observed in the replication team's data is an artifact of insufficient sample size or random variation. However, the consistency across 20 experiments—with the same direction of reversal in 11 out of 20—makes a strong case that the protocol difference is the causal factor. Moreover, the effect sizes under delay conditioning were consistently larger, which aligns with the known neurobiology: delay conditioning relies on a more direct amygdala pathway, while trace conditioning engages a more distributed network that may produce more variable responses.

Another counter-argument is that the 2-second gap is too short to be considered true trace conditioning; some researchers reserve that term for intervals of 5 seconds or more. However, even a 1-second gap has been shown to recruit hippocampal activity in rodents, so the distinction is not purely categorical. The replication team's data suggest that even a short gap can shift the neural substrates of learning.

It is also worth noting that the original 2008 study used a different mouse strain (C57BL/6) than some of the replication experiments (which also used C57BL/6, but from a different supplier). Subtle genetic differences between substrains could interact with protocol effects. The replication team controlled for this by using the same supplier for all experiments, but the possibility remains that the results might differ in other strains.

Finally, the replication team's findings are based on a single lab. While the internal consistency is high, the generalizability to other labs using the same protocols is unknown. A multi-lab replication that systematically varies the protocol would be the next logical step.

Systematic Variation as a Path Forward

Systematic variation, not exact replication, may be the path forward. By deliberately varying a key parameter—like the shock-tone interval—and observing how the effect changes, researchers can map the boundary conditions of a phenomenon. That map is more valuable than any single point estimate. The fear conditioning case shows that a single procedural detail, once considered too minor to report, can reverse a finding. The lesson is that in neuroscience, the devil is not just in the details—the devil is the details.

To make this approach practical, journals could encourage "parameter sweep" studies that report results across a range of values for a critical variable. For example, a fear conditioning study could include experiments with 0, 1, 2, and 5-second trace intervals, and report how the effect size changes. Such data would not only strengthen the original finding but also provide a template for future replication attempts. Funding agencies could support this by recognizing that systematic variation, while more resource-intensive, produces more robust knowledge.

In summary, the reversal of 11 out of 20 memory effects due to a 2-second change in shock timing is a wake-up call. It demonstrates that procedural details are not minor footnotes but central determinants of experimental outcomes. By embracing systematic variation, transparent reporting, and multi-protocol designs, the field can build a more reliable foundation for understanding the neural basis of memory.

Related Articles