How Do Scientific Explanations Go Wrong?
Scientific explanations can go wrong in different ways. They may compare accounts that answer different questions, claim causation from insufficient evidence, overstate model validation, generalise beyond an established domain, conceal exploratory flexibility, substitute narrative coherence for evidential discrimination, or present integration as demonstrated explanatory advance. This article distinguishes these failures and shows why correction should be calibrated to the particular defect rather than treated as wholesale scientific rejection.
Key Points
- Explanatory failure has distinguishable forms and should not be reduced to a single general verdict.
- Explanations can be miscompared when they address different targets, contrasts or grains.
- Association, plausibility and model fit do not by themselves establish stronger causal or representational claims.
- Local evidence does not determine its own wider scope, and exploratory discovery should not be presented as independent confirmation.
- A model, concept or framework may remain scientifically useful after an overextended explanatory claim is narrowed.
Where This Article Fits
The preceding article identified the positive, target-sensitive dimensions by which explanatory success can be assessed. This article addresses the corresponding diagnostic question: how can an explanatory claim go wrong? It distinguishes failures such as target mismatch, weak evidential discrimination, causal overreach, model-validation inflation, scope inflation, hidden flexibility, narrative excess and overclaimed integration.
The aim is not to reduce every defect to total explanatory failure. A model, hypothesis or framework may remain scientifically useful after a stronger claim has been withdrawn or narrowed. The next article develops the comparative procedure needed when explanations genuinely address sufficiently matched targets.
Introduction
Scientific explanations can fail in different ways. An account may answer the wrong question, claim more than its evidence supports, extend beyond the conditions in which it was established, or present a useful model or synthesis as though it had proved more than it has.
These are not interchangeable defects. Each requires a diagnosis matched to the claim being made. Nor does identifying an overclaim necessarily invalidate the underlying evidence, model or scientific contribution. Often the appropriate response is to narrow the claim until it matches what the evidence warrants.
Explanatory Failure Is Not One Thing
An explanation is not simply successful or unsuccessful in every respect. It may accurately identify one dependency while overstating its scope. It may predict reliably without uniquely representing the system that produces the prediction. It may organise existing findings without yet demonstrating a new explanatory contribution.
This means that criticism should identify exactly what has failed. Has the account answered a different question from the one it was supposed to answer? Does the evidence support association but not causation? Has a model been tested only within a restricted domain? Was an exploratory finding presented as though it had survived an independent test?
The diagnosis matters because different failures require different corrections. Treating them all as instances of “bad explanation” obscures both the defect and what remains scientifically valuable.
A Useful Idea Can Still Be Overclaimed
Scientific usefulness and evidential entitlement are different. A model, concept or framework may remain valuable after a stronger causal, generality, validation or novelty claim is withdrawn—but its usefulness does not establish that stronger claim.
The First Failure: Comparing Different Questions
Two explanations may concern the same phenomenon without explaining the same aspect of it. One might identify a molecular mechanism, another reconstruct an evolutionary history, and a third describe a mathematical dependency. Their shared subject matter does not by itself make them direct competitors.
A winner–loser comparison becomes misleading when the accounts differ in explanatory target, contrast, grain or intended use. Before one account is said to defeat another, it must be clear that they address sufficiently matched questions. The plurality of explanation in biology makes this especially important because causal, mechanistic, historical, mathematical and organisational accounts can illuminate different features of the same system (Ross, 2025).
Target mismatch becomes more serious when a claimed advance depends on weakening or misdescribing the alternative. An account may appear superior only because its competitor has been required to answer a question it was not designed to answer. The appropriate diagnosis may therefore be neither victory nor failure, but target difference or complementarity.
PA-04 identifies malformed comparison as a source of explanatory overclaim. The complete procedure for comparing sufficiently matched explanations belongs to the next article.
When the Evidence Cannot Distinguish the Claim
An explanatory claim requires evidence capable of distinguishing it from relevant alternatives. Evidence may be reliable and scientifically important while remaining insufficient for the stronger interpretation placed upon it.
This problem is especially clear when an association is treated as though it established a causal relationship. Observational evidence can reveal patterns, identify risk factors and motivate causal hypotheses. It does not automatically distinguish a treatment effect from confounding, selection effects, reverse causation or differences between the populations being compared.
This problem is illustrated by research on postmenopausal hormone therapy and coronary heart disease. Earlier observational findings had suggested reduced coronary risk among hormone users. The Women’s Health Initiative randomized trial instead found increased coronary risk among women assigned to the estrogen-plus-progestin regimen and concluded that the regimen should not be used for primary prevention of coronary heart disease (Rossouw et al., 2002).
The discrepancy did not show that observational research was scientifically worthless. A subsequent reanalysis of the Nurses’ Health Study showed that estimates became substantially more consistent with the randomized trial when the observational study was analysed more like a sequence of treatment-initiation trials. Differences in time since menopause and length of follow-up accounted for much of the earlier disagreement (Hernán et al., 2008). The methodological lesson is therefore calibrated: causal interpretation depends on whether study design and analysis distinguish the proposed treatment effect from relevant alternative explanations.
Causal overreach can also arise in mechanistic explanation. A pathway may be biologically plausible and consistent with an observed outcome without having been shown to produce that outcome under the conditions in question. Compatibility is weaker than causal discrimination. Interventions, explicit causal models, natural experiments and sensitivity analyses can strengthen a causal claim when they are appropriate to the system and question (Pearl, 2009; Woodward, 2003).
The failure therefore lies in the distance between the evidence and the claim, not necessarily in the evidence itself.
When Models Are Asked to Prove Too Much
Models are indispensable because they select, idealise and simplify. Their scientific value does not depend on reproducing every feature of the systems they represent. Difficulty arises when successful performance for a specified purpose is inflated into proof that the model is uniquely true or completely validated.
Oreskes, Shrader-Frechette and Belitz (1994) drew attention to this problem in numerical models of open natural systems. Such systems cannot ordinarily be closed to all relevant influences, and multiple model structures may remain compatible with the available observations. Agreement between a model and observed data can therefore support confidence in a bounded application without establishing that the model is a complete or uniquely correct representation.
The appropriate language should identify what has actually been achieved. A model may be well confirmed for an intended use, predictively successful within a defined domain, robust across specified assumptions, or informative about a particular dependency. These are substantial scientific accomplishments. They should not be converted into unrestricted validation.
Model-validation inflation is corrected by narrowing the claim to the model’s demonstrated performance, domain and purpose. This preserves the model’s scientific value while avoiding an inference that the evidence cannot sustain.
Scope Inflation
Evidence is always obtained under particular conditions. Participants, organisms, environments, measurements, interventions and historical circumstances delimit what has actually been examined. Scope inflation occurs when a conclusion is extended beyond those conditions without testing or defending the bridge from the studied case to the broader claim.
Research based disproportionately on Western, educated, industrialised, rich and democratic populations showed how conclusions drawn from unusually restricted samples could be presented as claims about human psychology generally (Henrich, Heine, & Norenzayan, 2010). The problem was not merely sample size. It concerned whether variation relevant to the general claim had been adequately represented.
The same structure occurs elsewhere. A result established in one species, laboratory environment, developmental stage or measurement regime may not automatically generalise to other domains. Even a highly reproducible local effect does not determine its own range of application.
Generalisability cannot be inferred solely from confidence in the original result. It requires evidence or argument connecting the local finding to the wider population or domain. Broad claims made without that bridge exemplify scope inflation (Yarkoni, 2022).
Correcting the problem need not require abandoning the finding. The justified conclusion may simply be narrower than the one initially announced.
Hidden Flexibility
Scientific inquiry often involves exploration. Researchers examine data, notice patterns, revise hypotheses and develop new questions. This is an important source of discovery. The problem arises when an analysis developed after inspecting the results is presented as though it had been specified in advance and subjected to an independent test.
Kerr (1998) called one form of this practice HARKing: hypothesising after the results are known. The defect is not retrospective interpretation itself. It is the concealment of the distinction between generating a hypothesis and testing it.
When many analytical choices are available, a result may appear stronger than it is if the unsuccessful alternatives remain undisclosed. Flexible stopping rules, outcome selection, subgroup choice, model specification and reporting decisions can all affect the apparent evidential force of a finding. If these choices are hidden, readers cannot tell how severely the hypothesis was tested.
Concerns about undisclosed flexibility form part of the wider reproducibility problem addressed by Munafò et al. (2017). Transparency about exploratory status, disclosure of analytical decisions, independent data and preregistration where appropriate can help distinguish discovery from confirmation.
Preregistration is not a universal remedy, and confirmatory inquiry is not the only scientifically legitimate form of investigation. The narrower requirement is that the evidential status of the result should be represented accurately.
Discovery Is Not Confirmation
Exploration after seeing the data can generate valuable hypotheses. The defect is not post-hoc inquiry itself, but presenting an exploratory result as though it had already survived an appropriately independent test.
Narrative Coherence Without Sufficient Discrimination
An explanation can be coherent, plausible and attractive without being well distinguished from alternatives. Narrative coherence helps scientists organise evidence and generate hypotheses, but it cannot substitute for evidence that bears differentially on competing possibilities.
Evolutionary explanation is particularly susceptible to this problem because plausible adaptive benefits can often be proposed retrospectively. Gould and Lewontin’s critique of adaptationism warned against treating every trait as though a plausible adaptive story were sufficient evidence that the trait had been selected for that function (Gould & Lewontin, 1979).
Their criticism does not imply that adaptationist explanation is inherently defective. Adaptive hypotheses can be powerful when supported by comparative, historical, developmental, ecological or experimental evidence. The failure occurs when plausibility is allowed to perform the work of discrimination.
The lesson extends beyond evolutionary biology. Mechanistic stories, historical reconstructions and functional narratives can all become overconfident when their coherence is mistaken for evidential support. A good story may identify a serious possibility; it does not establish that possibility merely by making the outcome intelligible.
When Integration Is Overclaimed
Scientific understanding often advances by bringing previously separated findings into a common framework. Integration can clarify relationships, expose unresolved tensions, coordinate research and improve communication. These are genuine achievements.
But assembling known factors does not automatically produce a new explanation. A synthesis may organise what is already understood without identifying an additional dependency, resolving an evidential conflict or supporting a consequence unavailable from the component accounts. Likewise, new terminology may improve conceptual clarity without establishing a new explanatory result.
The diagnostic error is integration inflation: treating unification, organisation or redescription as sufficient evidence of explanatory advance. The correction is not to dismiss integration but to state precisely what it contributes.
Whether an integrated framework provides explanatory gain beyond the strongest relevant existing accounts is a separate positive question. PA-06 will establish the evidential burden for making that claim. PA-04 establishes only that integration and novelty cannot be assumed to demonstrate it.
Diagnostic Map of Explanatory Overclaim. Scientific explanations can exceed their warrant in different ways. The appropriate response depends on the particular mismatch between target, claim and evidence; identifying an overclaim does not automatically eliminate the contribution that remains supported.
Failure Should Be Calibrated
The most important discipline is to match the correction to the defect. A causal claim may need to become an associational claim. A universal conclusion may need to become population-specific. A model described as validated may instead be well confirmed for a particular use. An apparent confirmation may need to be reclassified as exploratory evidence. An integrated framework may retain organisational value even if its stronger novelty claim is withdrawn.
This calibration prevents two opposite errors. The first is overclaim: presenting evidence as warranting more than it does. The second is over-rejection: treating the failure of a strong interpretation as though it erased every weaker contribution.
Good scientific criticism identifies both what must be withdrawn and what remains warranted. Explanatory evaluation is improved when the claim is brought into proportion with its evidential support.
What This Article Establishes
Explanatory overclaim has distinguishable forms involving different relationships among target, claim and evidence. These include target mismatch, weak evidential discrimination, causal overreach, model-validation inflation, scope inflation, hidden flexibility, narrative excess and overclaimed integration. Corrective controls should be matched to the diagnosed defect rather than collapsed into a single general verdict.
What This Article Does Not Establish
This article does not claim that the identified failure modes form an exhaustive taxonomy or that they are equally common or consequential. It does not imply that observational evidence cannot contribute to causal inquiry, that simplified models must uniquely represent reality to be scientifically valuable, that narrow findings are defective merely because their scope is limited, or that exploratory investigation is inherently illegitimate.
Nor does the article reject adaptationist explanation or scientific integration as such. It does not provide the complete procedure for comparing explanations or the positive test for explanatory gain. Those tasks remain reserved for later articles.
Related Terms and Next Step
Related glossary terms include biological explanation, explanation, explanandum, explanatory target and mechanism.
Next, How Should Scientific Explanations Be Compared? develops a disciplined procedure for determining when explanations address sufficiently matched targets and can therefore be compared without manufacturing rivalry or explanatory victory.
See Also
Related Articles
References
- (1979). The Spandrels of San Marco and the Panglossian Paradigm: A Critique of the Adaptationist Programme. Proceedings of the Royal Society of London. Series B. Biological Sciences, 205(1161), 581–598 . https://doi.org/10.1098/rspb.1979.0086
- (2010). The Weirdest People in the World?. Behavioral and Brain Sciences, 33(2–3), 61–83 . https://doi.org/10.1017/S0140525X0999152X
- (2008). Observational Studies Analyzed Like Randomized Experiments: An Application to Postmenopausal Hormone Therapy and Coronary Heart Disease. Epidemiology, 19(6), 766–779 . https://doi.org/10.1097/EDE.0b013e3181875e61
- (1998). HARKing: Hypothesizing After the Results Are Known. Personality and Social Psychology Review, 2(3), 196–217 . https://doi.org/10.1207/S15327957PSPR0203_4
- (2017). A Manifesto for Reproducible Science. Nature Human Behaviour, 1, 0021 . https://doi.org/10.1038/s41562-016-0021
- (1994). Verification, Validation, and Confirmation of Numerical Models in the Earth Sciences. Science, 263(5147), 641–646 . https://doi.org/10.1126/science.263.5147.641
- (2009). Causality: Models, Reasoning, and Inference. Cambridge University Press. https://doi.org/10.1017/CBO9780511803161
- (2025). Explanation in Biology. Cambridge University Press. https://doi.org/10.1017/9781009300940
- (2002). Risks and Benefits of Estrogen Plus Progestin in Healthy Postmenopausal Women: Principal Results From the Women’s Health Initiative Randomized Controlled Trial. JAMA, 288(3), 321–333 . https://doi.org/10.1001/jama.288.3.321
- (2003). Making Things Happen: A Theory of Causal Explanation. Oxford University Press. https://doi.org/10.1093/0195155270.001.0001
- (2022). The Generalizability Crisis. Behavioral and Brain Sciences, 45, e1 . https://doi.org/10.1017/S0140525X20001685