Where This Article Fits

The preceding article, A Practical Guide to Building Better Scientific Explanations, showed how an individual explanation can be constructed and assessed so that its explanandum, evidence, comparator, claimed contribution and scope remain appropriately aligned.

This article addresses a different question.

A scientific framework may contain many kinds of claim: conceptual, methodological, explanatory, modelling, empirical, predictive, integrative, programmatic and claims about scope or generality. The scientific standing of the framework cannot therefore be inferred simply by asking whether one local explanation succeeds.

Earlier articles established the controls that PA-10 now inherits. How Should Scientific Explanations Be Compared? required sufficiently matched explanatory targets and strong relevant alternatives. What Counts as Explanatory Gain? established that gain is bounded, target-relative and comparator-relative. Which Explanatory Framework Is Best? showed that controlled comparison need not produce a global winner.

This article brings those controls to framework-scale assessment. It does not evaluate APS or any other particular scientific framework, and nothing in the procedure depends on APS substantive biology.

Introduction

A new scientific framework should be given its strongest fair test—and a real opportunity to fail.

That requirement is more demanding than either enthusiasm for novelty or reflexive defence of established approaches. A framework should be reconstructed strongly enough to receive full credit for what it actually proposes, but its claims should also remain exposed to evidence capable of supporting, limiting or counting against them.

A Framework Must Be Testable Without Being Tested Unfairly

Framework assessment should avoid two opposite errors.

The first is premature endorsement. A framework can appear persuasive when its terminology is novel, its conceptual architecture is elegant, its formalism is sophisticated or its motivating examples are favourable. None of those features, by itself, establishes the broader scientific standing of the framework.

The second error is premature dismissal. New frameworks should not be tested only against weak formulations, caricatured claims or standards that established approaches themselves are not required to meet.

The appropriate standard is demanding in both directions:

Give the framework its strongest defensible formulation, but preserve genuine evidential risk.

Scientific appraisal has long involved several considerations rather than one mechanical measure. Kuhn emphasised that theory choice draws on values such as accuracy, consistency, scope, simplicity and fruitfulness, while Laudan analysed scientific progress through changing empirical and conceptual problem-solving performance (Kuhn, 1977; Laudan, 1977).

PA-10 does not turn those traditions into a universal scoring system. It uses the more limited conclusion that framework assessment should remain sensitive to what kind of scientific claim is actually being made and what evidence would bear on it.

1. Separate the Framework’s Claims

A framework is rarely one claim.

It may introduce terminology, reorganise existing concepts, propose methodological rules, supply new models, make substantive empirical claims, generate predictions, integrate previously separate phenomena or claim broad applicability across different explanatory targets.

These achievements should not be bundled into one undifferentiated verdict.

Conceptual clarification may be scientifically valuable without establishing an empirical hypothesis. A useful methodology may guide inquiry without supplying additional explanatory capacity on every target. A predictive model may perform well without establishing every interpretive claim associated with the framework. An integrative synthesis may organise evidence effectively without showing that it explains something unavailable to the strongest existing account.

A practical assessment may therefore begin with a claim map or claim register. Such a device is optional rather than universal. Its purpose is simply to prevent different kinds of claim from borrowing evidential credit from one another.

The first question is:

What claim is actually seeking scientific standing?

2. Match the Evidential Burden to the Claim

Different claims require different forms of support.

A conceptual claim may be assessed for coherence, clarity, discrimination and relation to established usage. A methodological claim may be assessed by whether it improves inquiry, comparison, measurement or reasoning. A modelling claim may require formal adequacy, empirical fit, robustness or successful application to a defined problem. A substantive empirical claim requires evidence bearing directly on the scientific proposition at issue.

These achievements can coexist, but they are not interchangeable.

Conceptual coherence does not establish empirical validation. Formal tractability does not establish explanatory superiority. Integration does not automatically demonstrate additional explanatory capacity. Usefulness does not become empirical confirmation merely because it improves scientific practice.

This distinction is especially important in model-based science. Oreskes, Shrader-Frechette and Belitz argued that verification and validation do not amount to establishing the truth of open natural-system models in an unrestricted sense (Oreskes, Shrader-Frechette, & Belitz, 1994). The broader lesson for framework assessment is that evidence should support the kind of claim actually made, not a stronger claim silently inferred from it.

The second question is therefore:

What evidence would genuinely bear on this kind of claim?

3. Specify the Target and the Claimed Contribution

Where a framework claims explanatory advance, the relevant explanandum should be specified before superiority is adjudicated.

The framework should state what it proposes to explain and what additional capacity is claimed. That contribution might involve a mechanism, dependency, developmental route, discriminating prediction, counterfactual relation, mathematical constraint or some other target-relevant consequence.

The governing question remains:

What does the candidate explanation enable us to explain, discriminate, constrain or infer that the strongest relevant comparator does not already enable us to do?

This burden applies when explanatory gain is claimed. It is not a universal demand that every legitimate scientific contribution must take the form of explanatory gain.

Where feasible, material targets, expected contributions, relevant comparators and possible adverse outcomes should be specified before the result used for adjudication is known. Transparency and prospective specification can reduce retrospective fitting and strengthen reproducibility (Munafò et al., 2017).

Formal preregistration may be valuable in appropriate research designs, but this article does not treat it as a universal requirement for framework assessment.

4. Use the Strongest Relevant Comparator

Comparative claims are only as informative as the alternative against which they are tested.

A framework should not earn superiority by defeating a historically weak version of a rival, a simplified textbook description or an account stripped of resources actually used in contemporary scientific practice.

Where comparative superiority is claimed, the comparator should be selected from the explanatory target outward.

Lipton’s work on inference to the best explanation emphasises explicit comparison among explanatory alternatives, while contemporary philosophy of biological explanation reinforces the importance of recognising the diversity of explanatory targets and resources available in biology (Lipton, 2004; Ross, 2025).

The practical implication is straightforward:

Reconstruct the strongest relevant account actually available for the target under assessment.

Sometimes that comparator will be a single account. Sometimes established scientific practice legitimately combines several resources. A composite comparator is appropriate when those resources are actually integrated in the science being assessed.

But the opposite error must also be avoided. Combining every imaginable explanatory resource into an artificial super-comparator would make additional contribution impossible by construction.

The comparator should be strong enough to make the test informative, but real enough to represent actual scientific practice.

5. Make Adverse and Null Outcomes Possible

A framework cannot be tested fairly if every possible result is interpreted as support.

Motivating examples are often useful for developing a framework, but broader scientific standing should not depend only on cases selected because they already appear favourable.

Where the claim permits it, assessment should therefore include genuine possibilities of outcomes such as:

  • no additional explanatory gain;
  • failure of a predicted difference;
  • comparator advantage;
  • unsuccessful transfer to a new target;
  • unresolved assessment;
  • non-comparability;
  • scope restriction;
  • evidence that counts positively against the specified claim.

An adverse case is not merely a difficult case. It is a case capable of bearing against a material claim.

Similarly, a null result is not automatically a global refutation. It may show only that additional gain was not demonstrated on one target, that the available evidence remains insufficient, or that the original scope claim was too broad.

The essential condition is simpler:

The testing design must permit outcomes that do not count as framework support.

A Framework Must Be Allowed to Lose Locally

A fair test must permit outcomes that do not count as support. Failure to demonstrate additional explanatory gain on one target does not by itself refute an entire framework. Insufficient evidence is not positive refutation, and a failed transfer may narrow a scope claim while leaving a local result intact. But evidence may also genuinely count against a specified framework claim. The important discipline is to preserve these distinctions rather than converting every outcome into either framework-wide success or framework-wide failure.

6. Test Transfer Only Where Broader Scope Is Claimed

Not every scientific framework needs to apply everywhere.

A framework may make a legitimate bounded contribution to one class of problems without claiming universality. Breadth is not itself explanatory quality.

The burden changes when the framework claims more.

If it claims broad applicability, generality, unification or framework-wide standing across multiple targets, the evidence should extend beyond the case on which the initial success was demonstrated.

The governing principle is:

The breadth of the evidential burden should track the breadth of the framework claim.

A transfer test may therefore ask:

  • Does the relevant explanatory or methodological capacity survive on a sufficiently different target?
  • Is the same framework commitment doing the work?
  • What remains stable across applications?
  • Has substantial target-specific machinery been added?
  • What evidence justifies moving from local success to a broader inference?

Problems of generalisation arise when conclusions outrun the populations, measures, situations or targets on which the original evidence was obtained (Yarkoni, 2022). Framework-level claims require the same discipline.

A failed transfer need not erase the original result. It may instead show that the framework’s legitimate scope is narrower than previously claimed.

7. Keep Local Results Local

The scope of the conclusion should not exceed the scope of the evidence.

A local explanatory gain establishes a local explanatory gain.

It does not automatically establish:

  • framework-wide superiority;
  • replacement of established frameworks;
  • universal applicability;
  • superiority on unrelated targets;
  • superiority of every other claim made by the framework.

The reverse inference is equally invalid.

A local null result does not show that the entire framework is false. One unsuccessful transfer does not invalidate every successful application. One unresolved comparison does not demonstrate global scientific failure.

Pluralist work on explanation has long shown that explanatory relations need not reduce to exclusive winner-takes-all competition (Marchionni, 2008; Gijsbers, 2016).

The appropriate question is therefore not:

Did the framework win?

It is:

What standing does this result warrant, for this claim, on this target, with this evidence?

Broader standing requires bridge evidence commensurate with the broader claim.

8. Make Framework Revision Visible

Scientific frameworks change.

That is not a defect. Revision in response to evidence, criticism, new methods or expanded scientific knowledge is part of normal scientific development.

The methodological problem arises when the framework tested is silently replaced by a revised version after an adverse, null or unresolved result.

A transparent assessment should therefore retain answers to four questions:

  1. What version of the framework made the claim?
  2. What exactly was tested?
  3. What changed after the result?
  4. Does the revision preserve, narrow, replace or supersede the original claim?

A revised framework may be stronger than its predecessor. It may also require a new test.

The governing rule is:

Scientific revision is legitimate; silent retrospective substitution is not.

Framework-version control keeps adverse evidence meaningful because the claim that encountered the evidence remains identifiable even when later work improves the framework.

Framework Assessment Need Not Produce One Verdict

Once claims are separated, their scientific standing may diverge.

One part of a framework may be well supported. Another may supply bounded explanatory gain. A methodological component may be useful without establishing an empirical proposition. A scope claim may need narrowing. A model may remain promising but unresolved. A substantive claim may encounter evidence that counts against it.

Possible outcomes therefore include:

  • supported;
  • bounded explanatory gain;
  • useful without demonstrated explanatory gain;
  • no additional gain;
  • unresolved;
  • narrowed;
  • adversely tested.

These outcomes are not stages on a ladder and should not be converted into a numerical score.

A scientifically assessable framework need not receive one global verdict.

Testing a Scientific Framework

The framework-level procedure can now be seen as a sequence of controls rather than an algorithm for producing a winner.

Framework-testing workflow showing claim separation, claim-specific evidential burdens, target and contribution specification, strong relevant comparison where warranted, genuine adverse and null possibilities, transfer testing where broader scope is claimed, bounded conclusions, visible revision, and differentiated framework outcomes.

Testing a Scientific Framework. A framework should be assessed through the distinct claims for which scientific standing is sought. Evidential burdens should match those claims, relevant comparisons should use strong alternatives, adverse and null outcomes must remain possible, and broader claims require evidence commensurate with their scope. The resulting standing may therefore differ across claims rather than resolving into a single framework-wide verdict.

What Makes the Test Fair?

A fair framework test is neither protective nor punitive.

It gives the framework its strongest defensible formulation.

It identifies what kind of claim is actually being assessed.

It asks what evidence would bear on that claim.

Where comparison is warranted, it reconstructs strong relevant alternatives.

It leaves genuine room for adverse, null, unresolved and comparator-favourable outcomes.

It tests broader transfer only when broader scope is claimed.

It prevents local results from being promoted into global conclusions.

And it records material framework revisions rather than allowing the tested claim to disappear retrospectively.

The central symmetry is therefore methodological rather than mechanical.

Comparable claims should face comparable standards of evidential responsibility, but different claim types need not be tested by identical procedures.

A framework receives a genuine opportunity to succeed because its test also leaves specified claims a genuine opportunity not to succeed.

What This Article Establishes

New scientific frameworks can be assessed more rigorously when their distinct claims are separated, each claim is matched to an appropriate evidential burden, relevant comparisons use sufficiently strong alternatives, adverse and null outcomes remain genuinely possible, broader scope is supported by evidence commensurate with that scope, and material framework revisions remain visible.

Such assessment permits differentiated outcomes.

Some framework claims may be supported. Some may provide bounded explanatory gain. Some may be scientifically useful without demonstrating additional explanatory capacity. Others may remain unresolved, require narrowing or encounter evidence that counts against them.

Framework assessment therefore need not terminate in either global endorsement or wholesale rejection.

What This Article Does Not Establish

This article does not provide a universal framework-testing algorithm.

It does not require every new framework to defeat established frameworks globally.

It does not require every framework to demonstrate broad transfer.

It does not require the same form of evidence for conceptual, methodological, explanatory, modelling and empirical claims.

It does not treat novelty, breadth, integration, usefulness, conceptual richness or formal sophistication as substitutes for evidence appropriate to the claim.

It does not infer framework-wide superiority from local explanatory gain.

It does not infer framework-wide rejection from one adverse, null or unresolved case.

And it does not evaluate APS or any other particular scientific framework.

Completing the Methodology and Explanation Sequence

This article completes the ten-article Methodology and Explanation sequence.

Across that sequence, the task has moved from identifying what scientific explanation is and how explanatory practices developed, through criteria of success, failure, comparison and explanatory gain, to contemporary explanatory frameworks, controlled comparative adjudication, practical construction of explanations, and finally framework-level testing.

Completion of this article does not itself establish that the sequence is globally coherent or that any particular scientific framework satisfies the methodology developed here.

Those are separate questions.

See the glossary entries for Biological Explanation, Explanation, Explanandum and Explanatory Target.