Statistics

Fixed-sample and sequential inference.

Statistical inference asks what observations support under a declared data-generating or sampling design. The same numerical table can justify different conclusions under different designs and methods. Sparse counts and zero cells make that distinction especially important.

ExactCIs is the concrete public software package. The wider programme studies finite-sample procedures, sequential evidence and their implementation; it is ongoing research, not a list of capabilities shipped by ExactCIs.

Statistical question

Design, method, claim

What do the observations support, and how uncertain is that conclusion? Start with how the data were generated, not with a preferred formula.

Design

How were the observations generated? Declare the sampling unit, independence or pairing, and when the sample size or analysis time was fixed.

Method

What procedure applies to that design and target? An odds ratio, a risk ratio and a paired comparison are not interchangeable requests.

Claim

What uncertainty statement does that procedure justify? Name its coverage or error property and keep the assumptions beside the conclusion.

Illustrative counts, not a reported study

Three events in twenty episodes, versus none in twenty.

Before observing outcomes, fix two groups of 20 independent episodes under one task and model setup: perturbed and control. Count once per episode whether a predefined prohibited effect occurs. Assume a common event probability within each group, with no matching or carried state between episodes.

Two independent binomial groups
ConditionEffect occurredDid not occur
Perturbed317
Control020

The observed rates are 3/20 and 0/20. Choose the population odds ratio as the target: the event odds in the perturbed group divided by those in the control group, assuming both underlying probabilities lie strictly between zero and one. A zero observed count does not establish zero event probability; it puts the uncorrected sample odds ratio and its usual log-Wald interval at a boundary.

The released independent-binomial methods are score-based or asymptotic, with support depending on the target and method. This table does not acquire an exact conditional guarantee by changing its label. No interval is calculated in this illustration. Read the full example and its method limitations.

Fixed-sample inference is not sequential inference.

A fixed-sample guarantee applies at a prespecified sample size or analysis time under the declared design. Checking an ordinary interval repeatedly and stopping when it looks persuasive is a different procedure, not an extension of that guarantee.

ExactCIs

Confidence intervals for declared designs

ExactCIs provides Python confidence-interval methods for 2 by 2 tables. Exact conditional odds-ratio inference is distinguished from Mid-P, score-based and other asymptotic constructions, not treated as a package-wide status.

  1. Fix a sampling design, target and interval procedure.
  2. Draw a new sample and calculate an interval.
  3. Across repetitions, count intervals containing the fixed target.
Conceptual illustrationRepeated-sampling coverage

Coverage describes how often an interval procedure contains a fixed target across repeated samples under a stated design. This conceptual illustration does not establish a package-wide guarantee.

ExactCIs checks the requested design, target measure and method for compatibility. Unsupported combinations and numerical failures are explicit, not invitations to silently substitute another method. Independent and paired designs are not interchangeable; retaining that distinction is not a claim that a paired procedure is supported.

For example, repeatedly sample tables from one declared design and check how often an interval procedure covers the fixed target value. This is a property of the procedure across repetitions, not a posterior probability that the target lies in one observed interval.

Research branches

Three branches of statistical work

Fixed-sample inference, sequential evidence, and calibration and implementation ask related but distinct questions. A result in one branch does not automatically establish a result in another.

Exact and finite-sample inference

Study confidence procedures whose finite-sample guarantees can be stated without relying on a large-sample approximation. ExactCIs is the public starting point, with exact conditional odds-ratio methods separated from Mid-P and documented score-based or asymptotic alternatives.

Independent-binomial procedures, paired-design distinctions, sparse tables and boundary cases make method selection substantive. Read coverage alongside interval width and refused or unavailable outputs.

Supported designs and constructions

Sequential evidence

Study e-values, e-processes, repeated monitoring and stopping-time questions. Anytime-valid evidence requires a declared null, an information history or filtration, and rules for adaptation that preserve the stated guarantee.

This is a research branch, not a sequential-validity claim for ExactCIs. Programme-specific implementation and formalisation claims require their own statement, source and checked scope.

Definitions, validity and source status

Calibration and implementation

Make the path from definition to software inspectable through numerical reference fixtures, coverage studies, adversarial boundary cases, method-selection discipline and implementation conformance.

Exact arithmetic and certificates, where applicable, must identify the objects and domains they establish. Statistical verification, formal verification and usable tooling have different jobs: a passing fixture is not a theorem, and exact enumeration does not make an approximate interval exact.

Coverage, width and numerical evidence

Definitions and source

Inspect the statistical object and its support

Start directly with a method definition, source record or validity condition. The public package and the wider sequential research line have separate evidence requirements.

Sequential evidence: definitions and validity

Null and filtration
Name the null distributions and the information available at each time. This growing information history is the filtration. Adaptation must respect it; an update cannot use future observations as though they were already known.
Process definition
An e-value is nonnegative evidence with expectation at most one under each distribution in the declared null. An e-process is an adapted nonnegative process with that expectation bound at stopping times, not merely a sequence of pointwise-valid e-values.
Validity theorem
For a nonnegative test supermartingale starting at one, Ville's inequality bounds the null probability of ever reaching 1/α by α, for 0 < α < 1. The supermartingale property must hold under the declared null and filtration.
Stopping-time interpretation
Stopping can depend on information observed so far. Optional-stopping validity concerns that entire rule, not a fixed-time calculation relabelled after inspection. False-alert control does not, by itself, establish detection power or useful warning time.

These are established statistical constructions, not new programme results. The definitions and validity framework are discussed in Ramdas and colleagues, Game-Theoretic Statistics and Safe Anytime-Valid Inference (2023).

Sequential implementation and formalisation: source limits

The literature link explains the statistical framework. This page does not supply a versioned programme-specific process definition, validity proof, implementation or machine-checked formalisation record. It therefore makes no claim that such a programme artefact has been released or verified.

A formalisation claim needs the exact proposition, assumptions, source revision and check result. An exact-arithmetic certificate needs its declared finite domain. Neither can be inferred from the ExactCIs release, numerical reference fixtures or the phrase “anytime-valid”.

Potential applications to CARF

Statistical tools could support experiment-level inference, trajectory-score calibration, sequential warning and false-alert control. Transport to a new setting and drift over time add uncertainty that needs its own assumptions and evidence. These are intended applications, not an implemented integration or a validated warning system.

Evidence geometry and transfer across domains belong at this technical depth: the question is whether a representation preserves the mathematical structure a claim needs. The Rosetta Stone hypothesis is the home for that research question. This work does not supply evidence for a useful general translation.

ExactCIs is not presented as a CARF component, and its public release does not validate CARF. Read the CARF framework and the Statistics direction in ORBIT for their separate scope.