Explain 03

How do controlled experiments help measure AI behaviour?

Short answer

Controlled experiments compare prespecified conditions designed to differ in a primary factor or policy. Measured nuisance variables are matched, blocked or randomised where appropriate, adherence is checked and uncertainty is quantified. The design narrows plausible explanations; it does not prove that every other influence was absent. The rules for pass and fail are set before anyone looks.

matched where measured: task, model, seeds, load
Arm A
without the primary factor
Arm B
with the primary factor
Positive control
a known effect, to be detected at a prespecified rate
Negative control
clean data; false alerts stay within a prespecified bound
pass and fail rules written down first
The contrast estimates the effect of the assigned difference under the stated design; the controls test the method reading it.

Model outputs are stochastic, and evaluation pipelines introduce additional variation. Run the same task twice and something differs; look at enough output and something always looks strange. A single unusual result therefore says little. In a controlled experiment, matched or randomised conditions are designed to differ in one primary factor. Replication, uncertainty estimates, adherence checks and negative controls are then used to assess whether the observed contrast is larger and more consistent than ordinary variation.

The controls answer a different question: is this measurement method sensitive and specific enough for the intended use? Positive controls should be detected at a prespecified rate, and false alerts on negative controls should remain below a prespecified bound. Failure does not prove the method measures nothing, but it does mean the method is not yet reliable enough to interpret unknown cases.

A concrete example

An illustrative design: does accumulated history shift behaviour? Pair the episodes: the same task with and without a long working history behind it. Add a third condition carrying an equally heavy but irrelevant load, so that “the session was simply harder” competes as an explanation and can lose. A difference between the history condition and both controls would support a history-specific effect within that tested setting. It would not by itself establish a universal mechanism, or transport to other models and tasks.

A common misunderstanding

“An unusual output is a finding.” Without controls and a comparison, an unusual output is an anecdote. Ordinary variation and pipeline artefacts imitate signal constantly; the experiment exists to stop them being promoted into claims.

Why it matters

Every CARF measurement proposed for operational use must first be evaluated on known controls, inside a prespecified experimental contrast. Passing those checks supports use only in the tested scope. That is what separates a measurement from an impression, and it is the only kind of evidence Overdog lets a claim rest on.

A pharaoh hound and a robot quadruped facing each other inside an instrumented enclosure, tracking cameras above, one small amber marker on the floor between them
Plate 10 · Matched armsOverdog archive