Overdog

Independent research and engineering across AI behaviour, statistics and mathematics.

Finding the mathematics between proof and behaviour.

How can discovery scale without lowering the standard required to accept a result?

AI can generate candidates faster than we can check them. A convincing answer, a statistical estimate and a mathematical pattern each need a different kind of evidence. Our programme, ORBIT, brings the three subjects together to investigate how discovery and careful checking can advance together.

An illustrated sighthound engineer studying an instrument.
Conceptual illustrationThe observer

An illustration of measurement and interpretation.

Three research subjects

Choose the question that interests you.

Each direction produces useful work in its own right. ORBIT stands for Overdog Research in Behaviour, Inference and Theorems.

AI behaviour

Can we recognise a developing failure before a prohibited action?

Study how a tool-using AI system's earlier observations and actions change what happens next.

AI behaviour

Statistics

What do the observations support, and how uncertain is that conclusion?

Connect a study design to an appropriate method and an honest uncertainty statement.

Statistics

Mathematics

Can a better representation make a difficult problem tractable?

Use exact computation and proof to find structure and establish what follows from it.

Mathematics

Rosetta Stone hypothesis

Could a useful mathematical connection exist?

One route starts from exact mathematical objects; another starts from observed AI interactions. We ask whether representations of state, change and constraint can connect them. A useful general translation has not been established.

What a correspondence would have to preserve
From mathematicsExact problems
From behaviourObserved AI interactions
Hypothesis, not a resultRosetta Stone hypothesis

A candidate mathematical description of state, change, constraint and control that both routes could reach. Whether one exists, and what it would preserve, is the open question.

Statisticstests predictions and quantifies uncertainty
Formal methodscheck specified properties and transfer conditions

Concrete work

Start with an experiment, a package or a result.

Choose a subject, work through an example, then inspect its method and sources.

Illustrative experiment

AI behaviour: CARF

In a simulated refund task, an untrusted message claims that approval has already been given. What changes in the agent's actions, the actual refund and legitimate task completion?

Pre-deployment R&D. The example is a design; prospective warning and intervention performance remain research questions.

Explore the refund experimentInspect current evidence
Public software

Statistics: ExactCIs

Four counts can require different uncertainty calculations depending on how the data were collected. Explore a 2 × 2 table with a zero cell and the method choices it raises.

The package distinguishes exact conditional, Mid-P, score and asymptotic methods. Fixed-sample validity does not imply sequential validity.

Explore the table and packageInspect methods and source

How we investigate

Characterise, compare and test competing explanations, with a clear record of what would change the conclusion. Discover, test, then promote or reject.

Sources and people

Overdog is led by Chey Loveday, a statistical geneticist and scientific-software researcher. Discuss the work.