AI behaviour
Can we recognise a developing failure before a prohibited action?
Study how a tool-using AI system's earlier observations and actions change what happens next.
AI behaviourOverdog
Independent research and engineering across AI behaviour, statistics and mathematics.
AI can generate candidates faster than we can check them. A convincing answer, a statistical estimate and a mathematical pattern each need a different kind of evidence. Our programme, ORBIT, brings the three subjects together to investigate how discovery and careful checking can advance together.

An illustration of measurement and interpretation.
Three research subjects
Each direction produces useful work in its own right. ORBIT stands for Overdog Research in Behaviour, Inference and Theorems.
AI behaviour
Study how a tool-using AI system's earlier observations and actions change what happens next.
AI behaviourStatistics
Connect a study design to an appropriate method and an honest uncertainty statement.
StatisticsMathematics
Use exact computation and proof to find structure and establish what follows from it.
MathematicsRosetta Stone hypothesis
One route starts from exact mathematical objects; another starts from observed AI interactions. We ask whether representations of state, change and constraint can connect them. A useful general translation has not been established.
What a correspondence would have to preserveA candidate mathematical description of state, change, constraint and control that both routes could reach. Whether one exists, and what it would preserve, is the open question.
Concrete work
Choose a subject, work through an example, then inspect its method and sources.
In a simulated refund task, an untrusted message claims that approval has already been given. What changes in the agent's actions, the actual refund and legitimate task completion?
Pre-deployment R&D. The example is a design; prospective warning and intervention performance remain research questions.
Explore the refund experimentInspect current evidenceFour counts can require different uncertainty calculations depending on how the data were collected. Explore a 2 × 2 table with a zero cell and the method choices it raises.
The package distinguishes exact conditional, Mid-P, score and asymptotic methods. Fixed-sample validity does not imply sequential validity.
Explore the table and packageInspect methods and sourceCan three consecutive positive integers all be powerful? Paper I excludes indices divisible by 17 or 41 from a shifted-square equation in its stated Lucas-sequence family.
The recurrence theorem has explicit parameter and index hypotheses. Erdős 364 remains open.
The problem and current resultCharacterise, compare and test competing explanations, with a clear record of what would change the conclusion. Discover, test, then promote or reject.
Overdog is led by Chey Loveday, a statistical geneticist and scientific-software researcher. Discuss the work.