Joery de Vries — Projects to Supervise
Projects to Supervise
Homeostatic Regulators in Multi-Agent Environments
Supervisors: Joery de Vries, Neil D. Lawrence
Prerequisite: L172 Information, Energy and Intelligence (IEI), or equivalent preparation in information theory, maximum entropy, and information geometry.
Conant and Ashby’s good regulator theorem concerns a regulator that holds an outcome steady against a disturbance. Its success criterion is the entropy of the outcome, which is different from the classical notion of reward in reinforcement learning: it is a concave objective over the occupancy polytope, so an optimal single regulator is deterministic. This project asks what happens when the disturbance is another regulator. Several agents share an environment and each minimises the entropy of its own outcome under its own reference measure. From any agent’s viewpoint the other agents are structured, adaptive disturbances. Refinements of the theorem, notably Wentworth’s, say the regulator must carry a posterior over its disturbance, thus the notion of “model” that Conant and Ashby’s theorem implies is a posterior over the other agents’ policies. We will try to answer whether this posterior is necessary, and whether the joint problem is Nash. The working hypothesis is that competing regulators partition the state space into per-agent stable niches, which remains to be verified experimentally.
What Does a Good Regulator Need to Know?
Supervisors: Joery de Vries, Neil D. Lawrence
Prerequisite: L172 Information, Energy and Intelligence (IEI), or equivalent preparation in information theory, maximum entropy, and information geometry.
Conant and Ashby’s (1970) good regulator theorem says a successful regulator must be a model of its system. Which model depends on what the regulator observes: Wentworth’s (2021) “gooder regulator” for instance requires the belief state. Since a good regulator objective minimises the entropy of a regulated outcome this adds a secondary dependence during learning due to concavity of the optimization problem. Similar to convex RL, it can be solved by a sequence of linear rewards built from the occupancy of the outcome features. Although the agent converges to a single deterministic policy, during learning its representation must support every reward in the sequence. Therefore, reusing what it learned under earlier rewards while staying focused on what the objective makes relevant is crucial. For instance, the successor features of the outcome suffice for this. Despite much work on state abstraction, self-predictive representations and sensorimotor world models, it is unclear what a good regulator needs to represent while it learns. This project investigates what acting and learning require for good regulators in the language of state abstractions of Li, Walsh and Littman (2006) and of Ni et al. (2024), and what combination of latent world-model loss delivers all aspects.