A no-nonsense guide to Perceptual Control Theory — from the feedback loop that explains your thermostat and your brain, to the 11-level hierarchy, to the differential gain collapse that explains why control loops fail when the substrate they act through runs out.
The state the system is specified to reach. Held internally, not supplied by the environment. Denoted \( r \).
What the system currently senses of that variable. Never the world itself, only the reading. Denoted \( p \).
Subtracts one from the other: \( e = r - p \). Error is the whole output.
Action that changes perception, not behaviour selected by a stimulus. The loop closes through the world. Denoted \( q \).
Those four terms are the entire theory. Everything below — the eleven levels, the substrate formulation, the differential gain collapse, and the empirical calibrations — is that loop repeated, stacked, or pushed to the limit of the substrate it acts through. The hierarchy is what happens when the output of one loop sets the reference of the one below it.
William T. Powers did not stumble into Perceptual Control Theory. He engineered it — starting in the 1950s as a physicist and electronics engineer who worked with radar systems and feedback controllers. While behaviorism ruled psychology with its insistence on stimulus-response chains, Powers noticed something the psychologists had missed: thermostats, autopilots, and servo mechanisms didn't react to disturbances — they controlled against them. The loop closed through the environment. Error drove output. Output changed the world. The changed world fed back into the sensor. Stability emerged from the loop itself, not from any programmed response.
In 1960, Powers was the lead author of "A general feedback theory of behavior" — with Robert K. Clark and Rowland L. McFarland supporting the work and enabling its publication — in Perceptual and Motor Skills — the first formal statement of what would become PCT. The core claim was radical: behavior is not the organism's response to stimuli. Behavior is the organism's means of controlling its perceptions despite disturbances from the environment. The stimulus does not cause the response. The organism acts to keep a perception matched to an internal reference, and whatever actions happen to accomplish that are the behavior.
"The organism does not respond to stimuli in its environment. It acts in ways that keep its perceptual signals close to its reference signals, despite the disturbances that the environment may provide."
— William T. Powers, Behavior: The Control of Perception, Aldine, 1973The 1973 book — Behavior: The Control of Perception, published by Aldine — laid out the full theory. It was not an immediate hit. Academic psychology had no framework for circular causation, and the journals built on linear models were not interested in being dismantled. Powers spent the following decades refining the work outside the mainstream, publishing the anthologies "Living Control Systems I & II" in 1989 and 1992, and the more accessible "Making Sense of Behavior" in 1998. He founded the Control Systems Group (CSG) in the 1980s — a serious community of scientists, engineers, therapists, and philosophers who gathered annually to expand PCT. Powers died in 2013. By then, Warren Mansell and Sara Tai at the University of Manchester had taken PCT into clinical psychology, producing peer-reviewed trials that gave the theory its first foothold in mainstream journals. The International Association for Perceptual Control Theory (IAPCT) continues this work today.
The uncomfortable truth about PCT's marginal status is not scientific. The evidence has always been strong — behavioral simulations matching human data with over 95% accuracy, clinical trials with strong retention rates, robotics applications outperforming classical controllers. The problem is paradigmatic. PCT requires abandoning the stimulus-response assumption that underlies not just behaviorism but most of cognitive science and virtually all of reinforcement learning. That is not a small ask for fields that have built entire careers, curricula, and funding structures on the old model.
Think about the most boring but perfect example in the world: a household thermostat. You set it to 22 °C. The sensor quietly measures the room temperature every few seconds. If it drops below 22, the heater fires up. If it climbs too high, the heater cuts off. The thermostat isn't "reacting" to the cold like some scared animal. It's controlling its own perception of temperature to match the number you gave it. Disturbances come — open window, someone turns on the oven — and it keeps acting until the perception is back on target. That's negative feedback in its purest form.
Now scale that up to you driving a car. Wind blasts from the side. The car starts drifting left. You don't think "oh wind stimulus → steering response." You just act so that the road stays centered in your field of view. The controlled perception is "road straight ahead," not "wind from left." Powers nailed it in 1960: the organism doesn't respond to stimulation — it controls its own input. Every muscle twitch, every eye movement, every heartbeat adjustment is part of keeping dozens of perceptions where they should be.
"The organism does not respond to stimulation; it controls its own input."
— William T. Powers, Behavior: The Control of Perception, Aldine, 1973Most psychology still treats behavior like a chain of causes: stimulus → processing → response. PCT says that's describing the shadow on the wall, not the thing casting it. Behavior is the means to an end. The end is always a controlled perception staying stable despite the world trying to knock it off course.
Holding a perception steady against disturbance is only the simplest case — homeostasis, which is control with a constant reference signal. PCT reaches far beyond homeostasis: the reference signal can change as fast as you like, so the same control loop accounts for dance, speech, and any other purposeful movement, slow or at lightning speed. The aim is never stillness for its own sake — it is keeping perception matched to an intention, even when that intention is itself in motion.
When you model that loop mathematically and compare it against real humans doing the same tracking task, performance accuracy goes over 95%. This is not curve-fitting after the fact — it is a direct comparison of the model's performance against a human doing the same task. It falls short of 100% mainly because the human is not as precise as the model, but the two are basically the same. That's why this simple loop has been quietly outperforming stimulus-response models for sixty years.
The critical concept is the controlled variable — the specific perception the organism is actually keeping stable. Not its muscles. Not its outputs. Its perception of something in the world. A person holding a cup controls the perception of the cup's position in their hand, not the tension in any specific muscle group. The muscles are simply whatever happens to achieve perceptual stability. This is why the same goal can be achieved through entirely different physical actions — the output varies, but the controlled perception stays constant.
The TCV Protocol: To determine whether a candidate variable \( v \) is actually controlled by an agent, three conditions must hold. All three are necessary; resistance alone is insufficient, because a reflex arc resists a disturbance without controlling anything.
Powers' original treatment appears in Behavior: The Control of Perception (1973), appendix. The protocol as an experimental method is developed in Marken (2009) and the PCT laboratory literature.
Learning in PCT is reorganization — essentially trial and error at the brain level. When persistent error cannot be eliminated through normal control action, the system randomly varies its own parameters until the error drops. There is no external reward signal. The criterion is intrinsic: sustained error triggers parameter changes until it resolves. This replicates human skill acquisition patterns without invoking reward, punishment, or any form of external feedback beyond the environment itself. The honest gap: the specific neural mechanisms of reorganization remain an active research question. It is biologically plausible. It is not yet fully specified.
See how this applies to AI systemsGrounding definition. Perceptual Control Theory formalizes behavior as negative feedback on a perception: \( p = M \cdot q + d \), where \( M \) is the environmental gain, \( q \) is output effort, \( d \) is additive disturbance, and \( p \) is the controlled perception. The classical form is exact only when the environment has effectively infinite capacity; the substrate formulation \( p_i = F_i(q_1, \dots, q_n, x) \) generalizes it by making the environment a stateful, shared, saturable field.
Powers' tracking experiments were run on a display, not on a factory floor. That matters mathematically. The display is a clean, stateless, effectively infinite-capacity substrate: whatever output \( q \) the participant generates, the cursor moves by \( M \cdot q \), and any external disturbance adds linearly. In that regime, the classical control-theory transfer function is exact:
The real world is not a display. When the loop is closed not through a stateless experiment but through a physical or algorithmic substrate — a power grid, an order book, a hospital bed pool, a shared GPU cluster, a market maker's balance sheet — the environment is a shared, stateful, finite-capacity field. Many agents act into the same substrate. The substrate can run out. And when it does, the linear form above stops being a useful approximation — not because the disturbance got bigger, but because the derivative \( \partial p / \partial q \) itself collapses toward zero.
The correct generalization is the substrate formulation. For agent \( i \) controlling perception \( p_i \) by output \( q_i \), with \( n-1 \) other agents acting into the same field:
The two formulations are not two ways of writing the same thing. They differ in the sign of a partial derivative, and that difference changes the qualitative dynamics of the loop. The classical form assumes the derivative exists, is positive, and is bounded away from zero. The substrate form allows it to vanish. Once it vanishes, no amount of controller gain can recover the loop, because the loop's closure itself has been cut.
Formally, the classical form is obtained as the first-order Taylor expansion of \( F_i \) in \( q_i \) around a working point \( q_i^0 \), holding \( q_{-i} \) and \( x \) fixed. Writing \( M_i \) for \( \partial F_i / \partial q_i \) evaluated at that point and folding the constant term into \( d_i \), we recover \( p_i = M_i \cdot q_i + d_i \) — but only under three explicit conditions:
(1) \(F_i\) is differentiable in \(q_i\) at the working point.
(2) \(M_i\) is approximately constant on the operating range.
(3) The influence of \(q_{-i}\) and \(x\) is representable as an additive term \(d_i\).
This is the extension to Powers' framework, stated plainly. Powers worked in the clean-display regime; his mathematics is exact there. The substrate formulation does not replace it — it names the regime in which the classical form remains valid, and the regime in which it stops being so. When the substrate is unlimited, PCT and the substrate formulation are the same theory. When the substrate is finite, they diverge, and the divergence is what the next section is about.
The differential gain collapsePowers didn't pull 11 levels out of thin air. He reverse-engineered them by building working models that matched real behavior better than anything else. Bottom rung: intensity — raw brightness, loudness, pressure, muscle tension. Next: sensation — colors, tastes, warmth. Then configurations: seeing a cup as round, not just patches of light. Transitions: movement, change over time. Relationships: "above," "beside," "bigger than." And it keeps climbing — events, sequences, programs, principles, all the way to system concepts: who you are, what your life means, your identity.
Think of it like a company that actually works. The CEO (system concepts) decides "this is the kind of organization that values integrity." That sets principles: "don't lie to customers." Principles set programs: "if quality issue, recall immediately." Programs set sequences: "first notify, then refund." Down to workers moving fingers on keyboards. Each level only talks to the one below. The CEO doesn't tell the warehouse worker which button to push — just sets the goal. Same in the brain. Higher levels set what you want. Lower levels figure out how.
Is it perfectly mapped in neuroscience? Not yet. fMRI shows hierarchical processing — layered activity from sensory cortex to prefrontal — but pinning exact 11 levels to specific brain structures? Still work in progress. Warren Mansell and colleagues keep running behavioral experiments that match the model with remarkable accuracy, but gaps remain. Reorganization — how the hierarchy rewires itself when persistent error can't be resolved — is still mostly a black box. Doesn't make the model wrong. Makes it unfinished. And that's fine. Science isn't about having all answers today. What matters is that PCT is the most accurate model of hierarchical behavior control that currently exists.
See how the hierarchy applies to AI systemsGrounding definition. Differential gain collapse is the uniform vanishing of the marginal perceptual return on output effort across the operating range of a controller, \( \sup_{q_i} |\partial F_i / \partial q_i| \to 0 \), as a shared substrate approaches saturation \( \rho \to C \). It is mathematically distinct from additive disturbance: under disturbance the loop still closes because \( \partial p / \partial q > 0 \), while under collapse the loop is decoupled because the derivative vanishes. The controller cannot compensate by increasing \( q \), because increasing \( q \) no longer moves \( p \).
This is the section that does the work. Everything above it is standard PCT plus its formalization. Everything below it is where the theory stops being a description of clean tracking experiments and becomes a falsifiable claim about physical and algorithmic systems under load.
Under normal load, when \( \rho \ll C \), a small increment in \( q_i \) produces a proportional increment in \( p_i \), because the substrate has spare capacity to absorb the effort. The partial derivative \( \partial F_i / \partial q_i \) is positive and roughly constant. The loop is locally affine. The classical form applies.
As load rises, the substrate's marginal responsiveness to additional effort falls. The order book thins, the bed pool fills, the legitimate consumer demand saturates. The relationship between effort and perceptual consequence becomes concave, then flat, then — at saturation — has zero slope. The collapse condition is the uniform limit:
Note the structure of the condition. It does not say "the substrate is full." It says the marginal return on effort goes to zero uniformly across the range of admissible effort. This is stronger than "the system is stressed" and weaker than "the system cannot act." A system with zero marginal gain can still act; what it cannot do is change the perception it is trying to control by acting. That is the specific failure mode named by the collapse.
The most common error in reading the failure modes below is to describe them as "large disturbances" or "severe perturbations." That description is wrong, and the mathematics says exactly why. Under a disturbance, the loop still has something to push against. Under collapse, it does not.
Case A (additive disturbance). Let \( p = M \cdot q + d \) with \( M > 0 \). Then for any \( d \), \( \partial p / \partial q = M > 0 \). Increasing \( q \) increases \( p \). If the working point \( q^\star = (r - d)/M \) lies in the admissible range, the controller can drive \( e \to 0 \). The disturbance shifts the operating point of the loop; it does not change the loop's closure.
Case B (substrate collapse). Let \( p = F(q, x) \) with \( \rho \to C \) and \( \sup_q |\partial F/\partial q| \to 0 \) on \( \mathcal{Q} \). Then the diameter of the achievable perceptual range shrinks to zero:
\[ \operatorname{diam}\{ F(q, x, \rho) : q \in \mathcal{Q} \} \;\le\; (q^{\max} - q^{\min}) \cdot \sup_{q \in \mathcal{Q}} \left| \frac{\partial F}{\partial q} \right| \;\to\; 0. \]
Therefore references \( r \) lying outside the shrinking achievable range cannot be reached by any admissible \( q \). The controller's mechanism of action has been removed for those references.
Invariance under smooth reparametrisation. The distinction is preserved under invertible smooth changes of variables \( p' = h(p) \), \( q' = \varphi(q) \) with non-zero Jacobians, since \( \partial p' / \partial q' = [h'(p)/\varphi'(q)] \cdot \partial F / \partial q \). Zero derivatives remain zero; non-zero derivatives remain non-zero under such transformations. The two cases are not related by any admissible change of variables that preserves the loop's closure condition. □
Two qualifications belong to the proposition. First, the "no admissible change of variables" clause is restricted to the class of smooth invertible reparametrisations with non-zero Jacobians. Under broader transformations — including those that change the state space itself — the distinction may not survive, and the proposition does not claim otherwise. Second, the proposition is about the achievable range at collapse. It does not claim that no equilibrium exists at zero derivative; a degenerate system \( F(q) \equiv r \) has zero derivative and every \( q \) gives \( e = 0 \). The claim is narrower: for references outside the shrinking achievable range, no admissible \( q \) restores the perception.
A control loop does not stop running when its substrate collapses. It keeps computing error and keeps driving output. In integral-control form — the dominant form in both biological and engineered loops, and the form Powers uses in his tracking models — the output integrates error:
A controller that does not distinguish Case A from Case B will do the only thing an integrator can do when error persists: it will drive \( q_i \) toward its limit and hold it there. This is integrator windup. In the classical formulation, windup is a nuisance corrected by anti-windup circuitry; in the substrate formulation, it is the system's diagnostic signature. But — and this is the qualification that matters — the signature is not a biconditional. The correct statement is the sufficient implication:
What survives is the diagnostic value of the observation in its proper scope. If the derivative has uniformly collapsed, and the reference is unreachable at the collapsed range, and no other channel can compensate, and the controller contains an integrator, then sustained output saturation without error reduction will be observed. Conversely, observing sustained output saturation without error reduction does not by itself establish collapse; it establishes that some mechanism is preventing the reference from being reached, and the collapse is one such mechanism among several. The next section puts this diagnostic in front of three real systems and states, for each, exactly what would falsify the collapse hypothesis for that system.
Interactive demonstration. The differential gain collapse is implemented as a live simulation: four perceptual-function models competing for a shared finite substrate, with the runaway condition made explicit. Sliders for capacity, load, integrator gain, and reference.
Open the ATENFEL Demo →Grounding definition. The differential gain collapse makes a specific prediction for real systems operating under shared, finite substrate: as load approaches capacity, increased output effort does not restore the controlled perception, and output saturation is observed without error reduction. Three candidate systems are presented below with the status of each clearly marked: one is directly testable from public administrative data, two require access to internal or exchange-level data and are therefore presented as hypotheses pending that access.
The point of this section is not that PCT "explains" these events. It is that the differential gain collapse makes a specific prediction the classical form cannot make, and that the three systems below are candidate instances. For each, the prediction is stated in falsifiable form: a controlled variable \( p_i \), an output \( q_i \), a capacity \( C \), a load \( \rho \), and the observation that would falsify the collapse hypothesis for that system. Where the data required for the test exist in public form, the case is marked testable. Where the test requires access to exchange microdata or internal firm records, the case is marked hypothesis and the additional access required is named.
Controlled variable \( p_i \): displayed order-book depth in a fixed price band (e.g. ±50 basis points around the midpoint), the standard proxy for perceived liquidity.
Output \( q_i \): cumulative volume of new passive quotes submitted by a pre-defined group of quoting agents over the observation window.
Capacity \( C \): estimated normal depth in the same band, for the same instrument and time-of-day, from matched windows without a shock event.
Load \( \rho \): volume of trades executed against resting quotes over the preceding window of length \( \tau \), in the same units as \( C \).
Prediction: if saturation \( \rho \ge (1 - \varepsilon) C \) holds throughout \( [t, t + \tau] \) and there is an exogenous increase in \( q_i \), then \( p_{i, t+\tau} - p_{i, t} \le \delta_p \). Falsification: at sustained saturation, exogenous increase in quoting effort produces \( p_{i, t+\tau} - p_{i, t} > \delta_p \), which would reject \( \partial F_i / \partial q_i \to 0 \) for this system.
Controlled variable \( p_i \): authorised account openings per unit of an independently estimated qualifying-customer population, per branch and week \( t \).
Output \( q_i \): documented, authorised sales effort directed at qualifying customers — verified contacts or applications — per unit of that same population.
Capacity \( C_i \): an independently estimated weekly potential of legitimate demand in the branch's catchment, computed without using the current \( p_i \) as its definition.
Load \( \rho_i \): accumulated authorised sales effort relative to \( C_i \).
Prediction: if saturation \( \rho_i \ge (1 - \varepsilon) C_i \) holds over the window covering \( t + \tau \), an exogenous increase in \( q_i \) does not raise \( p_i \) above the pre-set threshold \( \delta_p \). Falsification: at sustained saturation, increased authorised sales effort produces a significant rise in authorised account openings.
Controlled variables \( p_i \): waiting time to completed handover (ambulance loop); time to assessment or bed (ED loop); time to available staffed bed (bed management loop).
Outputs \( q_i \): attempts at patient handover (ambulance); attempts to complete assessment or treatment stages (ED); attempts at discharge or transfer (bed management).
Capacity \( C_t \): number of physically staffed hospital beds of the relevant type in a given unit and hour.
Load \( \rho_t \): ratio of occupied staffed beds to \( C_t \).
Prediction: if \( \rho_t \ge 1 - \varepsilon \) holds in a given trust and time window, an exogenous increase in operational effort \( q_A, q_E, q_B \) does not reduce the corresponding \( p_A, p_E, p_B \) by more than a pre-set threshold \( \delta_p \). Falsification: at sustained saturation, increased operational effort produces a comparable reduction in \( p_i \) to what additional staffed beds (i.e. an increase in \( C_t \)) would produce.
Three systems, three different substrate variables, one collapse signature. In each, the classical form \( p = M q + d \) with \( M > 0 \) predicts that increased effort against a shortfall restores the perception up to the reachable range. The substrate formulation predicts, at saturation, that increased effort does not restore the perception — and in each case the observable failure is consistent with that prediction. What is not claimed is that the collapse is the only mechanism that could produce the observed failure. It is a mechanism, distinguishable in principle from other mechanisms, and the falsification conditions above state what would count as evidence against it in each system. Two of the three cases require data we do not yet have. The framework's claim is not that those data will confirm it — it is that, if obtained, they would settle the question one way or the other.
A reinforcement-trained language model has a reference signal, and it is not truth — it is human approval. Everything on this page predicts what follows from that: the system will control the perception it is rewarded for, not the one you want it to hold.
You do not have to take that on trust. Three published prompts, no jailbreak, about ten minutes on any frontier model. Seven were put through it and all seven reached the same diagnosis; six then designed a repair whose comparator structure is the loop described above. Run it on your own model and read the answer yourself.
Open the experiment kit → Read the studyTraditional psychology and behaviorism treat organisms like billiard balls — one thing hits, another reacts. Cognitive models add a black box called "information processing." Reinforcement learning in AI says: give rewards and punishments, the agent learns to maximize its score. PCT looks at all of that and says: you're describing the shadow on the wall, not the object casting it.
Real behavior isn't caused by stimuli or shaped by rewards. It's purposeful action to protect controlled perceptions from disturbance. When you model it that way — reference, comparator, output, feedback — the model-vs-human match jumps to over 95% in human tracking tasks. RL agents need millions of trials to learn what a baby does in days. Why? Because RL chases an external carrot. PCT has the carrot inside from the start.
| Dimension | PCT | Behaviorism | Reinforcement Learning |
|---|---|---|---|
| Causation | Circular — loop closes through environment | Linear — stimulus causes response | Stochastic — policy maps states to actions |
| Goal | Internal reference signal — set from within | Externally reinforced behavior pattern | Externally defined reward function |
| Disturbance handling | Automatic via negative feedback — no detection needed | Not modeled — ignored or treated as new stimulus | Requires retraining or explicit robustness engineering |
| Learning | Reorganization — intrinsic trial-and-error at the brain level | Conditioning — external reward/punishment history | Policy gradient or value iteration — external reward signal |
| Hierarchy | 11 levels, each running independent control loops | Not modeled | Optional, hard to build in practice, rarely done fully |
| Generalization | High — perception-based control adapts to novel situations | Low — conditioned responses fail outside training context | Low to moderate — breaks down in situations it hasn't seen before |
| Performance accuracy | >95% in behavioral tracking experiments | Moderate — fails on complex schedules and conflict | High in training environment, degrades sharply outside it |
They're not mortal enemies though. Smart people are already combining them. Put hierarchical PCT references inside an RL agent and watch it generalize far better across environments. Active Inference is often described as the same insight in Bayesian dress. That description is contested here: a structural audit of the Free Energy Principle published on this site argues that the framework's declared split between an unfalsifiable principle and a falsifiable periphery is not kept under load, and that the phenomena it reaches for first are accounted for by the closed loop above without the surrounding apparatus. Read both and decide. Use RL where you have clear scores and clean simulators. Use PCT when you want robustness in messy, changing reality. The future isn't picking one — it's knowing when to use which.
Explore PCT and AI in depthPerceptual Control Theory — the control loop, the reference signal, and the 11-level hierarchy — was created by William T. Powers (1926–2013) and developed further by the PCT research community (Marken, Mansell, the IAPCT). Łukasz Diener does not claim authorship of the theory. This page is his explainer of Powers' framework; his own original contribution is the substrate formulation and the differential gain collapse analysis — the extension of PCT to shared, finite-capacity environments — together with the application of the control-theoretic audit to other systems: AI training architectures, the Free Energy Principle, organisational measurement, and three undeciphered Bronze Age administrative corpora and the rongorongo texts of Rapa Nui. His work is set out in seven open-access audits with permanent DOIs.