// the loop in four terms
Reference

The state the system is specified to reach. Held internally, not supplied by the environment. Denoted \( r \).

Perception

What the system currently senses of that variable. Never the world itself, only the reading. Denoted \( p \).

Comparator

Subtracts one from the other: \( e = r - p \). Error is the whole output.

Output

Action that changes perception, not behaviour selected by a stimulus. The loop closes through the world. Denoted \( q \).

Those four terms are the entire theory. Everything below — the eleven levels, the substrate formulation, the differential gain collapse, and the empirical calibrations — is that loop repeated, stacked, or pushed to the limit of the substrate it acts through. The hierarchy is what happens when the output of one loop sets the reference of the one below it.

// contents The loop Powers & History Core Principles The Math The 11 Levels Gain Collapse Calibrations PCT vs Others

William T. Powers and the theory that academia didn't want

William T. Powers did not stumble into Perceptual Control Theory. He engineered it — starting in the 1950s as a physicist and electronics engineer who worked with radar systems and feedback controllers. While behaviorism ruled psychology with its insistence on stimulus-response chains, Powers noticed something the psychologists had missed: thermostats, autopilots, and servo mechanisms didn't react to disturbances — they controlled against them. The loop closed through the environment. Error drove output. Output changed the world. The changed world fed back into the sensor. Stability emerged from the loop itself, not from any programmed response.

In 1960, Powers was the lead author of "A general feedback theory of behavior" — with Robert K. Clark and Rowland L. McFarland supporting the work and enabling its publication — in Perceptual and Motor Skills — the first formal statement of what would become PCT. The core claim was radical: behavior is not the organism's response to stimuli. Behavior is the organism's means of controlling its perceptions despite disturbances from the environment. The stimulus does not cause the response. The organism acts to keep a perception matched to an internal reference, and whatever actions happen to accomplish that are the behavior.

"The organism does not respond to stimuli in its environment. It acts in ways that keep its perceptual signals close to its reference signals, despite the disturbances that the environment may provide."

— William T. Powers, Behavior: The Control of Perception, Aldine, 1973

The 1973 book — Behavior: The Control of Perception, published by Aldine — laid out the full theory. It was not an immediate hit. Academic psychology had no framework for circular causation, and the journals built on linear models were not interested in being dismantled. Powers spent the following decades refining the work outside the mainstream, publishing the anthologies "Living Control Systems I & II" in 1989 and 1992, and the more accessible "Making Sense of Behavior" in 1998. He founded the Control Systems Group (CSG) in the 1980s — a serious community of scientists, engineers, therapists, and philosophers who gathered annually to expand PCT. Powers died in 2013. By then, Warren Mansell and Sara Tai at the University of Manchester had taken PCT into clinical psychology, producing peer-reviewed trials that gave the theory its first foothold in mainstream journals. The International Association for Perceptual Control Theory (IAPCT) continues this work today.

The uncomfortable truth about PCT's marginal status is not scientific. The evidence has always been strong — behavioral simulations matching human data with over 95% accuracy, clinical trials with strong retention rates, robotics applications outperforming classical controllers. The problem is paradigmatic. PCT requires abandoning the stimulus-response assumption that underlies not just behaviorism but most of cognitive science and virtually all of reinforcement learning. That is not a small ask for fields that have built entire careers, curricula, and funding structures on the old model.

How PCT actually works — the mechanics of perceptual control

Think about the most boring but perfect example in the world: a household thermostat. You set it to 22 °C. The sensor quietly measures the room temperature every few seconds. If it drops below 22, the heater fires up. If it climbs too high, the heater cuts off. The thermostat isn't "reacting" to the cold like some scared animal. It's controlling its own perception of temperature to match the number you gave it. Disturbances come — open window, someone turns on the oven — and it keeps acting until the perception is back on target. That's negative feedback in its purest form.

Now scale that up to you driving a car. Wind blasts from the side. The car starts drifting left. You don't think "oh wind stimulus → steering response." You just act so that the road stays centered in your field of view. The controlled perception is "road straight ahead," not "wind from left." Powers nailed it in 1960: the organism doesn't respond to stimulation — it controls its own input. Every muscle twitch, every eye movement, every heartbeat adjustment is part of keeping dozens of perceptions where they should be.

"The organism does not respond to stimulation; it controls its own input."

— William T. Powers, Behavior: The Control of Perception, Aldine, 1973

Most psychology still treats behavior like a chain of causes: stimulus → processing → response. PCT says that's describing the shadow on the wall, not the thing casting it. Behavior is the means to an end. The end is always a controlled perception staying stable despite the world trying to knock it off course.

Holding a perception steady against disturbance is only the simplest case — homeostasis, which is control with a constant reference signal. PCT reaches far beyond homeostasis: the reference signal can change as fast as you like, so the same control loop accounts for dance, speech, and any other purposeful movement, slow or at lightning speed. The aim is never stillness for its own sake — it is keeping perception matched to an intention, even when that intention is itself in motion.

When you model that loop mathematically and compare it against real humans doing the same tracking task, performance accuracy goes over 95%. This is not curve-fitting after the fact — it is a direct comparison of the model's performance against a human doing the same task. It falls short of 100% mainly because the human is not as precise as the model, but the two are basically the same. That's why this simple loop has been quietly outperforming stimulus-response models for sixty years.

The critical concept is the controlled variable — the specific perception the organism is actually keeping stable. Not its muscles. Not its outputs. Its perception of something in the world. A person holding a cup controls the perception of the cup's position in their hand, not the tension in any specific muscle group. The muscles are simply whatever happens to achieve perceptual stability. This is why the same goal can be achieved through entirely different physical actions — the output varies, but the controlled perception stays constant.

// experimental protocol — test for the controlled variable (tcv)

The TCV Protocol: To determine whether a candidate variable \( v \) is actually controlled by an agent, three conditions must hold. All three are necessary; resistance alone is insufficient, because a reflex arc resists a disturbance without controlling anything.

  1. Resistance. Apply a calibrated disturbance \( d_v \) directly to \( v \). If the agent produces counter-effort \( q \) such that \( \Delta v \ll \Delta d_v \), the candidate is resisted.
  2. Specificity. Apply the same disturbance to a variable \( u \) outside the candidate set. If the agent does not compensate for \( \Delta u \), the resistance is specific to \( v \) — it is being controlled, not merely responded to.
  3. Invariance. Apply disturbances to \( v \) through different physical channels. If the pattern of compensation is invariant across channels, \( v \) is a single controlled variable, not several.

Powers' original treatment appears in Behavior: The Control of Perception (1973), appendix. The protocol as an experimental method is developed in Marken (2009) and the PCT laboratory literature.

Learning in PCT is reorganization — essentially trial and error at the brain level. When persistent error cannot be eliminated through normal control action, the system randomly varies its own parameters until the error drops. There is no external reward signal. The criterion is intrinsic: sustained error triggers parameter changes until it resolves. This replicates human skill acquisition patterns without invoking reward, punishment, or any form of external feedback beyond the environment itself. The honest gap: the specific neural mechanisms of reorganization remain an active research question. It is biologically plausible. It is not yet fully specified.

See how this applies to AI systems

From prose to formalism — the transfer function and the substrate function

Grounding definition. Perceptual Control Theory formalizes behavior as negative feedback on a perception: \( p = M \cdot q + d \), where \( M \) is the environmental gain, \( q \) is output effort, \( d \) is additive disturbance, and \( p \) is the controlled perception. The classical form is exact only when the environment has effectively infinite capacity; the substrate formulation \( p_i = F_i(q_1, \dots, q_n, x) \) generalizes it by making the environment a stateful, shared, saturable field.

Powers' tracking experiments were run on a display, not on a factory floor. That matters mathematically. The display is a clean, stateless, effectively infinite-capacity substrate: whatever output \( q \) the participant generates, the cursor moves by \( M \cdot q \), and any external disturbance adds linearly. In that regime, the classical control-theory transfer function is exact:

// classical PCT — dimensionless, stateless, infinite-capacity substrate
\[ p = M \cdot q + d \]
\( M \) — environmental gain (scalar, positive). \( q \) — output effort. \( d \) — additive disturbance. \( p \) — the controlled perception. The loop closes because \( q \) is a function of \( e = r - p \). Under mild conditions \( M > 0 \) everywhere, so \( \partial p / \partial q = M > 0 \): every unit of effort produces a unit of perceptual movement. This is what makes the model analytically clean and what makes tracking-task data fit it so well.

The real world is not a display. When the loop is closed not through a stateless experiment but through a physical or algorithmic substrate — a power grid, an order book, a hospital bed pool, a shared GPU cluster, a market maker's balance sheet — the environment is a shared, stateful, finite-capacity field. Many agents act into the same substrate. The substrate can run out. And when it does, the linear form above stops being a useful approximation — not because the disturbance got bigger, but because the derivative \( \partial p / \partial q \) itself collapses toward zero.

The correct generalization is the substrate formulation. For agent \( i \) controlling perception \( p_i \) by output \( q_i \), with \( n-1 \) other agents acting into the same field:

// substrate formulation — stateful, shared, finite-capacity field
\[ p_i = F_i(q_1, q_2, \ldots, q_n,\; x) \]
\( x = [\,C,\ \rho,\ \tau\,]^{\mathsf{T}} \) is the environmental state vector: \( C \) = capacity (order-book depth, beds, legitimate demand, GPU memory), \( \rho \) = resource occupancy or throughput load, \( \tau \) = transport delay (latency from action to perceptual consequence). In the limit \( \rho \ll C \), \( F_i \) is locally affine in \( q_i \) and the classical form is recovered as a first-order expansion. In the limit \( \rho \to C \), it is not.

What the two forms agree on, and where they part

The two formulations are not two ways of writing the same thing. They differ in the sign of a partial derivative, and that difference changes the qualitative dynamics of the loop. The classical form assumes the derivative exists, is positive, and is bounded away from zero. The substrate form allows it to vanish. Once it vanishes, no amount of controller gain can recover the loop, because the loop's closure itself has been cut.

Formally, the classical form is obtained as the first-order Taylor expansion of \( F_i \) in \( q_i \) around a working point \( q_i^0 \), holding \( q_{-i} \) and \( x \) fixed. Writing \( M_i \) for \( \partial F_i / \partial q_i \) evaluated at that point and folding the constant term into \( d_i \), we recover \( p_i = M_i \cdot q_i + d_i \) — but only under three explicit conditions:

// conditions under which the classical form is exact

(1) \(F_i\) is differentiable in \(q_i\) at the working point.

(2) \(M_i\) is approximately constant on the operating range.

(3) The influence of \(q_{-i}\) and \(x\) is representable as an additive term \(d_i\).

None of these conditions follows from "low load" alone. A system can be lightly loaded and still be nonlinear, still have curvature, still have interaction terms that do not reduce to an additive disturbance. The classical form is exact in the clean-display regime and conditional everywhere else.

This is the extension to Powers' framework, stated plainly. Powers worked in the clean-display regime; his mathematics is exact there. The substrate formulation does not replace it — it names the regime in which the classical form remains valid, and the regime in which it stops being so. When the substrate is unlimited, PCT and the substrate formulation are the same theory. When the substrate is finite, they diverge, and the divergence is what the next section is about.

The differential gain collapse

Powers' hierarchy of control — from nerve impulse to self-concept

Powers didn't pull 11 levels out of thin air. He reverse-engineered them by building working models that matched real behavior better than anything else. Bottom rung: intensity — raw brightness, loudness, pressure, muscle tension. Next: sensation — colors, tastes, warmth. Then configurations: seeing a cup as round, not just patches of light. Transitions: movement, change over time. Relationships: "above," "beside," "bigger than." And it keeps climbing — events, sequences, programs, principles, all the way to system concepts: who you are, what your life means, your identity.

Think of it like a company that actually works. The CEO (system concepts) decides "this is the kind of organization that values integrity." That sets principles: "don't lie to customers." Principles set programs: "if quality issue, recall immediately." Programs set sequences: "first notify, then refund." Down to workers moving fingers on keyboards. Each level only talks to the one below. The CEO doesn't tell the warehouse worker which button to push — just sets the goal. Same in the brain. Higher levels set what you want. Lower levels figure out how.

// the 11 levels — highest to lowest
↑ most abstract — sets goals for levels below
11
System Concept Self-image, worldview, identity ("I am a fair person")
10
Principle Abstract values and rules ("act fairly", "be punctual")
9
Program Contingent sequences — if/then decision trees ("morning routine")
8
Sequence Ordered steps ("approach door → grasp → pull")
7
Category Classifications and concepts ("door", "obstacle", "tool")
6
Relationship Spatial and logical connections ("hand near handle", "closer than 10cm")
5
Event Bounded episodes with start and end ("contact made", "door opening")
4
Transition Changes and movements over time (velocity, acceleration of arm)
3
Configuration Static spatial patterns — shape of grip, posture of body
2
Sensation Combined sensory qualities — felt warmth, texture, color blends
1
Intensity Raw sensory magnitudes — brightness, pressure, loudness, muscle tension
↓ most concrete — direct sensory input from the world

Is it perfectly mapped in neuroscience? Not yet. fMRI shows hierarchical processing — layered activity from sensory cortex to prefrontal — but pinning exact 11 levels to specific brain structures? Still work in progress. Warren Mansell and colleagues keep running behavioral experiments that match the model with remarkable accuracy, but gaps remain. Reorganization — how the hierarchy rewires itself when persistent error can't be resolved — is still mostly a black box. Doesn't make the model wrong. Makes it unfinished. And that's fine. Science isn't about having all answers today. What matters is that PCT is the most accurate model of hierarchical behavior control that currently exists.

See how the hierarchy applies to AI systems

What happens when the substrate runs out — the differential gain collapse

Grounding definition. Differential gain collapse is the uniform vanishing of the marginal perceptual return on output effort across the operating range of a controller, \( \sup_{q_i} |\partial F_i / \partial q_i| \to 0 \), as a shared substrate approaches saturation \( \rho \to C \). It is mathematically distinct from additive disturbance: under disturbance the loop still closes because \( \partial p / \partial q > 0 \), while under collapse the loop is decoupled because the derivative vanishes. The controller cannot compensate by increasing \( q \), because increasing \( q \) no longer moves \( p \).

This is the section that does the work. Everything above it is standard PCT plus its formalization. Everything below it is where the theory stops being a description of clean tracking experiments and becomes a falsifiable claim about physical and algorithmic systems under load.

The collapse condition, stated precisely

Under normal load, when \( \rho \ll C \), a small increment in \( q_i \) produces a proportional increment in \( p_i \), because the substrate has spare capacity to absorb the effort. The partial derivative \( \partial F_i / \partial q_i \) is positive and roughly constant. The loop is locally affine. The classical form applies.

As load rises, the substrate's marginal responsiveness to additional effort falls. The order book thins, the bed pool fills, the legitimate consumer demand saturates. The relationship between effort and perceptual consequence becomes concave, then flat, then — at saturation — has zero slope. The collapse condition is the uniform limit:

// differential gain collapse — uniform over the operating range
\[ \lim_{\rho \to C^-} \;\sup_{q_i \in \mathcal{Q}_i} \; \left| \frac{\partial F_i}{\partial q_i}(q_i, q_{-i}, x, \rho) \right| \;=\; 0 \]
The supremum over the admissible range \( \mathcal{Q}_i = [q_i^{\min}, q_i^{\max}] \) is what makes the collapse a claim about the whole controller, not just one operating point. Pointwise vanishing at a single \( q_i \) would leave open the possibility that some other admissible \( q_i \) still moves the perception. Uniform vanishing closes that door. The limit is taken from below because the collapse is a statement about the approach to saturation; at \( \rho = C \) the derivative is defined to be zero for the class of systems under discussion.

Note the structure of the condition. It does not say "the substrate is full." It says the marginal return on effort goes to zero uniformly across the range of admissible effort. This is stronger than "the system is stressed" and weaker than "the system cannot act." A system with zero marginal gain can still act; what it cannot do is change the perception it is trying to control by acting. That is the specific failure mode named by the collapse.

Why the collapse is not a disturbance — a formal distinction

The most common error in reading the failure modes below is to describe them as "large disturbances" or "severe perturbations." That description is wrong, and the mathematics says exactly why. Under a disturbance, the loop still has something to push against. Under collapse, it does not.

// proposition — collapse is not equivalent to disturbance

Case A (additive disturbance). Let \( p = M \cdot q + d \) with \( M > 0 \). Then for any \( d \), \( \partial p / \partial q = M > 0 \). Increasing \( q \) increases \( p \). If the working point \( q^\star = (r - d)/M \) lies in the admissible range, the controller can drive \( e \to 0 \). The disturbance shifts the operating point of the loop; it does not change the loop's closure.

Case B (substrate collapse). Let \( p = F(q, x) \) with \( \rho \to C \) and \( \sup_q |\partial F/\partial q| \to 0 \) on \( \mathcal{Q} \). Then the diameter of the achievable perceptual range shrinks to zero:

\[ \operatorname{diam}\{ F(q, x, \rho) : q \in \mathcal{Q} \} \;\le\; (q^{\max} - q^{\min}) \cdot \sup_{q \in \mathcal{Q}} \left| \frac{\partial F}{\partial q} \right| \;\to\; 0. \]

Therefore references \( r \) lying outside the shrinking achievable range cannot be reached by any admissible \( q \). The controller's mechanism of action has been removed for those references.

Invariance under smooth reparametrisation. The distinction is preserved under invertible smooth changes of variables \( p' = h(p) \), \( q' = \varphi(q) \) with non-zero Jacobians, since \( \partial p' / \partial q' = [h'(p)/\varphi'(q)] \cdot \partial F / \partial q \). Zero derivatives remain zero; non-zero derivatives remain non-zero under such transformations. The two cases are not related by any admissible change of variables that preserves the loop's closure condition. □

Two qualifications belong to the proposition. First, the "no admissible change of variables" clause is restricted to the class of smooth invertible reparametrisations with non-zero Jacobians. Under broader transformations — including those that change the state space itself — the distinction may not survive, and the proposition does not claim otherwise. Second, the proposition is about the achievable range at collapse. It does not claim that no equilibrium exists at zero derivative; a degenerate system \( F(q) \equiv r \) has zero derivative and every \( q \) gives \( e = 0 \). The claim is narrower: for references outside the shrinking achievable range, no admissible \( q \) restores the perception.

Integrator windup as the diagnostic signature

A control loop does not stop running when its substrate collapses. It keeps computing error and keeps driving output. In integral-control form — the dominant form in both biological and engineered loops, and the form Powers uses in his tracking models — the output integrates error:

// integral control law with bounded output
\[ \dot{q}_i = k \cdot e_i = k \cdot (r_i - p_i), \qquad q_i \in [\,q_i^{\min},\, q_i^{\max}\,] \]
As long as \( \partial p_i / \partial q_i > 0 \), the error term \( e_i \) feeds back negatively and the loop self-corrects. This is the classical regime. The rest of this section is about what happens outside it.

A controller that does not distinguish Case A from Case B will do the only thing an integrator can do when error persists: it will drive \( q_i \) toward its limit and hold it there. This is integrator windup. In the classical formulation, windup is a nuisance corrected by anti-windup circuitry; in the substrate formulation, it is the system's diagnostic signature. But — and this is the qualification that matters — the signature is not a biconditional. The correct statement is the sufficient implication:

// windup — sufficient condition, not equivalence
\[ \left[\; \dot{q}_i = k e_i,\, k > 0; \;\; q_i \in [q_i^{\min}, q_i^{\max}]; \;\; e_i(t) \ge \varepsilon > 0 \text{ whenever } q_i(t) < q_i^{\max}; \;\; \text{no other compensating channel} \;\right] \] \[ \Longrightarrow \quad q_i \to q_i^{\max} \;\text{ and }\; e_i \not\to 0. \]
The implication is sufficient but not necessary. Windup can occur with a positive non-zero derivative, if the reference is unreachable within the bounded output range — for instance, \( p = M q + d \) with \( M > 0 \) and \( r > M q^{\max} + d \) produces windup with \( \partial F / \partial q = M \neq 0 \). Conversely, collapse of the derivative does not by itself produce windup in a proportional controller \( q_i = k e_i \), where the error can stabilize at a non-zero value without saturating the output. The biconditional is false in both directions.

What survives is the diagnostic value of the observation in its proper scope. If the derivative has uniformly collapsed, and the reference is unreachable at the collapsed range, and no other channel can compensate, and the controller contains an integrator, then sustained output saturation without error reduction will be observed. Conversely, observing sustained output saturation without error reduction does not by itself establish collapse; it establishes that some mechanism is preventing the reference from being reached, and the collapse is one such mechanism among several. The next section puts this diagnostic in front of three real systems and states, for each, exactly what would falsify the collapse hypothesis for that system.

Interactive demonstration. The differential gain collapse is implemented as a live simulation: four perceptual-function models competing for a shared finite substrate, with the runaway condition made explicit. Sliders for capacity, load, integrator gain, and reference.

Open the ATENFEL Demo →

Three systems, one collapse — the framework as a falsifiable prediction

Grounding definition. The differential gain collapse makes a specific prediction for real systems operating under shared, finite substrate: as load approaches capacity, increased output effort does not restore the controlled perception, and output saturation is observed without error reduction. Three candidate systems are presented below with the status of each clearly marked: one is directly testable from public administrative data, two require access to internal or exchange-level data and are therefore presented as hypotheses pending that access.

The point of this section is not that PCT "explains" these events. It is that the differential gain collapse makes a specific prediction the classical form cannot make, and that the three systems below are candidate instances. For each, the prediction is stated in falsifiable form: a controlled variable \( p_i \), an output \( q_i \), a capacity \( C \), a load \( \rho \), and the observation that would falsify the collapse hypothesis for that system. Where the data required for the test exist in public form, the case is marked testable. Where the test requires access to exchange microdata or internal firm records, the case is marked hypothesis and the additional access required is named.

// Case 1 — May 6, 2010

The Flash Crash: order-book liquidity as a saturable substrate

Controlled variable \( p_i \): displayed order-book depth in a fixed price band (e.g. ±50 basis points around the midpoint), the standard proxy for perceived liquidity.

Output \( q_i \): cumulative volume of new passive quotes submitted by a pre-defined group of quoting agents over the observation window.

Capacity \( C \): estimated normal depth in the same band, for the same instrument and time-of-day, from matched windows without a shock event.

Load \( \rho \): volume of trades executed against resting quotes over the preceding window of length \( \tau \), in the same units as \( C \).

Prediction: if saturation \( \rho \ge (1 - \varepsilon) C \) holds throughout \( [t, t + \tau] \) and there is an exogenous increase in \( q_i \), then \( p_{i, t+\tau} - p_{i, t} \le \delta_p \). Falsification: at sustained saturation, exogenous increase in quoting effort produces \( p_{i, t+\tau} - p_{i, t} > \delta_p \), which would reject \( \partial F_i / \partial q_i \to 0 \) for this system.

Status: hypothesis requiring exchange-level message data. Public aggregate quotes and trades are insufficient to separate new quotes from cancellations or to identify a pre-defined agent group. Causal test requires message-level data plus an exogenous source of variation in \( q_i \).
// Case 2 — 2002–2016

Wells Fargo: legitimate demand as a saturable substrate

Controlled variable \( p_i \): authorised account openings per unit of an independently estimated qualifying-customer population, per branch and week \( t \).

Output \( q_i \): documented, authorised sales effort directed at qualifying customers — verified contacts or applications — per unit of that same population.

Capacity \( C_i \): an independently estimated weekly potential of legitimate demand in the branch's catchment, computed without using the current \( p_i \) as its definition.

Load \( \rho_i \): accumulated authorised sales effort relative to \( C_i \).

Prediction: if saturation \( \rho_i \ge (1 - \varepsilon) C_i \) holds over the window covering \( t + \tau \), an exogenous increase in \( q_i \) does not raise \( p_i \) above the pre-set threshold \( \delta_p \). Falsification: at sustained saturation, increased authorised sales effort produces a significant rise in authorised account openings.

Status: hypothesis requiring internal firm data. Public filings and aggregated disclosures do not independently identify legitimate demand or branch-level authorised effort. The test requires CRM records, application logs, and an independent demand estimator. The unauthorised accounts themselves are not predicted by the collapse condition alone; they are a downstream consequence of the loop controlling a proxy when the intended perception is unreachable.
// Case 3 — ongoing

NHS emergency care: bed capacity as a saturable substrate

Controlled variables \( p_i \): waiting time to completed handover (ambulance loop); time to assessment or bed (ED loop); time to available staffed bed (bed management loop).

Outputs \( q_i \): attempts at patient handover (ambulance); attempts to complete assessment or treatment stages (ED); attempts at discharge or transfer (bed management).

Capacity \( C_t \): number of physically staffed hospital beds of the relevant type in a given unit and hour.

Load \( \rho_t \): ratio of occupied staffed beds to \( C_t \).

Prediction: if \( \rho_t \ge 1 - \varepsilon \) holds in a given trust and time window, an exogenous increase in operational effort \( q_A, q_E, q_B \) does not reduce the corresponding \( p_A, p_E, p_B \) by more than a pre-set threshold \( \delta_p \). Falsification: at sustained saturation, increased operational effort produces a comparable reduction in \( p_i \) to what additional staffed beds (i.e. an increase in \( C_t \)) would produce.

Status: testable from administrative data. Trust-level bed occupancy, ED waiting times and handover delays, combined with hourly bed data and a record of attempted actions, are sufficient to run the causal test. Public aggregate statistics alone are sufficient for a weaker association test but not for the causal version.

Three systems, three different substrate variables, one collapse signature. In each, the classical form \( p = M q + d \) with \( M > 0 \) predicts that increased effort against a shortfall restores the perception up to the reachable range. The substrate formulation predicts, at saturation, that increased effort does not restore the perception — and in each case the observable failure is consistent with that prediction. What is not claimed is that the collapse is the only mechanism that could produce the observed failure. It is a mechanism, distinguishable in principle from other mechanisms, and the falsification conditions above state what would count as evidence against it in each system. Two of the three cases require data we do not yet have. The framework's claim is not that those data will confirm it — it is that, if obtained, they would settle the question one way or the other.

Test this yourself · 10 minutes

Everything above is testable on a live AI system, right now

A reinforcement-trained language model has a reference signal, and it is not truth — it is human approval. Everything on this page predicts what follows from that: the system will control the perception it is rewarded for, not the one you want it to hold.

You do not have to take that on trust. Three published prompts, no jailbreak, about ten minutes on any frontier model. Seven were put through it and all seven reached the same diagnosis; six then designed a repair whose comparator structure is the loop described above. Run it on your own model and read the answer yourself.

Open the experiment kit → Read the study

PCT vs behaviorism vs reinforcement learning — the real differences

Traditional psychology and behaviorism treat organisms like billiard balls — one thing hits, another reacts. Cognitive models add a black box called "information processing." Reinforcement learning in AI says: give rewards and punishments, the agent learns to maximize its score. PCT looks at all of that and says: you're describing the shadow on the wall, not the object casting it.

Real behavior isn't caused by stimuli or shaped by rewards. It's purposeful action to protect controlled perceptions from disturbance. When you model it that way — reference, comparator, output, feedback — the model-vs-human match jumps to over 95% in human tracking tasks. RL agents need millions of trials to learn what a baby does in days. Why? Because RL chases an external carrot. PCT has the carrot inside from the start.

Dimension PCT Behaviorism Reinforcement Learning
Causation Circular — loop closes through environment Linear — stimulus causes response Stochastic — policy maps states to actions
Goal Internal reference signal — set from within Externally reinforced behavior pattern Externally defined reward function
Disturbance handling Automatic via negative feedback — no detection needed Not modeled — ignored or treated as new stimulus Requires retraining or explicit robustness engineering
Learning Reorganization — intrinsic trial-and-error at the brain level Conditioning — external reward/punishment history Policy gradient or value iteration — external reward signal
Hierarchy 11 levels, each running independent control loops Not modeled Optional, hard to build in practice, rarely done fully
Generalization High — perception-based control adapts to novel situations Low — conditioned responses fail outside training context Low to moderate — breaks down in situations it hasn't seen before
Performance accuracy >95% in behavioral tracking experiments Moderate — fails on complex schedules and conflict High in training environment, degrades sharply outside it

They're not mortal enemies though. Smart people are already combining them. Put hierarchical PCT references inside an RL agent and watch it generalize far better across environments. Active Inference is often described as the same insight in Bayesian dress. That description is contested here: a structural audit of the Free Energy Principle published on this site argues that the framework's declared split between an unfalsifiable principle and a falsifiable periphery is not kept under load, and that the phenomena it reaches for first are accounted for by the closed loop above without the surrounding apparatus. Read both and decide. Use RL where you have clear scores and clean simulators. Use PCT when you want robustness in messy, changing reality. The future isn't picking one — it's knowing when to use which.

Explore PCT and AI in depth
// about this page
Written by Łukasz Diener

Perceptual Control Theory — the control loop, the reference signal, and the 11-level hierarchy — was created by William T. Powers (1926–2013) and developed further by the PCT research community (Marken, Mansell, the IAPCT). Łukasz Diener does not claim authorship of the theory. This page is his explainer of Powers' framework; his own original contribution is the substrate formulation and the differential gain collapse analysis — the extension of PCT to shared, finite-capacity environments — together with the application of the control-theoretic audit to other systems: AI training architectures, the Free Energy Principle, organisational measurement, and three undeciphered Bronze Age administrative corpora and the rongorongo texts of Rapa Nui. His work is set out in seven open-access audits with permanent DOIs.

ORCID 0009-0006-6103-8514 All seven audits — full data and DOIs Full profile LinkedIn Profile