sources · 1
Continual behavior adaptation in an ever-changing world involves continuous mode switching between moving to achieve a goal (“goal-directed”) and sensitively perceiving their environment (“sensory-focused”). Moving for obtaining food or water and examining the surrounding environment or danger clearly demonstrates the significance of the dual modes for survival. Notably, cognitive agents exhibit predictive regulation of nutritional (energy) status in anticipation of future environmental changes1,2,3, suggesting that the predicted future shapes present behavior. For example, hibernating or migrating animals increase their food intake prior to periods when food is scarce or unavailable4,5. This phenomenon, referred to as behavioral allostasis, shows that both modes play a profound role in maintaining homeostasis in a dynamic environment. However, trade-offs between goal-directed movement and sensory-focused perception (e.g., to carefully hear the surrounding sounds, one must stop) as well as the phenomenon of goal emergence complicate the mechanistic understanding of autonomous mode switching6,7. The mechanisms by which the brain resolves the computational conflict between movement and perception, the core elements of cognition, remain unclear. Therefore, this study proposes a multimodal neural network model of autonomous behavioral allostasis based on the Bayesian brain theory. It investigated how the mode switching between goal-directed movement and sensory-focused perception can emerge through autonomous behavioral allostasis. Exploring these things may be important for understanding the relationship between life and cognition8,9.
Autonomous mode switching between goal-directed movement and sensory-focused perception exposes critical issues in the computational modeling of the brain and the construction of autonomous intelligence. The notion that the brain is a predictive or generative machine has been formalized in neuroscience as predictive coding and active inference (or the free-energy principle (FEP))10,11,12,13,14. Additionally, it is applied to architectures in machine learning (deep learning) and robotics15,16,17 as well as computational understanding of psychiatric disorders18,19,20,21,22. The neurobiologically and systematically plausible computational basis converges into the Bayesian framework, or prediction error minimization (minimization of the mismatch between real and predicted sensations). Within which a Bayesian agent infers the causes of sensations as internal beliefs (latent variables), influenced by prior experiences. The conflict between the computational processes of perception and movement is a critical issue in explaining autonomous switching between both modes. Specifically, in the context of perception (perceptual inference), Bayesian beliefs should be adapted to reduce sensory prediction errors by fitting sensations. In contrast, during the execution of movements (active inference), these beliefs should remain unchanged and be realized by changing sensations. In other words, Bayesian beliefs serve as “recognition” in sensory-focused perception but “intention” in goal-directed movement. Furthermore, goal-directed movement itself is a risk factor for prediction errors due to interactions with the environment. This may be related with the argument of “the dark-room problem”23.
Previous computational modeling and neurorobotic studies have heuristically avoided these problems by operating or modeling one process of movement or perception at a time24,25, explicitly training an agent to move in a specific manner26,27,28, and/or preparing separate latent variables and cost functions for perception and movement (action planning)29,30,31. However, these heuristic solutions neglect autonomous mode switching between goal-directed movement and sensory-focused perception in the face of computational conflicts experienced in daily life. Furthermore, the same latent variable (Bayesian belief) can serve as both “recognition” and “intention” (e.g., the experience of seeing or recognizing a certain food in a forest leads to the formation of the intention to find that food in the forest later on). Behavioral allostasis through autonomous switching between both modes may require an account beyond prediction error minimization for explaining how “recognition” changes into “intention” and vice versa. This may involve the dynamic self-organization of future sensorimotor goals and their precision (or strength), which modulates the balance between sensory information and goal representation. However, the mechanisms behind autonomous goal generations for behavioral allostasis and the mode switching between goal-directed movement and sensory-focused perception within autonomous behavioral allostasis remain unclear.
In this study, we propose that a cognitive agent not only minimizes the sensory prediction errors at hand but also minimizes the predicted future sensory entropy (uncertainty). This hypothesis aims to explain the continual adaptation of homeostatic behaviors by linking the Bayesian brain hypothesis with a physical view of homeostasis. As physical systems, biological cognitive agents naturally decay without a cycle of nutrition or excretion, eventually leading to death. This inevitable tendency toward disorder may be associated with an increase in entropy owing to irreversible physical processes within the system (body)32. Thus, homeostatic behaviors of cognitive agents may be regarded as means of resisting the tendency toward disorder, as Schrödinger originally stated that life feeds on “negative entropy” or free-energy33. We consider this physical constraint in brain information processing given that sensation is the interface between the brain information processing system and the bodily physical system. Then, minimization of predicted future sensory entropy can be considered a “meta-goal” concerning second-order prediction (of uncertainty), which may explain the dynamic self-organization of future sensorimotor goals and their precision or strength for behavioral allostasis.
The basic premise of this hypothesis is that appropriate physiological (interoceptive) conditions in the body are required to maintain homeostasis. Conversely, unusual physiological conditions increase the influence of the physical tendency towards disorder (physiological entropy increase). This may be supported by that certain food habits (e.g., inappropriate caloric content of diet) and thermal stress increase entropy production34,35,36,37. Unusual physiological conditions can affect or disrupt all aspects of the body including sensory organs, leading to noisy sensory signals, which may be experienced as bad physical conditions such as dizziness and sluggishness in severe cases. Therefore, sensory (information) entropy, or uncertainty, is a potential source from which the brain can recognize the homeostatic context of the system through sensorimotor experience.
To test our hypothesis, we developed a Bayesian neural network framework for homeostatic behavioral adaptation via hierarchical multimodal integration. Behavioral allostasis is associated with various aspects of cognitive processing, including learning of predictive models, perception of internal and external environments, movement generation, and goal generation. A computational model that can explain behavioral allostasis requires a principle that can explain various aspects of cognitive processing and multimodal integration involving interoception, exteroception, and proprioception. Thus, we aimed to incorporate the Bayesian brain theory and multimodal predictive processing into the neural network model. In addition, we showed that the meta-goal of Bayesian allostasis (minimizing entropy) can be derived by extending the framework of variational Bayes, a formal tractable solution of Bayesian inference, into future predictive processing. Our simulation experiment demonstrates how autonomous mode switching emerges through Bayesian allostasis. Our modeling framework provides a basis for exploring continual behavioral adaptation and the underlying computational mechanisms in the brain network.
We designed a navigation survival task in a dynamic environment in which a cognitive agent was required to survive through predictive interoceptive (energy) regulation. In the environment of the simulated square space (\left[-1.0,1.0\right]\times \left[-1.0,1.0\right]), food was placed at a random position (Fig. 1a). The agent’s energy state increased, if the agent approached it. In a 100 timestep cycle, the nutritional value of the food gradually decreased, and the food reappeared at another random position (Fig. 1b). It roughly mimics the changes in locations and the amount of food in the natural environment. In addition, a cue was placed at a fixed position (0.0, 0.4) that temporarily informed the agent of the x-y coordinate of the current food position if the agent approached it. This cue setting enables to consider behavior adaptation in an uncertain environment38.
Fig. 1: Constraints on environment and sensation.
a Environmental setup. In the environment of simulated square space (\left[-\mathrm{1.0,1.0}\right]\times \left[-\mathrm{1.0,1.0}\right]), a food is placed at a random position, and a cue is placed at a fixed position (0.0, 0.4). Unusual interoceptive state causes large sensory (information) uncertainty, and movements lead to energy expenditure. b Timeseries of proprioceptive state, exteroceptive state, interoceptive state, and food nutrition during random-exploration learning. Note that the agent does not sense the current food position or food nutrition directly. Each sensation takes values between -1.0 and 1.0, corresponding to the range of neural network output. The first 200 steps of learning data (100,000 steps) are shown.
At each time step, the agent received two-dimensional x-y coordinates of the body position, a two-dimensional cue signal, and a one-dimensional energy state, which were regarded as proprioceptive, exteroceptive, and interoceptive sensations, respectively. Each sensation took a value between −1.0 and 1.0, corresponding to the range of the neural network output. The exteroceptive state is meaningful only when the agent is near the cue, taking the value (−0.8, −0.8) when it is not. If the food is located at (−0.8, -0.8), the agent can deduce this because the exteroceptive state remains unchanged while near the cue. As a physiological constraint of the agent, a low or high interoceptive state (negative or positive deviations from 0.0 value) causes large sensory (information) uncertainty implemented as Gaussian noise added to all sensory modalities (Fig. 1a). We assume that this corresponds to the effects of deviations from homeostasis or an increase in physiological entropy due to excessively low or high interoceptive states. We regard the agent as dead when the interoceptive state exceeds 1.0 or drops below −1.0. In addition, we assume that the movements of the agent lead to energy expenditure (a decrease in the interoceptive state).
For behavioral allostasis, the agent must have an internal predictive model of the environment that represents how multimodal sensations and their uncertainties are changed by internal and external causes. Here, we propose the multimodal Bayesian homeostatic recurrent neural network (MBH-RNN), which integrates multimodal sensorimotor processing and homeostatic processing based on variational Bayes (Fig. 2a). The MBH-RNN has a structural and temporal hierarchy consisting of sensorimotor modules (lower-perceptual modules distributed for each modality and multimodal-associative modules) and homeostatic modules (unexpected-uncertainty-cause module and higher-cognitive module). Sensorimotor modules process information from each sensory modality and its associations. They may be associated with the functionality of brain regions such as sensory areas and the insula cortex39. In contrast, the higher-cognitive module integrates multimodal information with uncertainty (homeostatic) information, which may be analogous to cognitive control regions in the brain, such as the prefrontal and anterior cingulate cortices, which are thought to control a wide range of motivational, goal-directed, and uncertainty-related behaviors40,41. Moreover, the unexpected-uncertainty-cause module is primarily involved in the perception of the cause of uncertainty, including unexpected causes, and generates predictions about sensory uncertainty related to all sensory modalities. The function of this module is essential for computing the predicted future sensory entropy related with the homeostatic context and driving feeding behavior through interactions with the higher-cognitive module. We hypothesize that the unexpected-uncertainty-cause module may be associated with brain regions, such as the amygdala and basal forebrain, which are thought to represent a range of uncertainty information, including unexpected uncertainty, and modulate feeding behavior42,43,44,45,46.
Fig. 2: Multimodal Bayesian homeostatic recurrent neural network.
a Structure and information processing of multimodal Bayesian homeostatic recurrent neural network (MBH-RNN). Each module (m) is based on a predictive-coding-inspired variational RNN (PV-RNN). Prior and posterior distributions of latent state are represented as single Gaussian distributions for clarity, although each of them is a multivariate Gaussian distribution. NLL: negative log-likelihood. KLD: Kullback-Leibler divergence. b An example of the effect of ablating a latent variable in the multimodal-associative module on generations of predictions about proprioceptive state, exteroceptive state, interoceptive state, and sensory uncertainty. c Probability of trained latent variables representing either proprioceptive, exteroceptive, interoceptive, or sensory uncertainty information. A simultaneous display of colors or lines means the representation of multiple pieces of information. Proprio proprioceptive module, Extero exteroceptive module, Intero interoceptive module, Multi multimodal-associative module, Uncer unexpected-uncertainty-cause module, Cog higher-cognitive module.
Each module of the MBH-RNN is based on a predictive-coding-inspired variational RNN (PV-RNN)47,48. A brief explanation of the information processing in the MBH-RNN is as below (see Fig. S1 and the “Methods” section for details). At each time step (t), the latent states ({{\boldsymbol{z}}}{t}^{\left(m\right)}) of the MBH-RNN represent Bayesian beliefs regarding the causes of sensations or their uncertainty as (multivariate) Gaussian distributions individually assigned to each module (m). Based on the Bayesian brain hypothesis10, each latent state has prior and posterior probability distributions that correspond to the estimated hidden causes before and after observing the current sensations, respectively. Based on the latent states, the MBH-RNN generates top-down predictions about the mean ({\hat{{\boldsymbol{x}}}}{t}) and standard deviation (uncertainty) ({\widehat{{\boldsymbol{\sigma }}}}{x,t}) of sensations ({{\boldsymbol{x}}}{t}) as a Gaussian distribution. Here, the deterministic recurrent states ({{\boldsymbol{d}}}{t}^{\left(m\right)}) transform latent states into sensory predictions via synaptic connections that represent the dynamic relationships between sensations and their causes. The MBH-RNN uses a multiple timescale RNN (MT-RNN) as the transformation function, which represents the temporal hierarchy by a multiple timescale property in neural activation, inspired by the biological brain26. Specifically, higher-level modules (e.g., multimodal-associative module) have slower neural dynamics than lower-level modules have (e.g., lower-perceptual modules). Importantly, ({\widehat{{\boldsymbol{x}}}}{t}) and ({\widehat{{\boldsymbol{\sigma }}}}{x,t}) are generated through different network pathways, where the higher-cognitive module serves as the information hub. This distinction of the prediction pathways enables us to dissociate the sensorimotor processing from homeostatic processing and assures that the predicted sensory uncertainty ({\hat{{\boldsymbol{\sigma }}}}{x,t}) is the uncertainty caused by homeostatic context that affects all sensory modalities.
The cost function of the MBH-RNN is variational free-energy (VFE) that is introduced as a tractable quantity within variational Bayes that bounds the surprisal (or the negative log model evidence) for sensations48 (see “Methods” for the derivation).
$${{\rm{VFE}}}{{\rm{t}}}=\underbrace{\frac{{({{\boldsymbol{x}}}{{{t}}}-{{\widehat{\boldsymbol{x}}}}{{{t}}})}^{2}}{2{{\widehat{\boldsymbol{\sigma }}}}{{{x}},{{t}}}^{2}}+\frac{{\rm{l}}{\rm{n}}(2\pi {\widehat{{\boldsymbol{\sigma }}}}{x,t}^{2})}{2}}{{\rm{Negative}},{\rm{accuracy}}}+\underbrace{\mathop{\sum}\limits_{m}{D}{KL}[q({{\boldsymbol{z}}}{t}^{(m)}|{{\boldsymbol{e}}}{t:T})||p({{\boldsymbol{z}}}{t}^{(m)}|{{\boldsymbol{d}}}{t-1}^{(m)})]}{{\rm{Complexity}}}$$
(1)
The first term (negative accuracy term) of the VFE describes the negative log-likelihood (NLL) or precision-weighted sensory prediction error. The second term (complexity term) describes Kullback-Leibler divergence (KLD) between the posteriors ({\boldsymbol{q}}({{\boldsymbol{z}}}{t}^{\left(m\right)}|{{\boldsymbol{e}}}{t:T})) and priors ({\boldsymbol{p}}({{\boldsymbol{z}}}{t}^{\left(m\right)}|{{\boldsymbol{d}}}{t-1}^{\left(m\right)})). The priors are generated from prior experience through previous deterministic recurrent states ({{\boldsymbol{d}}}{t-1}^{\left(m\right)}) whereas the posteriors are determined by the backpropagated errors ({{\boldsymbol{e}}}{t:T}) ((T): the last time step of the time window).
The learning of the MBH-RNN was performed using data acquired through random exploration of the environment for 100 food nutrition cycles (10,000 time steps). Figure 1b shows the time series of the proprioceptive, exteroceptive, and interoceptive sensations during random exploration. The learning data did not contain any organized movements, and the proprioception data for random movements were generated automatically using a predefined algorithm. For convenience, in the learning phase, the agent was assumed to be rescued when it was near death, in which the interoceptive state was constrained so that it could not exceed 1.0 or drop below −1.0. The MBH-RNN learned to reconstruct the experienced sensations ({{\boldsymbol{x}}}{1:T}) ((T) = 10,000) by repeating the prediction generations and parameter updates within time steps (1:T) 100,000 times offline (Fig. S1). In the learning process, posteriors at each time step ({{\boldsymbol{z}}}{q,1:T}) and synaptic weights ({\boldsymbol{\omega }}) are updated through the gradient descent method for the accumulated VFE over time steps, where the partial derivative of the VFE with respect to each parameter is calculated by the back propagation through time (BPTT) algorithm49.
We investigated how latent states in the trained MBH-RNN represented multimodal information. It was analyzed through an ablation study by removing each latent variable one-by-one and evaluating the influence on the generations of ({\hat{{\boldsymbol{x}}}}{t}) and ({\hat{{\boldsymbol{\sigma }}}}{x,t}) (see “Methods” for the evaluation). Figure 2b shows an example of the effect of ablating a latent variable in the multimodal-associative module. The ablation of the latent variable impacted the generation of predictions about the proprioceptive and interoceptive states (the solid red lines deviate from the black dashed lines), but not those about the exteroceptive state and sensory uncertainty. This suggests that the ablated latent variable represents bimodal proprioceptive and interoceptive information. Based on the ablation analysis for all latent variables of 10 trained MBH-RNNs, we categorized the latent variables depending on the information they represented. As shown in Fig. 2c, each of the lower-perceptual modules developed a corresponding modality-specific representation, whereas the multimodal-associative module developed bimodal and trimodal latent variables and represented interoceptive information and its integration with proprioceptive and exteroceptive information. However, the unexpected-uncertainty-cause module was only involved with sensory uncertainty predictions, while the higher-cognitive module represented the integration of multimodal information and sensory uncertainty. This suggests that the unexpected-uncertainty-cause module represents the estimated cause of sensory uncertainty that is independent of the internal and external environmental states, whereas the higher-cognitive module represents the internal and external environmental causes of sensory uncertainty. Thus, we confirmed that the MBH-RNN can be used to develop a hierarchical multimodal predictive model of the environment.
Using the trained MBH-RNNs, we analyzed the autonomous behavior generation of the agent in a dynamic environment, in which the random changes in food position in a 100 timestep cycle were different from the learning experience. Ten test trials were performed by each of the 10 trained MBH-RNNs. Autonomous behavior generation was implemented as a simultaneous operation of perception, action generation, and goal modulation based on VFE minimization from past to future (Figs. 3a and S2). Specifically, at the current sensorimotor time step ({t}{c}) ((\ge! 0)), the MBH-RNN, with fixed synaptic weights, repeats prediction generations and posterior updates within time steps from ({t}{c}-{{win}}{p}+1) to ({t}{c}+{{win}}{f}) for 200 times in an online manner. The length of the past time window is ({{win}}{p}=10)(or win**p = t**c + 1 if t**c < 10) while that of the future time window is ({{win}}{f}=200). The cost function in the past time window is the accumulated VFE of the past (VFEP), which is calculated based on the observed sensations ({{\boldsymbol{x}}}{{t}{c}-{{win}}{p}+1:{t}_{c}}). This drives the postdictive perceptual inference of internal and external states. In addition, for future predictive processing, we extend the VFE into the future time window while considering that sensations in the future are unknown variables. Inspired by a previous mathematical study50, we derive a form of the VFE of the future (VFEF) by evaluating VFE under the expectation of predictions about future sensations (that bounds the expected surprisal, or the expected negative log model evidence). This changes the NLL term of the VFE into the predicted conditional (Shannon) entropy of sensation (see “Methods” for derivation).
$${{\mathrm{VFEF}}}{t}=\underbrace{\frac{1}{2}+\frac{{\mathrm{ln}}\left(2{{\uppi }}{{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}\right)}{2}}{{\mathrm{Predicted}},{\mathrm{sensory}},{\mathrm{entropy}}}+\underbrace{\mathop{\sum }\limits{m}{D}{KL}\left[q({{\boldsymbol{z}}}{t}^{\left(m\right)}|{{\boldsymbol{e}}}{t:{t}{c}+wi{n}{f}})||p({{\boldsymbol{z}}}{t}^{\left(m\right)}|{{\boldsymbol{d}}}{t-1}^{\left(m\right)})\right]}{{\mathrm{Predicted}},{\mathrm{complexity}}}$$
(2)
Here, ({{\boldsymbol{e}}}{t:{t}{c}+{{win}}{f}}) indicates the back-propagated errors calculated using the BPTT algorithm, which modulates the posteriors (({t}{c}+{{win}}{f}) is the last time step of the future time window). This derivation of the VFEF naturally describes the meta-goal of the proposed Bayesian allostasis (the minimization of entropy). At each sensorimotor time step ({t}{c}), the MBH-RNN updates the posteriors ({{\boldsymbol{z}}}{q,{t}{c}-{{win}}{p}+1:{t}{c}+{{win}}{f}}) to minimize the accumulated free-energy ({F}{{allostasis}}) through BPTT.
$${F}{{allostasis}}=\mathop{\sum }\limits{t={t}{c}-{{win}}{p}+1}^{{t}{c}}{{\rm{VFE}}}{t}+\mathop{\sum }\limits_{t={t}{c}+1}^{{t}{c}+{{win}}{f}}{{\rm{VFEF}}}{t}$$
(3)
An important aspect of the posterior updates through the back propagation is that the past posteriors are determined through interactions among previous experiences, sensory prediction errors, and goal-directed errors (Fig. 3a), in which we regard future beliefs as future “goal” that is adapted towards minimizing predicted future sensory entropy. The action generation of the agent is performed via the realization of proprioceptive predictions generated at the current sensorimotor time step ({t}_{c}) before the posterior updates using a proportional-integral-derivative controller. We regard sequential future predictions about proprioceptive, exteroceptive, and interoceptive sensations as “planning.”
Fig. 3: Autonomous Bayesian allostasis model.
a Online inference of latent states is performed via minimization of variational free energy from the past to the future. ({t}{c}), ({{win}}{p}), and ({{win}}{f}) indicate the current sensorimotor time step, the length of past time window, and the length of the future time window, respectively (({{win}}{p}=10,{{win}}{f}=200).) Action is generated via a proportional-integral-derivative (PID) controller based on proprioceptive prediction at the current sensorimotor time step ({t}{c}). Latent states in the past were determined through interaction among previous experiences, sensory prediction errors, and goal-directed errors. KLD: Kullback-Leibler divergence. b An example of timeseries for proprioceptive state (proprio), exteroceptive state (extero), interoceptive state (intero), and predicted sensory uncertainty (uncer) during an autonomous survival task. The estimated sensory sigma is the average value over all sensations. The white-colored background indicates the time window of network predictive processing from the past to the future. In the future time window, the proprioceptive, exteroceptive, and interoceptive states are predicted values (planning), not real sensations. See Fig. S3 for the corresponding timeseries of latent states.
Figure 3b shows an example of how the agent achieved continual behavioral adaptation for homeostasis in the test trial (see also Fig. S3 and Video S1). At ({t}{c}=90) in Fig. 3b, the agent kept the interoceptive state at ~0.0 (appropriate level) by previously moving for food intake. The MBH-RNN planned to continue moving (see predicted proprioceptive states in future planning). In addition, the predicted cue information (exteroceptive state) in the future changed, implying that the agent predicted that the food position would change in the future. In response to the adequate interoceptive state, the agent rested for a while around time steps from (90) to (220) (see the timeseries at ({t}{c}=220) in Fig. 3b). At ({t}{c}=220), the agent planned to restart its movements while predicting increased sensory uncertainty at future time steps of ~400. At ({t}{c}=500), the agent experienced a reduced interoceptive level and large sensory uncertainty. Future planning was generated to increase the interoceptive state and resolve deviations from homeostasis. As shown in Fig. S3, the unexpected-uncertainty-cause module adapted the posterior in the face of actual danger at ({t}{c}=500), although it did not change its latent state when predicting large sensory uncertainty in the future (e.g., ({t}{c}=90)). This suggests that the posterior adjustment of the unexpected-uncertainty-cause module may represent the perception of uncertainty that could not be expected in advance. Eventually, by ({t}_{c}=900), the agent successfully survived the danger by actively obtaining food and resting again. These results show that, based on the internal predictive model, the MBH-RNN can achieve autonomous behavioral allostasis through mode switching between moving (feeding) and resting.
To analyze the mechanisms underlying behavior switching between moving and resting, we observed the behavior switching and averaged the detected network behaviors. Figure 4 shows the changes in the posterior distributions based on the values of the 25 steps before the start of movement or rest (see “Methods”). The changes of posterior sigma (standard deviation) were calculated as differences from the corresponding states at the time step (-25) (which can be positive or negative), while the changes of posterior mean were calculated as absolute differences from the corresponding states at the time step (-25) because positive or negative sign of the change of mean states can vary depending on the different trained networks. The components of the VFEP (NLL and KLD terms) as well as the predicted sensory uncertainty, VFEF, and interoceptive state were also plotted. The network behaviors displayed as solid lines in Fig. 4 were analyzed using data after the end of the posterior updating process throughout the trial, meaning that the posteriors at time step (t) are the postdictive values at the tail of the past time window when the sensorimotor time step is ({t}{c}=t+{{win}}{p}-1) (({{win}}{p}=10)). For reference, we plotted the immediate network behaviors recorded immediately after posterior updating at the sensorimotor time step (({t=t}{c})) as dashed lines. The postdictive posteriors displayed as solid lines at time step (t) and the immediate posteriors displayed as dashed lines at time step (t+{{win}}{p}-1) are the values recorded at the same sensorimotor time step ({t}{c}=t+{{win}}_{p}-1). In the following analyses, we focus primarily on postdictive network behavior at the tail of the past time window. This focus arises because the postdictive posteriors reflect the general network behavior across the time window, influencing the generation of forward predictions and receiving backward error signals throughout this period. Additionally, in the immediate network behaviors in Fig. 4a, b, the KLDs between the posteriors and priors were generally small in all modules throughout the time steps, where the priors were determined by past posteriors, including those at the tail of the past time window. Thus, we conclude that immediate network behaviors are largely controlled by postdictive posteriors.
Fig. 4: Autonomous goal-directed and sensory-focused mode switching.
a Changes of posterior distributions, predicted sensory uncertainty, interoceptive state, components of variational free-energy, i.e., negative log-likelihood (NLL) and Kullback-Leibler divergence (KLD), and variational free-energy of the future (VFEF) before and after the switching from sensory-focused rest to goal-directed movement. The blue arrow indicates earlier posterior adjustments during rest in the higher-cognitive module. The red arrows indicate quick posterior changes and increased KLD signals in the higher-cognitive and interoceptive modules that triggered the goal-directed movement. The red vertical dashed line indicates increased NLL and KLDs in the multimodal-associative, proprioceptive, and exteroceptive modules but decreased KLDs in the higher-cognitive and interoceptive modules with a strong postdictive posterior in the higher-cognitive module. b The changes in the posterior distributions, predicted sensory uncertainty, interoceptive state, components of variational free-energy (NLL and KLD), and VFEF before and after switching from goal-directed movement to sensory-focused rest. The red vertical dashed line indicates the moment when the sigma of the postdictive posterior in each module started to increase. In (a, b), the solid lines indicate the postdictive network behaviors (({t}{c}=t+{{win}}{p}-1)), and the dashed lines indicate the immediate network behaviors (({t}{c}=t)) (({t}{c}): sensorimotor time step, ({{win}}{p}): the length of past time window). The posterior distributions and KLDs are averaged over all dimensions of latent variables. The NLL and predicted sensory uncertainty are averaged over all sensations. The VFEF at time step (t) (({t}{c}=t)) is the value accumulated over time steps within future time window ({t}{c}+1:{t}{c}+{{win}}{f}) (({{win}}{f}): the length of future time window). These values are also averaged over all sequences extracted from timeseries of 10 test trials in each network. Line shadings represent the standard error over 10 different trained networks. Uncer: predicted sensory uncertainty. Proprio: proprioceptive module. Extero: exteroceptive module. Intero: interoceptive module.
Figure 4a shows the behavior switching from resting to moving. First, before the start of movement, the postdictive posterior (solid line) in the higher-cognitive module was slightly adjusted around (t=-)18 as indicated by the blue arrow (see also the slight adjustment of the immediate posterior displayed as a dashed line in that module at around the corresponding sensorimotor time step ({t}{c}=-9)). During rest, the levels of NLL (sensory prediction errors) were very small because of the limited interaction with the environment, suggesting that earlier posterior adjustments may result from goal-directed errors for food intake (via the VFEF). Then, right before the start of moving, quick adjustments of postdictive posteriors in the higher-cognitive and lower interoceptive modules were observed from around (t=-9), with increases of KLDs in the corresponding modules. These postdictive posterior adjustments may be ascribed to the gradual increase in VFEF after the corresponding sensorimotor time step ({t}{c}=0). The start of the movement for food intake was reflected in the change in the posterior of the proprioceptive module and an increased interoceptive state after (t=0). During movement, the postdictive posteriors in the higher-cognitive and interoceptive modules were maintained, and the sigma of the postdictive posterior in the higher-cognitive module decreased. The time step of the posterior sigma in the higher-cognitive module hit the bottom (around (t=10)), coinciding with the moment when the VFEF reached its peak (around the corresponding sensorimotor step ({t}_{c}=19)). This suggests that large VFEF formed a strong (postdictive) belief with a low estimated sigma. Importantly, the movement and resultant interaction with the environment after (t=0) led to increases in the NLL and KLDs in the multimodal-associative, proprioceptive, and exteroceptive modules, whereas the KLDs in the higher-cognitive and interoceptive modules decreased. In other words, the MBH-RNN modulates which components of the VFE should be minimized to realize future homeostatic and interoceptive goals. Mode switching from resting to moving may involve the dynamic regulation of the strength (precision) of hierarchical latent states, in which the higher-cognitive and interoceptive modules maintain strong postdictive posteriors, ignoring sensory prediction errors during movement. As such, the emergence of strong beliefs or “intention” enabled movement generation at the expense of large sensory prediction errors.
In contrast, Fig. 4b shows the behavior switching from moving to resting. In this case, the interoceptive state was regulated to an appropriate level before the agent began to rest. At the moment of the start of resting, the sigma of the postdictive posterior in each module was generally increased after around (t=-10) (the corresponding sensorimotor step is ({t}_{c}=-1)), as highlighted by the vertical red line. The start of rest was reflected in the invariant posterior position of the proprioceptive module after (t=0). In addition, the rest greatly decreased the NLL, which may be ascribed to a reduction in the interaction with the environment. These observations suggest that the weak postdictive posteriors with high estimated sigma served as sensory-focused “recognition” of internal and external states that were adapted towards minimizing sensory prediction errors, leading to the rest. In summary, Fig. 4a, b illustrate how autonomous switching between goal-directed and sensory-focused modes emerges through minimization of VFE from the past to the future.
To clarify general network behavior underlying the switching between sensory-focused and goal-directed modes, we investigated the dynamics of VFEP and VFEF during the mode switching. Figure 5 plots pairs of VFEP and VFEF as a two-dimensional space during the mode switching shown in Fig. 4. Figure 5a shows the VFE dynamics during the switching from sensory-focused mode to goal-directed mode (Fig. 4a). Around time steps from -25 to 0 during sensory-focused mode, VFEP and VFEF stably remained at low levels, in which all components of VFEP were minimized (Fig. 4a). Then, VFEP and VFEF gradually increased around time steps from 0 to 10, indicating that the increased VFEF formed strong goals in the future that drive goal-directed movement. Specifically, Fig. 4a shows that the increased VFEF led to generations of large KLD signals in the higher-cognitive module and interoceptive module because of the gap between the goal and the real. The goal-directed movement was generated towards minimizing the increased KLDs in the higher-cognitive module and interoceptive module, corresponding to the realization of strong beliefs (“intention”) emerging in the higher-cognitive module and interoceptive module that enables the actual movement generations via changing proprioceptive predictions. However, in the process of goal-directed movement, the resulting interaction with the environments generated large sensory prediction errors (negative log likelihood) and KLD signals in the multimodal-associative, exteroceptive, and proprioceptive modules, and the VFEP and VFEF peaked at around time step 10. Then, around time steps from 10 to 24, VFEP was reduced after successfully realizing “intention” by goal-directed movement. These results indicate that generations of large VFEF may underlie the switching from sensory-focused mode to goal-directed mode, in which the inevitable risk of perdition error generations required the mode switching and the emergence of strong intention for realizing the goals. Notably, around time steps from 0 to 10, both VFEP and VFEF increased, meaning that the mode switching does not merely occur for minimizing the sum of VFEP and VFEF. Rather, the mode switching was involved with the qualitative shift of VFE minimization itself, or the modulation of which components of VFE the agent minimized.
Fig. 5: Dynamics of variational free-energy during mode switching.
a The dynamic of variational free-energy of the past (VFEP) and variational free-energy of the future (VFEF) during the switching from sensory-focused mode to goal-directed mode shown in Fig. 4a. b The dynamic of VFEP and VFEF during the switching from goal-directed mode sensory-focused mode shown in Fig. 4b. In (a, b), the VFEP and VFEF are the values accumulated over time steps within past time window and future time window, respectively. The crosses and circles indicate sensory-focused mode and goal-directed mode, respectively.
In addition, Fig. 5b shows the VFE dynamics during the switching from goal-directed mode to sensory-focused mode shown in Fig. 4b. In this case, VFEF and VFEP alternately decreased while gradually converging to a low and stable free-energy state. The mode switching was involved with weakened beliefs (Fig. 4b). This indicates that the reduction of VFEF and weakened internal beliefs led to the gradual shift from goal-directed to sensory-focused modes, resulting in rest. As such, the dynamic of VFEF and VFEP characterized the cognitive mode switching of the agent, where the components of the VFE agent minimized were internally modulated through the dynamic interaction among the neural network, body, and the environment.
We investigated whether the Bayesian allostasis model could explain the increase in food intake in preparation for food shortage periods, as observed in biological cognitive agents1,4,5. For clarity, we compared the proposed allostasis model with a setpoint model (control condition) that had an explicit homeostatic setpoint (target value) of the future interoceptive state set to 0.0. In the setpoint model, the predicted sensory entropy term of the VFEF is replaced by a NLL for the interoceptive sensation, given a fixed setpoint. Equivalent interoceptive prediction errors between the predicted future interoceptive sensation and the setpoint have historically been explored to explain goal-directed homeostatic processing within various contexts, including cybernetics, active inference (with the concept of expected free energy), and homeostatic reinforcement learning51,52,53,54. Note that, in the setpoint model, the target interoceptive state (setpoint) is predicted as a probabilistic distribution, and the interoceptive prediction error (calculated as NLL) is weighted by the estimated precision (inverse of standard deviation) (see “Methods”). We used the same trained MBH-RNNs for both the allostasis and setpoint models, meaning that the difference was only in the autonomous survival process after learning.
We first evaluated the difference in the survival time steps between the allostasis and setpoint models (Fig. 6a). A paired t-test revealed that the allostasis model survived significantly longer than the setpoint model did ((t\left(9\right)=3.38,p=0.0082)). To analyze the general characteristics of interoceptive regulation in the survival task, the cycle of food nutrition was divided into five periods: the food recovery period (I), the food gradually decreasing periods (II-IV), and the food shortage period (V) (Fig. 6b). Figure 6c shows the difference between the two models in interoceptive regulation by plotting the changes in the interoceptive state during the food cycle based on the values in period I. The allostasis model increased the interoceptive level until Period III, and the interoceptive level in Period V was similar to that in Period I. Simultaneously, the number of food acquisitions peaked in Period III, following an increased acquisition of cue information in Period II (Fig. 6d). It indicates that the allostasis model exhibits predictive interoceptive regulation by capturing the food nutrition cycle and exploring food positions through information gain. However, in the setpoint model, the interoceptive level, as well as the acquisition of food and cue information, monotonically decreased during the food gradually decreasing periods II-IV, and the interoceptive level in Period V was smaller than that in Period I (Fig. 6c, e). Furthermore, the allostasis model showed a smaller amount of energy expenditure throughout the food nutrition cycle than the setpoint model did, which may have increased the effectiveness of energy intake (Fig. 6f). Differences between the allostasis and setpoint models in the regulation of energy intake and expenditure may result in significant differences in survival performance.
Fig. 6: Predictive interoceptive regulation.
a Box-and-whiskers plots showing the difference in autonomous survival time step between the allostasis model and the setpoint model (control condition). Each plot is an average of the 10 test trials by each of 10 trained networks (center line, median; cross, mean; box limits, upper and lower quartiles; whiskers, (\times 1.5) interquartile range). (* * p < 0.01.) b the food nutrition cycle is equally divided into five periods: from food recovery period (I) to food gradually decreasing periods (II–IV), and food shortage period (V). c Change of interoceptive state based on the value in the period I in each model. d, e Changes in the amounts of food and cue acquisitions per time step in the allostasis model and the setpoint model. f Changes of energy intake and energy expenditure per time step in the allostasis model and the setpoint model. In (c, d, e, f), the gray shadow represents the shape of the cycle of averaged food nutrition for reference, and error bars represent the standard error.
We propose that an autonomous cognitive agent does not just minimize sensory prediction errors but also minimizes predicted future sensory entropy. We demonstrated that this Bayesian allostasis model can be derived by extending the variational Bayesian framework to future predictive processing. We validated the Bayesian allostasis model by simulating a structurally and temporally hierarchical MBH-RNN that exhibited self-organizing behavior, switching between goal-directed movement (feeding) and sensory-focused rest, based on the internal predictive model. The continuous mode switching resulted in predictive interoceptive regulation in a dynamic environment, which may explain behavioral allostasis in biological cognitive agents1,4,5. To our knowledge, our model is the first predictive processing model of autonomous behavioral allostasis with goal generation mechanism. Our Bayesian neural network framework opens new avenues for exploring brain information processing to anchor continual behavioral adaptation.
Goal-directed and sensory-focused mode switching were involved in switching the roles of Bayesian beliefs in VFE minimization. During the goal-directed movement, strong postdictive posteriors emerged in the higher-cognitive module and interoceptive module. The strong beliefs were maintained and served as “intention” that enabled movement generations in the face of prediction error generations. On the other hand, in the sensory-focused rest, postdictive posteriors were generally weakened and adapted towards minimizing sensory prediction errors, serving as “recognition” of internal and external states. These results suggest that mode switching correspond to meta-control of which components of the VFE are minimized by the agent. This self-organization of mode switching may explain how “recognition” changes into “intention” and vice versa in a Bayesian brain.
In particular, switching from sensory-focused rest to goal-directed movement may demonstrate how cognitive agents actively interact with the external environment, even if it is unpredictable or surprising. At the moment of the mode switching, generations of large KLD signals (due to goal-directed signals for food intake) changed the higher-level postdictive posteriors into strong beliefs (“intention”), which triggered movements for feeding. Importantly, adjustments of the higher-level postdictive posteriors began during rest before the start of movement (blue arrow in Fig. 4a). This observation might be associated with biological observations of the readiness potential in advance of conscious intention to action55,56,57. More specifically, in our model, the earlier readiness posterior adjustments and the subsequent generation of the strong beliefs or “intention” might be associated with different internal experiences of the agent in terms of the experienced level of goal-directed signals. This may have computational and phenomenological implications for the emergence of intentions.
Our Bayesian allostasis model suggests that the homeostatic behaviors of cognitive agents may primarily be guided by the aspiration of a certain computation concerning second-order prediction (low predicted future sensory entropy) rather than by maintaining physical states at the setpoints. We would like to emphasize that in our model, goal-directed movement and sensory-focused rest were not simply optimized through learning. Rather, they emerged as behavioral patterns autonomously generated by the agent itself through allostatic process. Our primary focus was on the agent’s autonomy—how it constructs its own behaviors beyond mere optimization through learning. This perspective is distinct from models where behavior is explicitly shaped by reinforcement or genetic algorithms. For example, a previous study showed that a genetic-learning-based neural network model exhibited behavioral attractor transitions depending on its energy level (interoceptive state) through an evolved neurocontroller58. This self-organized process was ultimately optimized via genetic learning for maximizing predefined fitness, as determined by factors such as survival duration and collision avoidance. In their model, both exploratory movements associated with low energy level and reduced movement associated with appropriate energy level served the same predefined goal—maximizing the fitness. However, our model introduces a crucial difference: goal-directed movement and sensory-focused rest were not simply different strategies toward a single fixed goal (reward or fitness). Instead, the components of the VFE agent minimized were dynamically regulated, as explained above. In addition, our model did not just react to the current energy level but predictively regulated the energy level by modulating future goals and their precision. This allostatic interoceptive control has not been thoroughly explored in previous studies. From this perspective, the previously proposed setpoint-driven models for homeostasis and allostasis (explained as adjustments of setpoints)51,52,53,54 can be considered a specific case of the goal-directed mode in our model. In addition, our model explains how strong beliefs regarded as “intention” (or “setpoint” in the setpoint-driven models) can emerge in the hierarchical neural network. From another perspective, our model integrates the concept of reward in reinforcement learning into the Bayesian brain hypothesis by considering interoceptive predictive processing in the face of increased physiological entropy owing to irreversible physical processes within the body. In biological cognitive agents, the optimization of homeostatic behavior through learning, as proposed in previous studies, and allostatic behavioral adaptation with goal generation, as proposed in our study, may be integrated or may vary depending on the species.
Autonomous mode switching remains a largely unexplored phenomenon in computational modeling studies of the brain. However, it may be a key for studying “intention” and “self” in the computational machine. Biological and theoretical studies have suggested that different kinds of sensory suppression processes, such as sensory attenuation and gating, may be associated with self-nonself distinction and a sense of agency25,48. Sensory attenuation refers to fewer perceptual and neural responses to self-generated sensations than to externally generated sensations, whereas sensory gating refers to fewer responses to sensations during movement than during rest59. Sensory attenuation can be explained by self-organized mode switching in the process of VFE minimization, depending on the inferred causality of sensations (self or nonself)22,48. Specifically, in the mechanism of sensory attenuation through VFE minimization, constantly strong Bayesian beliefs at a lower perceptual level were observed during the self-generated context. This led to a higher weight of the corresponding KLD terms at the lower perceptual level than those at the higher network levels, resulting in reduced prediction-error-induced responses at the lower perceptual level. However, the current model study showed that emergence of strong beliefs at the interoceptive and higher-cognitive modules enabled movement generation by ignoring sensory prediction errors, where the corresponding KLD terms (associated with the strong beliefs) of VFE were weighted higher than NLL term as well as KLD terms in other modules. Ignoring sensory prediction errors due to strong beliefs in the goal-directed mode may explain sensory gating during movement. These observations provide mechanistic implications for the difference between sensory attenuation and gating, although both can be explained as autonomous mode switching in VFE minimization. The self-organization of autonomous mode switching may pave the way for studying the computational structures underlying autonomy in terms of perception and behavior.
The concept of VFE minimization has been widely studied as a unifying framework for understanding cognition, self-organization, and biological systems. Recent literature has explored its theoretical underpinnings, applications, and limitations across various domains, including theoretical neuroscience, physics, artificial intelligence, and cognitive science. In our model, VFE minimization is used as a computational principle driving cognitive processing such as learning, perception, and goal generation within the RNN. Notably, we do not assume that temporal change of neural activity directly minimizes VFE, as seen in standard formulations of FEP that attempt to uniformly explain state changes of self-organizing system scoping from neural activity to sentient behavior as random dynamical system described in terms of a Langevin stochastic differential equation based on the premise of convergence to a non-equilibrium steady state12,60,61. Instead, our framework treats VFE minimization as a computational principle guiding predictive processing of the RNN, in which the neural dynamics is generated through forward calculation of the RNN. This implementation allows us to incorporate complex neural dynamics, multimodal sensorimotor learning, and dynamic brain-body-environment interaction, making our approach aligned with enactivist perspectives62,63. This also helps reconcile our framework with critiques that argue standard FEP formulations may not generalize well to complex cognitive and biological phenomena because of the assumptions such as non-equilibrium steady states, Markov blankets, and solenoidal flow constraints (key components of FEP)64,65,66.
Furthermore, there is another perspective of free-energy minimization that attempts to evaluate state changes of neural network models in terms of free-energy minimization. For example, a previous study defined a non-variational free-energy (different mathematical quantity from VFE proposed in FEP) in terms of the relative difference between the energy and entropy and evaluated the non-variational free-energy on a spike-timing-dependent-plasticity-based spiking neural network model for exploring self-organizing neural states in the absence of sensing during a sleep-like transition67,68. Additionally, another study explores a correspondence between a class of cost functions shared for both neural activity and synaptic plasticity in neural network models and VFE minimization69. These approaches that evaluate free energy of neural networks from an observer’s perspective offer insights into how free-energy minimization can be analyzed in biological neural networks. In contrast to these studies, our approach treats VFE as the cost function of the neural agent that guides cognitive processing. This phenomenological perspective of VFE minimization allows us to explain autonomous behavior adaptation and goal generation. The FEP is a broad concept, leaving room for debate regarding its scope of application and which processes or aspects should be assumed to minimize free energy.
Neuroimaging studies suggest that allostatic-interoceptive processing involves a large-scale brain network that is composed of core regions of the salience and default mode networks, such as the insula, anterior cingulate cortex, prefrontal cortex, amygdala, thalamus, and hypothalamus70. Based on the lines of analysis of our trained model (Figs. 2–4 and S3), we speculate about a correspondence between the functional modules of our MBH-RNN and the allostatic-interoceptive system in a biological brain. Regarding the sensorimotor modules of the MBH-RNN, the interoceptive module represented modality-specific interoceptive information, whereas the multimodal-associative module integrated interoceptive information with exteroceptive and proprioceptive information. This hierarchical interoceptive neural representation may be analogous to the hierarchical functional organization of the insular cortex39. In contrast, regarding homeostatic modules, the higher-cognitive module of the MBH-RNN integrated multimodal information with sensory-uncertainty information (homeostatic context) and controlled goal-directed behavior. This functionality may be consistent with biological knowledge about cognitive control regions40,41. In addition, the unexpected-uncertainty-cause module of the MBH-RNN primary represented perception of unexpected cause of uncertainty and generated the top-down prediction about sensory uncertainty that forms the “meta-goal” essential for adapting future goals. Fig. S3 shows that the latent state of the unexpected-uncertainty-cause module did not change when predicting large sensory uncertainty in the future (e.g., ({t}{c}=90) in Fig. S3), but responded to actual increases of sensory uncertainty, or deviations from homeostasis (e.g., ({t}{c}=500) in Fig. S3). The unexpected-uncertainty-cause module might drive behavior for resolving unexpected cause of uncertainty. These characteristics of the unexpected-uncertainty-cause module might be associated with brain regions representing a range of uncertainty information including unexpected and fear-related events such as amygdala and basal forebrain that are also suggested to modulate feeding behavior through connections with other regions42,43,44,45,46. In summary, the self-organizing functional network of the MBH-RNN may provide insights into the functional architecture of the allostatic-interoceptive network71 and related regions in the brain.
As a limitation, we focused only on the cognitive aspect of allostasis, which requires behavioral interactions with the external environment. Further discussions may be required for interoceptive controls via reflex arcs such as cardiovascular properties (e.g., blood pressure), which seem to be involved with predefined reactions in the central nervous system based on “hard-wired” setpoints72. Additionally, future studies should explore the scalability of our model in a physical or simulated robot with high-dimensional sensors, as well as validate the in silico neural network model with a biological body. Our multimodal Bayesian RNN is computationally complex, and its scalability and real-time applicability remain open challenges. Advances in computational technology will help the realization of this direction. In addition, applications of our model into more complex navigation task will also be a potential direction. A previous study proposes a neural-network-based hippocampal navigation model that can navigate in cluttered environments with a set of obstacles by using a combination of grid cell-driven vector navigation, place cell-driven topological navigation, and border cell-driven local obstacle avoidance73. The navigation model achieved object avoidance and flexibility even in sparsely explored environments and various mazes with mechanism of subgoal selection via hippocampal replay events. Combining other approaches such as this framework and reinforcement learning74 with our model can be a potential future direction for achieving flexible allostatic navigation in more realistic cluttered environments. Specifically, in cluttered environments, the agent may need to carefully observe the environment to appropriately adjust subgoals and execute navigation via continuous switching between sensory-focused and goal-directed modes. This remains a future challenge in explaining allostasis in complex environments. Finally, the application of our model to explore the computational mechanisms underlying altered interoceptive processing in psychiatric disorders (e.g., depression, eating disorders, and autism spectrum disorder)75,76,77 and to investigate possible interventions for preventing or improving them will be a potential direction for future studies.
In this section, we explain the settings used to create the sensory data. The radius of the food and cue fields was set to 0.2, which was used to judge the acquisition. The relationship between the real standard deviation of sensory noise ({\sigma }{x,t}) and the interoceptive state ({x}{t,{intero}}) (the effect of deviations from homeostasis) is defined as the summation of sigmoid functions (the graph can be seen in Fig. 1a).
$${\sigma }{x,t}=\gamma \left{\frac{1}{1+{e}^{-20\left(-{x}{t,{intero}}-0.5\right)}}+\frac{1}{1+{e}^{-20\left({x}_{t,{intero}}-0.5\right)}}\right}$$
(4)
Here, (\gamma =\sqrt{0.1}) determines the maximum of ({\sigma }{x,t}). In biological agents, the relationship between ({\sigma }{x,t}) and ({x}_{t,{intero}}) may be determined by the bodily characteristics of the agent, such as muscles and metabolism, which are regulated through bodily changes (allodynamics). Although the relationship was fixed in our experiment, we confirmed that allostatic interoceptive regulation was observed even in a smaller (\gamma) setting ((\gamma =\sqrt{0.001})), showing the robustness of our allostasis model (Fig. S4).
The energy expenditure (decrease in the interoceptive state) ({{\rm{EE}}}{t}) owing to movement is calculated at each time step based on the moving distance ({{\rm{MD}}}{t}) (difference between the current agent’s position and the previous position) as ({{\rm{EE}}}{t}=k(0.1+{{\rm{MD}}}{t})). Here, (k=0.013) controls the level of expenditure and is heuristically determined such that the average interoceptive state in the learning data is nearly 0.
In this section, we describe the mathematical details of top-down prediction generation and bottom-up parameter updates using the MBH-RNN, which is based on PV-RNN48 (Fig. S1).
Prediction generation was performed in a top-down manner using a network hierarchy. The internal state ({h}{t,i}) and output ({d}{t,i}) of the ith deterministic variable (recurrent unit) in each network module at time step (t) ((t\ge 1)) is calculated as.
$${h}{t,i}=\left{\begin{array}{l}\frac{1}{\tau }\left(\mathop{\sum}\nolimits{j\in {I}{\rm{Cd}}}{\omega }{{ij}}{d}{t-1,j}+\mathop{\sum}\nolimits{j\in {I}{\rm{Cz}}}{\omega }{{ij}}{z}{t,j}+{b}{i}\right)+\left(1-\frac{1}{\tau }\right){h}{t-1,i}\left(i\in {I}{\rm{Cd}}\right)\ \frac{1}{\tau }\left(\mathop{\sum}\nolimits_{j\in {I}{\rm{Ud}}}{\omega }{{ij}}{d}{t-1,j}+\mathop{\sum}\nolimits{j\in {I}{\rm{Uz}}}{\omega }{{ij}}{z}{t,j}+\mathop{\sum}\nolimits{j\in {I}{\rm{Cd}}}{\omega }{{ij}}{d}{t,j}+{b}{i}\right)+\left(1-\frac{1}{\tau }\right){h}{t-1,i}(i\in {I}{\text{Ud}})\ \frac{1}{\tau }\left(\mathop{\sum}\nolimits_{j\in {I}{\rm{Ad}}}{\omega }{{ij}}{d}{t-1,j}+\mathop{\sum}\nolimits{j\in {I}{\rm{Az}}}{\omega }{{ij}}{z}{t,j}+\mathop{\sum}\nolimits{j\in {I}{\rm{Cd}}}{\omega }{{ij}}{d}{t,j}+{b}{i}\right)+\left(1-\frac{1}{\tau }\right){h}{t-1,i}(i\in {I}{\rm{Ad}})\ \frac{1}{\tau }\left(\mathop{\sum}\nolimits_{j\in {I}{\rm{Pd}}}{\omega }{{ij}}{d}{t-1,j}+\mathop{\sum}\nolimits{j\in {I}{\rm{Pz}}}{\omega }{{ij}}{z}{t,j}+\mathop{\sum}\nolimits{j\in {I}{\rm{Ad}}}{\omega }{{ij}}{d}{t,j}+{b}{i}\right)+\left(1-\frac{1}{\tau }\right){h}{t-1,i}\left(i\in {I}{\text{Pd}}\right)\ \frac{1}{\tau }\left(\mathop{\sum}\nolimits_{j\in {I}{\rm{Ed}}}{\omega }{{ij}}{d}{t-1,j}+\mathop{\sum}\nolimits{j\in {I}{\rm{Ez}}}{\omega }{{ij}}{z}{t,j}+\mathop{\sum}\nolimits{j\in {I}{\rm{Ad}}}{\omega }{{ij}}{d}{t,j}+{b}{i}\right)+\left(1-\frac{1}{\tau }\right){h}{t-1,i}\left(i\in {I}{\rm{Ed}}\right)\ \frac{1}{\tau }\left(\mathop{\sum}\nolimits_{j\in {I}{\rm{Id}}}{\omega }{{ij}}{d}{t-1,j}+\mathop{\sum}\nolimits{j\in {I}{\rm{Iz}}}{\omega }{{ij}}{z}{t,j}+\mathop{\sum}\nolimits{j\in {I}{\rm{Ad}}}{\omega }{{ij}}{d}{t,j}+{b}{i}\right)+\left(1-\frac{1}{\tau }\right){h}{t-1,i}\left(i\in {I}{\rm{Id}}\right)\end{array}\right.$$
(5)
$${d}{t,i}=\tanh \left({h}{t,i}\right)\left(i\in {I}{\text{Cd}},{I}{\text{Ud}},{I}{\text{Ad}},{{{I}{\text{Pd}},I}{\text{Ed}}I}{\text{Id}}\right)$$
(6)
where ({I}{\text{Cd}}), ({I}{\text{Ud}}), ({I}{\text{Ad}}), ({I}{\text{Pd}}), ({I}{\text{Ed}}), and ({I}{\text{Id}}), are the index sets of deterministic variables in the higher-cognitive module, unexpected-uncertainty-cause module, multimodal-associative module, proprioceptive module, exteroceptive module, and interoceptive module, respectively. ({I}{\text{Cz}}), ({I}{\text{Uz}}), ({I}{\text{Az}}), ({I}{\text{Pz}}), ({I}{\text{Ez}}), and ({I}{\text{Iz}}) are the index sets of probabilistic latent states in the higher-cognitive module, unexpected-uncertainty-cause module, multimodal-associative module, proprioceptive module, exteroceptive module, and interoceptive module, respectively. ({\omega }{{ij}}) is the weight of the synaptic connection from the jth neuron to the ith neuron; ({z}{t,j}) is the output of jth latent states at time step (t); (\tau) is the time constant of the deterministic variable; and ({b}{i}) is the bias of the ith deterministic variable. A deterministic variable with a small time constant (\tau) has a tendency to change its activity rapidly, while that with a large time constant has a tendency to change its activity slowly26. We set the initial internal states of the deterministic variables ({h}{0,i}) ((i\in {I}{\text{Cd}},{I}{\text{Ud}},{I}{\text{Ad}},{{{I}{\text{Pd}},I}{\text{Ed}},I}{\text{Id}})) to 0 (({d}_{0,i}) is also 0).
The probabilistic latent state ({{\boldsymbol{z}}}{t}) in each module is assumed to follow a multivariate Gaussian distribution with a diagonal covariance matrix, meaning ({z}{t,i}) and ({z}{t,j}) are independent ((i,j\in {I}{\text{Cz}},{I}{\text{Uz}},{I}{\text{Az}},{{{I}{\text{Pz}},I}{\text{Ez}},I}{\text{Iz}}\wedge i\ne j)). Here, the mean ({\mu }{p,t,i}) and sigma (standard deviation) ({\sigma }{p,t,i}) of the prior distribution (p({z}{t,i})) in all modules are calculated as below:
$$p\left({z}{t,i}\right)=p\left({z}{t,i}|{d}{t-1,j}\right)={\mathcal{N}}\left({z}{t,i};{\mu }{p,t,i},{\sigma }{p,t,i}\right)$$
(7)
$${\mu }{p,t,i}=\tanh \left(\sum {j}{\omega }{{ij}}{d}{t-1,j}\right)$$
(8)
$${\sigma }{p,t,i}=\exp \left(\sum {j}{\omega }{{ij}}{d}{t-1,j}\right)$$
(9)
Here, prior distributions in each module are calculated from the previous deterministic states of the same module, meaning ((i\in {I}{\text{Cz}}\wedge j\in {I}{\text{Cd}})\vee (i\in {I}{\text{Uz}}\wedge j\in {I}{\text{Ud}})\vee (i\in {I}{\text{Az}}\wedge j\in {I}{\text{Ad}})\vee (i\in {I}{\text{Pz}}\wedge j\in {I}{\text{Pd}})\vee (i\in {I}{\text{Ez}}\wedge j\in {I}{\text{Ed}})\vee (i\in {I}{\text{Iz}}\wedge j\in {I}{\text{Id}})).
The posterior distribution in each module is calculated as,
$$q\left({z}{t,i}|{{\boldsymbol{e}}}{t:T}\right)={\mathcal{N}}\left({z}{t,i};{\mu }{q,t,i},{\sigma }_{q,t,i}\right)$$
(10)
$${\mu }{q,t,i}=\tanh \left({a}{\mu ,t,i}\right)$$
(11)
$${\sigma }{q,t,i}=\exp \left({a}{\sigma ,t,i}\right)$$
(12)
$${z}{t,i}={\mu }{q,t,i}+{\sigma }_{q,t,i}\times \epsilon$$
(13)
Here, (T) represents the last time step of the time window. ({{\boldsymbol{a}}}{t}) is the adaptive internal state of the latent variables representing the posterior distributions. The adaptive variables ({{\boldsymbol{a}}}{t}) are determined by error signals ({{\boldsymbol{e}}}{t:T}) propagated using the BPTT algorithm. The adaptive variables ({{\boldsymbol{a}}}{t}) are initialized by the corresponding initial internal states of the latent variables that represent the prior distributions. In Eq. (13), by sampling (\epsilon) from ({\mathcal{N}}\left(\mathrm{0,1}\right)), the latent state ({z}_{t,i}) is obtained.
Finally, mean predictions about proprioceptive, exteroceptive, and interoceptive sensations were individually generated from the proprioceptive, exteroceptive, and interoceptive modules, respectively, whereas sigma predictions for each sensation were generated from the unexpected-uncertainty-cause module.
$${\hat{x}}{t,i}=\left{\begin{array}{c}\tanh \left(\sum {j\in {I}{\text{Pd}}}{\omega }{{ij}}{d}{t,j}\right)\left(i\in {I}{\text{Po}}\right)\ \tanh \left(\sum {j\in {I}{\text{Ed}}}{\omega }{{ij}}{d}{t,j}\right)\left(i\in {I}{\text{Eo}}\right)\ \tanh \left(\sum {j\in {I}{\text{Id}}}{\omega }{{ij}}{d}{t,j}\right)\left(i\in {I}{\text{Io}}\right)\end{array}\right.$$
(14)
$${\hat{\sigma }}{x,t,i}=\exp \left(\sum {j\in {I}{\text{Ud}}}{\omega }{{ij}}{d}{t,j}\right)\left(i\in {I}{\text{Po}},{I}{\text{Eo}},{I}{\text{Io}}\right)$$
(15)
Here, ({I}{\text{Po}}), ({I}{\text{Eo}}), ({I}_{\text{Io}}) are the index sets of the output units for the proprioceptive, exteroceptive, and interoceptive predictions, respectively.
The cost function of the MBH-RNN is derived based on variational Bayes or the FEP10. Under the FEP, minimizing the surprisal (or the negative log model evidence) for sensations ({\boldsymbol{x}}) over time is the ultimate goal of self-organizing systems. In the learning of the MBH-RNN, the surprisal over all time steps can be written as:
$$-{\rm{l}}{\rm{n}},p({{\boldsymbol{x}}}{1:T})=-{\rm{l}}{\rm{n}}\int p({{\boldsymbol{x}}}{1:T},{{\boldsymbol{z}}}{1:T})d{{\boldsymbol{z}}}{1:T}$$
(16)
$$=-{\mathrm{ln}}\int \mathop{\prod }\limits_{t=1}^{T}\left[p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)p\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{d}}}{t-1}\right)\right]d{{\boldsymbol{z}}}_{1:T}$$
(17)
Here, surprisal cannot be directly evaluated by the agent because it needs to know all hidden states ({\boldsymbol{z}}) of the environment that cause sensations, as described on the right side of the Eq. (16). Here, FEP introduces a tractable quantity, VFE, that bounds the surprisal, and minimization of the surprisal is replaced by minimization of the VFE. The bound of the surprisal in the MBH-RNN can be derived by utilizing Jensen’s inequality for a concave function (f:f(E[x])\ge E[f\left(x\right)]). Then, Eq. (17) can be modified as follows by introducing a dummy distribution for ({{\boldsymbol{z}}}{1}), ({q({\boldsymbol{z}}}{1})).
$$-{\mathrm{ln}}p\left({{\boldsymbol{x}}}{1:T}\right)=-{\mathrm{ln}}{\int{{{\boldsymbol{z}}}{1}}}d{{\boldsymbol{z}}}{1}q\left({{\boldsymbol{z}}}{1}\right){\int{{{\boldsymbol{z}}}{2:T}}}d{{\boldsymbol{z}}}{2:T}\frac{1}{q\left({{\boldsymbol{z}}}{1}\right)}\mathop{\prod }\limits{t=1}^{T}\left[p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)p\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{d}}}{t-1}\right)\right]$$
(18)
$$\le -{\int_{{{\boldsymbol{z}}}{1}}}d{{\boldsymbol{z}}}{1}q\left({{\boldsymbol{z}}}{1}\right){\mathrm{ln}}{\int{{{\boldsymbol{z}}}{2:T}}}d{{\boldsymbol{z}}}{2:T}\frac{1}{q\left({{\boldsymbol{z}}}{1}\right)}\mathop{\prod }\limits{t=1}^{T}\left[p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)p\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{d}}}{t-1}\right)\right]$$
(19)
In Eq. (19), Jensen’s inequality is used as a logarithmic function. The same procedure is performed for (t=2:T).
$$-{\mathrm{ln}},p\left({{\boldsymbol{x}}}{1:T}\right)\le -{\int{{{\boldsymbol{z}}}{1}}}d{{\boldsymbol{z}}}{1}q\left({{\boldsymbol{z}}}{1}\right)\cdots {\int{{{\boldsymbol{z}}}{T}}}d{{\boldsymbol{z}}}{T}q\left({{\boldsymbol{z}}}{T}\right){\mathrm{ln}}\mathop{\prod }\limits{t=1}^{T}\left[\frac{p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)p\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{d}}}{t-1}\right)}{q\left({{\boldsymbol{z}}}_{t}\right)}\right]$$
(20)
$$=-{\int_{{{\boldsymbol{z}}}{1}}}d{{\boldsymbol{z}}}{1}q\left({{\boldsymbol{z}}}{1}\right)\cdots {\int{{{\boldsymbol{z}}}{T}}}d{{\boldsymbol{z}}}{T}q\left({{\boldsymbol{z}}}{T}\right)\mathop{\sum }\limits{t=1}^{T}\left[\underbrace{{\mathrm{ln}},p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)}{\rm{A}}-\underbrace{{\mathrm{ln}}\frac{q\left({{\boldsymbol{z}}}{t}\right)}{p\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{d}}}{t-1}\right)}}_{\rm{B}}\right]$$
(21)
The first term in expression (21) is the expected NLL under (q) given that ({{\boldsymbol{d}}}{t}) depends on ({{\boldsymbol{z}}}{1:t}).
$${\rm{A}}=-{\int_{{{\boldsymbol{z}}}{1}}}d{{\boldsymbol{z}}}{1}q\left({{\boldsymbol{z}}}{1}\right)\cdots {\int{{{\boldsymbol{z}}}{T}}}d{{\boldsymbol{z}}}{T}q\left({{\boldsymbol{z}}}{T}\right)\mathop{\sum }\limits{t=1}^{T}{\mathrm{ln}},p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)$$
(22)
$$=-\mathop{\sum }\limits_{t=1}^{T}{E}{q({{\boldsymbol{z}}}{1:t})}\left[{\mathrm{ln}},p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)\right]$$
(23)
In addition, the second term can be deformed into forms of KLD between (q\left({{\boldsymbol{z}}}{t}\right)) and (p\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{d}}}{t-1}\right)) while ensuring that ({{\boldsymbol{d}}}{0}) is independent of ({{\boldsymbol{z}}}_{1:T}).
$${\rm{B}}={\int_{{{\boldsymbol{z}}}{1}}}d{{\boldsymbol{z}}}{1}q\left({{\boldsymbol{z}}}{1}\right)\cdots {\int{{{\boldsymbol{z}}}{T}}}d{{\boldsymbol{z}}}{T}q\left({{\boldsymbol{z}}}{T}\right)\mathop{\sum }\limits{t=1}^{T}{\mathrm{ln}}\frac{q\left({{\boldsymbol{z}}}{t}\right)}{p\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{d}}}_{t-1}\right)}$$
(24)
$$={\int_{{{\boldsymbol{z}}}{1}}}d{{\boldsymbol{z}}}{1}q\left({{\boldsymbol{z}}}{1}\right){\mathrm{ln}}\frac{q\left({{\boldsymbol{z}}}{1}\right)}{p\left({{\boldsymbol{z}}}{1}|{{\boldsymbol{d}}}{0}\right)}+{\int_{{{\boldsymbol{z}}}{1}}}d{{\boldsymbol{z}}}{1}q\left({{\boldsymbol{z}}}{1}\right)\cdots {\int{{{\boldsymbol{z}}}{T}}}d{{\boldsymbol{z}}}{T}q\left({{\boldsymbol{z}}}{T}\right)\mathop{\sum }\limits{t=2}^{T}{\mathrm{ln}}\frac{q\left({{\boldsymbol{z}}}{t}\right)}{p\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{d}}}_{t-1}\right)}$$
(25)
$$={D}{{KL}}\left[q\left({{\boldsymbol{z}}}{1}\right){||p}\left({{\boldsymbol{z}}}{1},|,{{\boldsymbol{d}}}{0}\right)\right]+\mathop{\sum }\limits_{t=2}^{T}{E}{q({{\boldsymbol{z}}}{1:t-1})}\left[{D}{{KL}}\left[q\left({{\boldsymbol{z}}}{t}\right){||p}\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{d}}}{t-1}\right)\right]\right]$$
(26)
A dummy distribution (q\left({{\boldsymbol{z}}}{t}\right)) can be replaced by the posterior distribution determined by the back-propagated prediction error (q\left({{\boldsymbol{z}}}{t},|,{{\boldsymbol{e}}}_{t:T}\right)). Thus, using Eqs. (23) and (26), the bound of the surprisal can be written as,
$$\begin{array}{l}-{\mathrm{ln}},p\left({{\boldsymbol{x}}}{1:T}\right)\le -\mathop{\sum }\limits{t=1}^{T}{E}{q\left({{\boldsymbol{z}}}{1:t}|{{\boldsymbol{e}}}{1:T}\right)}\left[{\mathrm{ln}}p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)\right]+{D}{{KL}}\left[q\left({{\boldsymbol{z}}}{1}|{{\boldsymbol{e}}}{1:T}\right){||p}\left({{\boldsymbol{z}}}{1}|{{\boldsymbol{d}}}{0}\right)\right]\+\mathop{\sum }\limits_{t=2}^{T}{E}{q\left({{\boldsymbol{z}}}{1:t-1}|{{\boldsymbol{e}}}{1:T}\right)}\left[{D}{{KL}}\left[q\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{e}}}{t:T}\right){||p}\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{d}}}{t-1}\right)\right]\right]\end{array}$$
(27)
In the experiment, we resolved the calculation of the expectations of the NLL and KLD under (q) by considering a single sampling to reduce computational costs.
$$-{\mathrm{ln}},p\left({{\boldsymbol{x}}}{1:T}\right)\le \mathop{\sum }\limits{t=1}^{T}(-{\mathrm{ln}},p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)+{D}{{KL}}\left[q\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{e}}}{t:T}\right){||p}\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{d}}}_{t-1}\right)\right])$$
(28)
Eventually, the cost function (the bound of the surprisal) in the learning ({F}{{learning}}) is the accumulation of variational free-energy ({{\rm{VFE}}}{t}) over time steps.
$${F}{{learning}}=\mathop{\sum }\limits{t=1}^{T}{{\rm{VFE}}}_{t}$$
(29)
$${{\rm{VFE}}}{t}=-{\mathrm{ln}},p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)+\sum {m}{D}{{KL}}\left[q({{\boldsymbol{z}}}{t}^{\left(m\right)}|{{\boldsymbol{e}}}{t:T}){||p}({{\boldsymbol{z}}}{t}^{\left(m\right)}|{{\boldsymbol{d}}}_{t-1}^{\left(m\right)})\right]$$
(30)
$$=-{\mathrm{ln}}\frac{1}{\sqrt{2\pi {{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}}\exp \left(-\frac{{\left({{\boldsymbol{x}}}{t}-{{\widehat{\boldsymbol{x}}}}{t}\right)}^{2}}{2{{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}\right)+\mathop{\sum}\limits {m}{D}{{KL}}\left[q({{\boldsymbol{z}}}{t}^{\left(m\right)}|{{\boldsymbol{e}}}{t:T}){||p}({{\boldsymbol{z}}}{t}^{\left(m\right)}|{{\boldsymbol{d}}}{t-1}^{\left(m\right)})\right]=\underbrace{\frac{{\left({{\boldsymbol{x}}}{t}-{\widehat{{\boldsymbol{x}}}}{t}\right)}^{2}}{2{{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}{\boldsymbol{+}}\frac{{\mathrm{ln}}\left(2\pi {{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}\right)}{2}}{{\rm{Negative}}, {\rm{accuracy}}}{\boldsymbol{+}}\underbrace{\mathop{\sum}\limits{m}{D}{{KL}}\left[q({{\boldsymbol{z}}}{t}^{\left(m\right)}|{{\boldsymbol{e}}}{t:T}){||p}({{\boldsymbol{z}}}{t}^{\left(m\right)}|{{\boldsymbol{d}}}{t-1}^{\left(m\right)})\right]}{\rm{Complexity}}$$
(31)
where (m) denotes a network module. The first term in expression Eq. (31) (negative accuracy term) is the NLL. For simplicity, we assume that each sensory state follows a Gaussian distribution. In the experiment, we divided the negative accuracy term by the dimensions of the proprioceptive, exteroceptive, and interoceptive sensations. However, the second term (complexity term) is the KLD between the posterior and prior distributions of the latent states. Assuming that the prior and posterior distributions follow a multivariate Gaussian distribution with a diagonal covariance matrix, as described above, the KLD is computed analytically as47,
$${\rm{Complexity}}=\mathop{\sum}\limits_{m}\left({\mathrm{ln}}\frac{{{\boldsymbol{\sigma }}}{p,t}^{(m)}}{{{\boldsymbol{\sigma }}}{q,t}^{(m)}}+\frac{{\left({{\boldsymbol{\mu }}}{p,t}^{(m)}-{{\boldsymbol{\mu }}}{q,t}^{(m)}\right)}^{2}+{\left({{\boldsymbol{\sigma }}}{q,t}^{(m)}\right)}^{2}}{2{\left({{\boldsymbol{\sigma }}}{p,t}^{(m)}\right)}^{2}}-\frac{1}{2}\right)$$
(32)
In the experiment, we divided the complexity term by the dimensions of the latent variables for each module.
In the learning phase, synaptic weights ({\boldsymbol{\omega }}) and adaptive variables ({{\boldsymbol{a}}}{1:T}) are updated to minimize the ({{\rm{VFE}}}{t}) over all time steps ((T=\mathrm{10,000})). We used gradient descent with a rectified Adam optimizer78 for parameter updates, where the partial derivative of the variational free energy with respect to each parameter was calculated using the BPTT algorithm.
By contrast, in the autonomous survival task, the MBH-RNN performed posterior updates within time steps from ({t}{c}-{{win}}{p}+1) to ({t}{c}+{{win}}{f}) at each sensorimotor time step ({t}{c}) (Fig. S2). The cost function in the past time window ({F}{{past}}) is the accumulated VFE of the past (VFEP) based on the observed sensations ({{\boldsymbol{x}}}{{t}{c}-{{win}}{p}+1:{t}{c}}).
$${F}{{past}}=\mathop{\sum }\limits{t={t}{c}-{{win}}{p}+1}^{{t}{c}}{{\rm{VFE}}}{t}$$
(33)
In contrast, VFEF is derived by considering a bound on the expected negative log model evidence (or the expected surprisal) under the distribution of predicted future sensory trajectories.
$$\begin{array}{ll}\displaystyle{{\rm{Expected}}; {\rm{suprisal}}=E}{p\left({{\boldsymbol{x}}}{{t}{c}+1:{t}{c}+{{win}}{f}}|{{\boldsymbol{z}}}{{t}{c}-{{win}}{p}+1:{t}{c}+{{win}}{f}},{{\boldsymbol{d}}}{{t}{c}-{{win}}{p}}\right)}\left[-{\mathrm{ln}}p\left({{\boldsymbol{x}}}{{t}{c}+1:{t}{c}+{{win}}{f}}\right)\right]\\qquad\qquad\qquad\qquad \displaystyle \le {E}{p\left({{\boldsymbol{x}}}{{t}{c}+1:{t}{c}+{{win}}{f}}|{{\boldsymbol{z}}}{{t}{c}-{{win}}{p}+1:{t}{c}+{{win}}{f}},{{\boldsymbol{d}}}{{t}{c}-{{win}}{p}}\right)}\left[\mathop{\sum }\limits_{t={t}{c}+1}^{{t}{c}+{{win}}{f}}\left[-{\mathrm{ln}}p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)+{D}{{KL}}\left[q\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{e}}}{t:{t}{c}+{{win}}{f}}\right){||p}\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{d}}}{t-1}\right)\right]\right]\right]\\qquad\qquad\qquad\qquad \displaystyle ={E}{p\left({{\boldsymbol{x}}}{{t}{c}+1}|{{\boldsymbol{d}}}{{t}{c}+1}\right)p\left({{\boldsymbol{x}}}{{t}{c}+2}|{{\boldsymbol{d}}}{{t}{c}+2}\right)\cdots p\left({{\boldsymbol{x}}}{{t}{c}+{{win}}{f}}|{{\boldsymbol{d}}}{{t}{c}+{{win}}{f}}\right)}\left[\mathop{\sum }\limits{t={t}{c}+1}^{{t}{c}+{{win}}{f}}\left[-{\mathrm{ln}}p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)+{D}{{KL}}\left[q\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{e}}}{t:{t}{c}+{{win}}{f}}\right){||p}\left({{\boldsymbol{z}}}{t},|,{{\boldsymbol{d}}}{t-1}\right)\right]\right]\right]\\qquad\qquad\qquad\qquad \displaystyle =\mathop{\sum }\limits_{t={t}{c}+1}^{{t}{c}+{{win}}{f}}\left[{E}{p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)}\left[-{\mathrm{ln}}p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)\right]+{D}{{KL}}\left[q\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{e}}}{t:{t}{c}+{{win}}{f}}\right){||p}\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{d}}}{t-1}\right)\right]\right]\\qquad\qquad\qquad\qquad \displaystyle =\mathop{\sum }\limits{t={t}{c}+1}^{{t}{c}+{{win}}{f}}\left[{\mathcal{H}}\left[p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)\right]+{D}{{KL}}\left[q\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{e}}}{t:{t}{c}+{{win}}{f}}\right){||p}\left({{\boldsymbol{z}}}{t}|{{\boldsymbol{d}}}{t-1}\right)\right]\right]\end{array}$$
(34)
Finally, the cost function in the future time window, ({F}_{{future}}) is derived by considering the KLD term in each module.
$${F}{{future}}=\mathop{\sum }\limits{t={t}{c}+1}^{{t}{c}+{{win}}{f}}{{\rm{VFEF}}}{t}$$
(35)
$${{\rm{VFEF}}}{t}={\mathcal{H}}\left[p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)\right]+\sum {m}{D}{{KL}}\left[q({{\boldsymbol{z}}}{t}^{\left(m\right)}|{{\boldsymbol{e}}}{t:{t}{c}+{{win}}{f}}){||p}({{\boldsymbol{z}}}{t}^{\left(m\right)}|{{\boldsymbol{d}}}_{t-1}^{\left(m\right)})\right]$$
(36)
Here, the first term is the conditional (Shannon) entropy of the sensations, and is calculated as follows by assuming a Gaussian distribution:
$$\begin{array}{l}{\mathcal{H}}\left[p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)\right]=-\int p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right){\mathrm{ln}}p\left({{\boldsymbol{x}}}{t}|{{\boldsymbol{d}}}{t}\right)d{{\boldsymbol{x}}}{t}\ =-\int \frac{1}{\sqrt{2\pi {{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}}\exp \left(-\frac{{\left({{\boldsymbol{x}}}{t}-{{\widehat{\boldsymbol{x}}}}{t}\right)}^{2}}{2{{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}\right){\mathrm{ln}}\frac{1}{\sqrt{2\pi {\hat{{\boldsymbol{\sigma }}}}{x,t}^{2}}}\exp \left(-\frac{{\left({{\boldsymbol{x}}}{t}-{{\widehat{\boldsymbol{x}}}}{t}\right)}^{2}}{2{{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}\right)d{{\boldsymbol{x}}}{t}\ =-\int \frac{1}{\sqrt{2\pi {{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}}\exp \left(-\frac{{\left({{\boldsymbol{x}}}{t}-{{\widehat{\boldsymbol{x}}}}{t}\right)}^{2}}{2{{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}\right)\left{-\frac{{\mathrm{ln}}\left(2\pi {{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}\right)}{2}-\frac{{\left({{\boldsymbol{x}}}{t}-{{\widehat{\boldsymbol{x}}}}{t}\right)}^{2}}{2{{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}\right}d{{\boldsymbol{x}}}{t}\ =\frac{{\mathrm{ln}}\left(2\pi {{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}\right)}{2}\int \frac{1}{\sqrt{2\pi {{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}}\exp \left(-\frac{{\left({{\boldsymbol{x}}}{t}-{{\widehat{\boldsymbol{x}}}}{t}\right)}^{2}}{2{{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}\right)d{{\boldsymbol{x}}}{t}\+,\int \frac{1}{\sqrt{2\pi {{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}}\exp \left(-\frac{{\left({{\boldsymbol{x}}}{t}-{{\widehat{\boldsymbol{x}}}}{t}\right)}^{2}}{2{{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}\right)\frac{{\left({{\boldsymbol{x}}}{t}-{{\widehat{\boldsymbol{x}}}}{t}\right)}^{2}}{2{{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}d{{\boldsymbol{x}}}{t}\ =\frac{{\mathrm{ln}}\left(2\pi {{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}\right)}{2}+\frac{1}{2{{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}\int {\left({{\boldsymbol{x}}}{t}-{{\widehat{\boldsymbol{x}}}}{t}\right)}^{2}\frac{1}{\sqrt{2\pi {{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}}\exp \left(-\frac{{\left({{\boldsymbol{x}}}{t}-{{\widehat{\boldsymbol{x}}}}{t}\right)}^{2}}{2{{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}\right)d{{\boldsymbol{x}}}{t}\ =\frac{{\mathrm{ln}}\left(2\pi {{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}\right)}{2}+\frac{1}{2{{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}{{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}\ =\frac{{\mathrm{ln}}\left(2\pi {{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}\right)}{2}+\frac{1}{2}\end{array}$$
(37)
Thus, the VFEF is described below.
$${{\mathrm{VFEF}}}{t}=\underbrace{\frac{1}{2}{\boldsymbol{+}}\frac{{\mathrm{ln}}\left(2{{\uppi }}{{\widehat{\boldsymbol{\sigma }}}}{{{x}},{{t}}}^{2}\right)}{2}}{{\mathrm{Predicted}},{\mathrm{sensory}},{\mathrm{entropy}}}{\boldsymbol{+}}\underbrace{\sum {m}{D}{{KL}}\left[q({{\boldsymbol{z}}}{t}^{\left(m\right)}|{{\boldsymbol{e}}}{t:{t}{c}+{{win}}{f}}){||p}({{\boldsymbol{z}}}{t}^{\left(m\right)}|{{\boldsymbol{d}}}{t-1}^{\left(m\right)})\right]}{{\mathrm{Predicted}},{\mathrm{complexity}}}$$
(38)
This derivation of VFEF naturally describes the minimization of the predicted future sensory entropy (uncertainty) or the meta-goal of Bayesian allostasis.
The MBH-RNN performs iterations of prediction generations and posterior updates within time steps from ({t}{c}-{{win}}{p}+1) to ({t}{c}+{{win}}{f}) to minimize the accumulated free-energy from the past to the future ({F}_{{allostasis}}) through BPTT.
$${F}{{allostasis}}={F}{{past}}+{F}_{{future}}$$
(39)
Furthermore, we prepared a setpoint model as the control condition in which the predicted sensory entropy term of the VFEF was replaced by a NLL (predicted negative accuracy term), or precision-weighted prediction error, given the explicit interoceptive target state (homeostatic setpoint) 0.0, which is equivalent to the prior preference for future outcomes within the concept of active inference applied to explain goal-directed interoceptive control52.
$$\underbrace{\frac{{\left({\boldsymbol{0}}-{{\widehat{\boldsymbol{x}}}}{t}\right)}^{2}}{2{{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}}{\boldsymbol{+}}\frac{{\mathrm{ln}}\left(2\pi {{\widehat{\boldsymbol{\sigma }}}}{x,t}^{2}\right)}{2}}{{\mathrm{Predicted}},{\mathrm{negative}},{\mathrm{accuracy}}}$$
(40)
The terms of VFE of the past and the predicted complexity term of VFEF are the same with the allostasis model.
The dimensions of the latent variables in the proprioceptive, exteroceptive, and interoceptive modules were ({N}{{\rm{Pz}}}=2,{N}{{\rm{Ez}}}=2,{N}{{\rm{Iz}}}=1), respectively, corresponding to the dimensions of the sensations in each modality. The dimension of latent variables in the unexpected-uncertainty-cause module was ({N}{{\rm{Uz}}}=1) as the minimal setting, while that in the higher-cognitive module was ({N}{{\rm{Cz}}}=3) as the dimension necessary for associating the information of each of the three modalities with sensory uncertainty. The dimension of the latent variables in the multimodal-associative module was ({N}{{\rm{Az}}}=8), which was determined by a preliminary experiment in which we increased the dimension of the latent variable in the multimodal-associative module individually and searched for the minimal setting for successful reconstruction of the learning data in all trained networks with different initial synaptic weights. For simplicity, the numbers of deterministic recurrent variables in the proprioceptive, interoceptive, unexpected-uncertainty-cause, and higher-cognitive modules were ({{N}{{\rm{Pd}}}={N}{{\rm{Id}}}={N}{{\rm{Ud}}}=N}{{\rm{Cd}}}=10) for simplicity, while those in the exteroceptive and multimodal-associative modules were ({N}{{\rm{Ed}}}={N}{{\rm{Ad}}}=20) required for reconstructing the learning data. We incorporated a temporal hierarchy (multiple timescale properties) into the MBH-RNN. Specifically, for each of the lower perceptual modules, we set the time constant as ({\tau }{{\rm{Pd}}}={\tau }{{\rm{Ed}}}={\tau }{{\rm{Id}}}=2) or (4) (for each of half the number of deterministic recurrent variables) as the simple multiple timescales setting. Then, the time constants in the multimodal-associative module and unexpected-uncertainty-cause module were set as ({\tau }{{\rm{Ad}}}={\tau }{{\rm{Ud}}}=4) or (8) (for each of the half-deterministic recurrent variables), while those in the higher-cognitive module were set as ({\tau }{{\rm{Cd}}}=8{\rm{or}}16) (for each of half deterministic recurrent variables). As such, the higher-level modules have a larger time constant (slower neural dynamics) than the lower-level modules have, thus implementing a temporal hierarchy. The synaptic weights were initialized with random values based on the Xavier method79. The biases of the deterministic recurrent variables were initialized with and fixed to random values following a Gaussian distribution ({\mathscr{N}}(\mathrm{0,10})), based on a previous study19. For quantitative evaluation, we trained 10 networks with different initial synaptic weights and biases for each hyper-parameter setting. In the learning phase, parameters including synaptic weights and adaptive (posterior) variables were updated ({{itr}}{{learning}}=\mathrm{100,000}) times with the rectified Adam optimizer (learning rate ({{lr}}{{learning}}=0.01), ({\beta }{1}=0.9), ({\beta }{2}=0.999)). In the autonomous survival task test, adaptive variables were updated ({{itr}}{{inference}}=200) times at each sensorimotor time step ({t}{c}) for a time window from the past time window of length ({{win}}{p}=10({{win}}{p}={t}{c}{if}{t}{c} < 10)) to the future time window of length ({{win}}{f}=200) with ({{lr}}{{inference}}=0.1). We chose the optimal parameter settings in the survival task from the combinations of ({{lr}}{{inference}}=\left{\mathrm{0.01,0.05,0.1,0.2,0.5,1.0}\right}\times {{itr}}{{inference}}=\left{\mathrm{100,200,500}\right}\times {{win}}{p}={\mathrm{10,20,50}}\times {{win}}{f}={\mathrm{50,100,200}}) by evaluating the autonomous survival time steps.
An analysis of the developed latent representation via the ablation of latent variables was conducted, as shown below (Fig. 2c). We regard the network predictions (({\hat{{\boldsymbol{x}}}}{1:T}) and ({\hat{{\boldsymbol{\sigma }}}}{x,1:T})) at the end of learning as the target of network predictions. First, for calculating the baseline variability of the network predictions, we made the trained normal network generating 10 different timeseries (({\hat{{\boldsymbol{x}}}}{1:T}) and ({\hat{{\boldsymbol{\sigma }}}}{x,1:T})) with different sets of random sampling of latent states. The differences between the target timeseries and the 10 re-generated timeseries were calculated, and the maximum value of the differences for each of the proprioceptive, exteroceptive, interoceptive, and uncertainty predictions was named the baseline difference (BD). Next, we removed one latent variable and made the lesioned network generating 10 different timeseries (({\hat{{\boldsymbol{x}}}}{1:T}) and ({\hat{{\boldsymbol{\sigma }}}}{x,1:T})) with different random sampling. The differences between the target timeseries and the 10 timeseries generated by the lesioned network were calculated, and the average of the differences for each of the proprioceptive, exteroceptive, interoceptive, and uncertainty predictions was named the ablation effect (AE). If the AE for a certain modality (proprioceptive, exteroceptive, interoceptive, or uncertainty predictions) was larger than the corresponding BD was, we considered that the ablated latent variable represented the affected modality information. For example, if the ablation of a latent variable of the multimodal-associative module affects proprioceptive and interoceptive sensory predictions, the ablated latent variable is considered to represent both proprioceptive and interoceptive information and is a bimodal latent variable. We repeated the evaluation for all latent variables in all network modules for all 10 trained networks. Finally, we categorized the types of latent variables based on the information represented in each network module and calcula
[Content truncated]











