From Timescales to Information Processing
in Stochastic Biological Systems

Max Planck Gesellschaft

UNESCO

EPFL

SIFS SUMMER SCHOOL 2026
Giorgio Nicoletti, International Center for Theoretical Physics
18th September 2026

Università degli Studi di Padova

Laboratory of Interdisciplinary Physics

Living systems must sense their surroundings and use that information to respond

Across biological scales, living systems must sense their surroundings and use that information to respond
Schematic bacterium with surface receptors and a flagellum
Bacteria
Chemical conditions influence movement
Molecular signaling connects sensing to changes in swimming
Schematic neighboring cells communicating through extracellular signals
Cells in an organism
Signals from other cells guide cellular responses
Regulatory networks connect extracellular cues to gene activity
Schematic connected neurons representing a sensory neural circuit
Organisms with brains
Sensory signals inform behavior
Neural circuits combine incoming evidence with internal state
Different biological systems face the same physical question:
how can a noisy internal response tell the system about its surroundings?
Tkačik & Bialek, ARCMP 7, 89 (2016); Tkačik & ten Wolde, Annu. Rev. Biophys. 54, 249 (2025)

What does it mean to process information?

An environmental change becomes detectable when it leaves a distinguishable trace inside the system
ENVIRONMENT
A condition
to sense
Chemical concentration,
light, or other cues
MEASUREMENT
Physical
interactions
Binding events,
sensory activity
PROCESSING
Internal
dynamics
Integration, memory,
adaptation
RESPONSE
Change
in activity
Movement, regulation,
neural responses

Environmental fluctuations Complex environments are both intrinsically noisy and change over time, and are often characterized by an different conditions

Information processing A system must be able to estimate environmental conditions quickly and reliably, while being able to distinguish an ensemble of inputs

How can a system respond reliably to one or more inputs,
and what do timescales have to do with it?
Tkačik & Bialek, ARCMP 7, 89 (2016); Tkačik & ten Wolde, Annu. Rev. Biophys. 54, 249 (2025)

Sensing and responding takes time, but the environment evolves in the meantime

Collecting evidence takes time, while the environment and the system keep changing
An illustrative changing signal, a short-memory response that follows changes but fluctuates, and a long-memory response that smooths noise but delays and attenuates changes
Same signal and noisy measurements; two response memories
Illustrative filtered records, not experimental data
Short memory
> Follows environmental changes quickly
> Not robust to estimation noise
Long memory
> Robust estimation via noise integration
> May not resolve environmental changes
Relative timescales matter
How long does the signal persist?
How quickly does the system respond?
Relative timescales matter: how long does the signal persist, and how quickly does the system respond
ten Wolde et al., J. Stat. Phys. 162, 1395 (2016); Mora & Nemenman, PRL 123, 198101 (2019)

How precisely can one receptor measure concentration?

ENVIRONMENT (\(\tau_H\))
\(H(t)\)
External signal(s)

\(\longrightarrow\)

RECEPTOR (\(\tau_R\))
\(R(t)\in\{0,1\}\)
Bound or unbound

\(\longrightarrow\)

READOUT (\(\tau_U\))
\(U(t)\)
Readout or response

Illustrative signal, receptor and readout trajectories

Observe for a time \(T\) to estimate the environment. We want: - \(\tau_R\ll T\): average effectively over receptor memory
- \(T\ll\tau_H\): the environment changes little during the estimate

ten Wolde et al., J. Stat. Phys. 162, 1395 (2016); Tkačik & Bialek, Annu. Rev. Condens. Matter Phys. 7, 89 (2016)

Diffusion sets the rate of receptor encounters

A RECEPTOR IN A REFLECTING MEMBRANE
Circular receptor disk of radius s in a reflecting membrane

\(s\): disk radius
\(D\): ligand diffusivity
\(c\): faraway concentration

\(H(t)=\{c = \mathrm{const}\}\) \(\tau_H=\tau_c\)

FROM DIFFUSION TO BINDING
\[J_{\rm empty}=4Dsc,\qquad k_+=4Ds\]

\(J_{\rm empty}\) counts captures per unit time while the receptor is not occupied

\[\overline N_{\rm bind}=T(1-p)\,4Dsc\]

\(p\) is the bound fraction, so \(1-p\) is the fraction of time available for capture

\(H(t)=\{c(t)\}\); here \(c(t)=c\) for \(T\ll\tau_c=\tau_H\)
\(c\): bulk concentration far from the receptor
Derivation: the diffusion flux into a receptor disk
Fixed \(H(t)=c(t)=c\) during the record (\(T\ll\tau_c\)); ideal diffusion-controlled disk; perfectly absorbing empty patch and reflecting surrounding plane
Berg & Purcell, Biophys. J. 20, 193 (1977)

The ligand diffuses and can get randomly captured by the receptor, which then switches from being inactive to active

\[ 0 \xrightleftharpoons[k_-]{k_+c} 1 \]

Observable
\(R(t)\in\{0,1\}\) over a window of duration \(T\)

Parameter
an unknown concentration \(c\), fixed within the window

Strategy
measure the fraction of time bound, then invert the mean response \(p(c)\)

\[ 0 \xrightleftharpoons[k_-]{k_+c} 1 \]

A RECEPTOR IN A REFLECTING MEMBRANE

Simulated binary receptor occupancy with binding and unbinding events

Integrating the disk flux to get the binding rate

\(c_{\rm loc}(\rho,z)\): stationary concentration; \(\rho\): radial distance in the membrane; \(z\): height above it
Stationary diffusion and boundary conditions

\[\nabla^2c_{\rm loc}=0,\qquad c_{\rm loc}\to c\quad\text{far from the patch}\]

\[c_{\rm loc}(\rho,0)=0\ (\rho<s),\qquad\partial_zc_{\rm loc}(\rho,0)=0\ (\rho>s)\]

\(z=0\): membrane surface. The disk of the receptor is an absorbing boundary, and the surrounding membrane a reflecting one
Solve in coordinates adapted to the disk

\[\rho=s\sqrt{(1+\xi^2)(1-\eta^2)},\qquad z=s\xi\eta\]

\[\frac{d}{d\xi}\!\left[(1+\xi^2)\frac{dc_{\rm loc}}{d\xi}\right]=0\quad\Longrightarrow\quad c_{\rm loc}=\frac{2c}{\pi}\arctan\xi\]

Disk: \(\xi=0\). Far field: \(\xi\to\infty\). Reflecting plane: \(\eta=0\), with zero normal derivative
Surface gradient → inward flux density

\[\left.\frac{\partial z}{\partial\xi}\right|_{\rho,\xi=0}=s\eta=\sqrt{s^2-\rho^2}\]

\[j(\rho)=D\,\partial_zc_{\rm loc}(\rho,0)=\frac{2Dc}{\pi\sqrt{s^2-\rho^2}}\]

\(j\): captured molecules per area and time
Integrate over rings of area \(2\pi\rho\,d\rho\)

\[J_{\rm empty}=\int_0^s j(\rho)\,2\pi\rho\,d\rho=4Dc\int_0^s\frac{\rho\,d\rho}{\sqrt{s^2-\rho^2}}\]

\[=4Dc\left[-\sqrt{s^2-\rho^2}\right]_0^s=4Dsc\]

\(k_+=J_{\rm empty}/c=4Ds\) has units volume/time
Stationary diffusion above a reflecting plane; absorbing disk of radius \(s\); fixed bulk concentration
Berg & Purcell, Biophys. J. 20, 193 (1977)

Estimate concentration from the fraction of time bound

Derivation: receptor memory

Observe the binary trajectory \(R(t)\) for a duration \(T\)

Berg & Purcell, Biophys. J. 20, 193 (1977); ten Wolde et al., J. Stat. Phys. 162, 1395 (2016)

A RECEPTOR IN A REFLECTING MEMBRANE

\(s\): disk radius
\(D\): ligand diffusivity
\(c\): faraway concentration
\(H=c\), \(\tau_H=\tau_c\)

Observable
\(R(t)\in\{0,1\}\) over a window of duration \(T\)

Parameter
an unknown concentration \(c\), fixed within the window

Strategy
measure the fraction of time bound, then invert the mean response \(p(c)\)

\[ 0 \xrightleftharpoons[k_-]{k_+c} 1 \]

CALIBRATION

\[ p=\langle R\rangle=\frac{k_+c}{k_+c+k_-} \]

so larger \(c\) raises the fraction of time bound

MEMORY

\[ C_R(u)=p(1-p)e^{-|u|/\tau_{\mathrm{corr}}} \]

\[ \tau_R=\tau_{\mathrm{corr}}=\frac{1}{k_+c+k_-} \]

\(C_R(u)\) measures how correlations at lag \(u\) decay

\(s\): disk radius
\(D\): ligand diffusivity
\(c\): faraway concentration

\(H(t)=\{c = \mathrm{const}\}\) \(\tau_H=\tau_c\)

Circular receptor disk of radius s in a reflecting membrane

Simulated binary receptor occupancy with binding and unbinding events

Flux balance gives the concentration–occupancy relation

Binding rate \(\alpha=k_+c\); unbinding rate \(\beta=k_-\); idea: estimate \(c\) from the measured bound fraction \(p\) (the stationary probability)

\[ \frac{dP_1}{dt}=\alpha(1-P_1)-\beta P_1 \]

At stationarity, probability flux into and out of the bound state balance:

\[ 0=\alpha(1-p)-\beta p \]

\[ p=\frac{\alpha}{\alpha+\beta} =\frac{k_+c}{k_+c+k_-} \implies c = \frac{p}{1-p}\frac{k_-}{k_+} \]

Idea: estimate \(c\) from the measured bound fraction
LIMITING CHECKS

\[ c\to0:\quad p\to0 \]

\[ c\to\infty:\quad p\to1 \]

Berg & Purcell, Biophys. J. 20, 193 (1977)

Conditional relaxation determines receptor memory

RELAXATION AFTER A BOUND STATE

\[ q(u)=P[R(u)=1\mid R(0)=1],\qquad q(0)=1 \]

The same probability balance holds after conditioning:

\[\dot q=\alpha(1-q)-\beta q\]

Subtract \(0=\alpha(1-p)-\beta p\):

\[ \frac{d}{du}(q-p)=-(\alpha+\beta)(q-p) \]

\[ q(u)=p+(1-p)e^{-u/\tau_{\mathrm{corr}}}, \qquad \tau_{\mathrm{corr}}=(\alpha+\beta)^{-1} \]

TWO-TIME COVARIANCE

\[ \begin{aligned} \langle R(0)R(u)\rangle &=P[R(0)=1]P[R(u)=1\mid R(0)=1]\\ &=p\,q(u) \end{aligned} \]

\[ C_R(u)=p(1-p)e^{-|u|/\tau_{\mathrm{corr}}} \]

\(C_R(0)=p(1-p)\) and \(C_R(u)\to0\) as \(u\to\infty\)
Stationarity gives \(C_R(-u)=C_R(u)\)
Berg & Purcell, Biophys. J. 20, 193 (1977)

Both transition rates determine the correlation time

ONE BOUND INTERVAL
\[P(T_b>u)=e^{-k_-u}\] \[\tau_{\mathrm{bound}}=\frac1{k_-}\]

which is the mean duration of a bound episode

Correlation decays through both transition rates

\[\tau_{\mathrm{corr}}=\frac1{k_+c+k_-}\] \[1-p=\frac{k_-}{k_+c+k_-}\] \[\tau_{\mathrm{corr}}=(1-p)\tau_{\mathrm{bound}}\]

Receptor correlation can be expressed in terms of both
\[ \begin{aligned} \langle R(0)R(u)\rangle &=p^2+p(1-p)e^{-|u|/\tau_{\mathrm{corr}}}\\[0.5em] &=p^2+p(1-p)e^{-|u|/[(1-p)\tau_{\mathrm{bound}}]} \end{aligned} \]

The second line is the paper’s Eq.49. The centered covariance subtracts \(p^2\)

Constant transition rates; stationary two-state receptor. \(t_b\): bound dwell time; \(\tau_{\rm corr}\): correlation time
Berg & Purcell, Biophys. J. 20, 193 (1977)

Long records average over receptor memory

Derivation: occupancy variance
The cell can only measure the fraction of time bound (measured in \(x=T/\tau_{\mathrm{corr}}\)) so its estimate is uncertain\[\bar R_T=T^{-1}\int_0^T R(t)\,dt \implies \operatorname{Var}(\bar R_T)=2p(1-p)\frac{x-1+e^{-x}}{x^2}\]
Exact variance at fixed concentration; \(x=T/\tau_{\mathrm{corr}}\)

\[\operatorname{Var}(\bar R_T)=2p(1-p)\frac{x-1+e^{-x}}{x^2}\]

SHORT RECORD
\[T\ll\tau_{\mathrm{corr}}:\quad \operatorname{Var}(\bar R_T)\to p(1-p)\]

Little changes during the record: effectively one observation

LONG RECORD
\[T\gg\tau_{\mathrm{corr}}:\quad \operatorname{Var}(\bar R_T)\simeq\frac{2p(1-p)\tau_{\mathrm{corr}}}{T}\]

\(N_{\mathrm{eff}}\simeq T/(2\tau_{\mathrm{corr}})\) independent samples give the same variance

\(x=T/\tau_{\mathrm{corr}},\quad \mathrm{Var}(\bar R_T)=2p(1-p)(x-1+e^{-x})/x^2\)
Fixed concentration; stationary receptor. Histogram: 1,800 records. Effective samples match the exact variance

Berg & Purcell, Biophys. J. 20, 193 (1977)

Occupancy variance measures finite-record error

\(E[\cdot]\) is the average with respect to repeated records at the same concentration, while \(\bar R_T\) is the time-average from one trajectory
\[\bar R_T=\frac1T\int_0^T R(t)\,dt\]
The fraction of the observation window spent bound
Stationarity gives an unbiased occupancy estimate
\[E[\bar R_T]=\frac1T\int_0^T E[R(t)]\,dt=p\]
\(E\) averages repeated records at fixed \(c\); \(E[R(t)]=p\)
Mean-squared error equals variance
\[E[(\bar R_T-p)^2]=\operatorname{Var}(\bar R_T) =\frac1{T^2}\int_0^T\!dt\int_0^T\!dt'\,\langle\delta R(t)\delta R(t')\rangle = \frac1{T^2}\int_0^T\!dt\int_0^T\!dt'\,C_R(t-t')\]
The bias term vanishes because \(E[\bar R_T]=p\)
Expand the squared time average, with \(\delta R=R-p\)
\[\operatorname{Var}(\bar R_T)=\frac1{T^2}\int_0^T\!dt\int_0^T\!dt'\,\langle\delta R(t)\delta R(t')\rangle\]
Every pair of observation times contributes its covariance
Use stationarity to express covariance by time lag

\[\operatorname{Var}(\bar R_T)=\frac1{T^2}\int_0^T\!dt\int_0^T\!dt'\,C_R(t-t')\]

Berg & Purcell, Biophys. J. 20, 193 (1977)

\[ \bar R_T=\frac1T\int_0^T R(t)\,dt, \qquad E[\bar R_T]=\frac1T\int_0^T E[R(t)]\,dt=p \]

\(T-u\): pair weight. Factor 2: the two time orderings. Stationarity makes \(C_R\) even

\[ \operatorname{Var}(\bar R_T) =\frac{2}{T^2}\int_0^T(T-u)C_R(u)\,du \]

Receptor memory sets the effective sample count

Substitute the receptor covariance
\[ \operatorname{Var}(\bar R_T) =\frac{2p(1-p)}{T^2}\int_0^T(T-u)e^{-u/\tau_{\mathrm{corr}}}\,du \]
Integration by parts
\[ \begin{aligned} \int_0^T(T-u)e^{-u/\tau_{\mathrm{corr}}}\,du &=\left[-\tau_{\mathrm{corr}}(T-u)e^{-u/\tau_{\mathrm{corr}}}\right]_0^T-\tau_{\mathrm{corr}}\int_0^T e^{-u/\tau_{\mathrm{corr}}}\,du\\ &=T\tau_{\mathrm{corr}}-\tau_{\mathrm{corr}}^2(1-e^{-T/\tau_{\mathrm{corr}}}) \end{aligned} \]
Exact finite-window variance
\[ \operatorname{Var}(\bar R_T) =\frac{2p(1-p)}{T^2} \left[T\tau_{\mathrm{corr}}-\tau_{\mathrm{corr}}^2(1-e^{-T/\tau_{\mathrm{corr}}})\right] \]
SHORT WINDOW
\[ T\ll\tau_{\mathrm{corr}}: \quad \operatorname{Var}(\bar R_T)\to p(1-p) \]
The receptor usually stays in its initial state
LONG WINDOW
\[ T\gg\tau_{\mathrm{corr}}: \quad \operatorname{Var}(\bar R_T)\simeq \frac{2p(1-p)\tau_{\mathrm{corr}}}{T} \]
Match \(p(1-p)/N\): \(N_{\rm eff}=T/(2\tau_{\rm corr})\)
Berg & Purcell, Biophys. J. 20, 193 (1977)

Integration by parts

\[ \begin{aligned} \int_0^T(T-u)e^{-u/\tau_{\mathrm{corr}}}\,du &=\left[-\tau_{\mathrm{corr}}(T-u)e^{-u/\tau_{\mathrm{corr}}}\right]_0^T-\tau_{\mathrm{corr}}\int_0^T e^{-u/\tau_{\mathrm{corr}}}\,du\\ &=T\tau_{\mathrm{corr}}-\tau_{\mathrm{corr}}^2(1-e^{-T/\tau_{\mathrm{corr}}}) \end{aligned} \]

Exact finite-window variance

\[ \operatorname{Var}(\bar R_T) =\frac{2p(1-p)}{T^2} \left[T\tau_{\mathrm{corr}}-\tau_{\mathrm{corr}}^2(1-e^{-T/\tau_{\mathrm{corr}}})\right] \]

The Berg–Purcell limit: diffusion and receptor size set sensing precision

Infer \(c\) by inverting the occupancy calibration: \(\hat c=(k_-/k_+)\bar R_T/(1-\bar R_T)\)
PRECISION OF OCCUPANCY AVERAGING
\[\frac{\operatorname{Var}(\hat c)}{c^2}\simeq\frac{2\tau_{\rm corr}}{Tp(1-p)}=\frac2{\overline N_{\rm bind}}\]

this is the long-\(T\) concentration error, which bounds the single-receptor precision

COUNT THE BINDING EVENTS
\[\overline N_{\rm bind}=Tk_+c(1-p)\]

the number of binding events scales linearly with the observation time, so the relative RMS error falls as \(T^{-1/2}\)

The estimator matters: full-trajectory likelihood reaches \(1/\overline N_{\rm bind}\) asymptotically

HOW QUICKLY EVENTS ARRIVE

\[k_+=4Ds\]

\[\overline N_{\rm bind}=4Dsc(1-p)T\]

\(D\): diffusivity; \(s\): the same disk radius
\(c\): faraway concentration

HOW PRECISELY OCCUPANCY SENSES

\[\frac{\operatorname{Var}(\hat c)}{c^2}\simeq\frac{2}{\overline N_{\rm bind}} = \frac1{2Dsc(1-p)T}\]

\[=\frac1{2Dsc(1-p)T}\]

diffusion and receptor geometry determine sensing of the fixed bulk concentration

What changes when concentration varies during the observation?
Derivations: concentration uncertainty and diffusion-limited precision
Occupancy-averaging estimator; many events; local error propagation; \(p\) away from 0 and 1
Ideal disk receptor; long-record occupancy averaging and local error propagation. Geometry and estimator determine the prefactor
Berg & Purcell, Biophys. J. 20, 193 (1977); Endres & Wingreen, PRL 103, 158101 (2009)

The calibration slope converts occupancy error to concentration error

Differentiate the calibration \(p(c)\)

\[p(c)=\frac{k_+c}{k_+c+k_-},\qquad c\frac{dp}{dc}=p(1-p) \quad \implies \quad \hat c=\frac{k_-}{k_+}\frac{\bar R_T}{1-\bar R_T}\]

\[\hat c=\frac{k_-}{k_+}\frac{\bar R_T}{1-\bar R_T}\]

For small errors, \(\delta p\simeq (dp/dc)\,\delta c\)
Linearize the inverse calibration

\[\frac{\hat c-c}{c}\simeq\frac{\bar R_T-p}{p(1-p)} \quad \implies \quad \frac{\operatorname{Var}(\hat c)}{c^2}\simeq\frac{\operatorname{Var}(\bar R_T)}{p^2(1-p)^2}\]

\[\frac{\operatorname{Var}(\hat c)}{c^2}\simeq\frac{\operatorname{Var}(\bar R_T)}{p^2(1-p)^2}\]

Divide occupancy variance by the squared calibration slope
Insert the long-window occupancy variance

\[\operatorname{Var}(\bar R_T)\simeq\frac{2p(1-p)\tau_{\rm corr}}{T} \quad \implies \quad \frac{\operatorname{Var}(\hat c)}{c^2}\simeq\frac{2\tau_{\rm corr}}{Tp(1-p)}\]

\[\frac{\operatorname{Var}(\hat c)}{c^2}\simeq\frac{2\tau_{\rm corr}}{Tp(1-p)}\]

Express that precision in observable event counts

\[\overline N_{\rm bind}=Tk_+c(1-p)=Tk_-p=\frac{T}{\tau_{\rm corr}}p(1-p) \quad \implies \quad \frac{\operatorname{Var}(\hat c)}{c^2}\simeq\frac2{\overline N_{\rm bind}}\]

\[\frac{\operatorname{Var}(\hat c)}{c^2}\simeq\frac2{\overline N_{\rm bind}}\]

Stationary binding and unbinding fluxes are equal. Full-trajectory likelihood gives half this asymptotic variance
Long-record asymptotic precision; many transitions; small errors under local inversion
Berg & Purcell, Biophys. J. 20, 193 (1977); Endres & Wingreen (2009)

Insert the capture rate to obtain the sensing limit

FROM THE RECEPTOR RECORD
\[\frac{\operatorname{Var}(\hat c)}{c^2}\simeq\frac{2\tau_{\mathrm{corr}}}{T p(1-p)}\]

\(p\) is the mean occupancy, \(T\) is the observation time, and we use the fact that \(\tau_{\mathrm{corr}}=1/(k_+c+k_-)\)

Eliminate the correlation time to explicit encounters
\[\frac{p}{\tau_{\mathrm{corr}}}=k_+c\] \[\frac{\operatorname{Var}(\hat c)}{c^2}\simeq\frac{2}{k_+c(1-p)T}\]

so the error falls as the number of binding events grows

Insert the Berg–Purcell patch arrival rate
\[k_+=4Ds\] \[\frac{\operatorname{Var}(\hat c)}{c^2}\simeq\frac{1}{2Dsc(1-p)T}\]

so faster diffusion, a larger receptor and longer observation supply more encounters, while occupancy reduces available capture time

The estimator in terms of physical parameters
\[\frac{\delta c}{c}\sim(DscT)^{-1/2}\]

The geometry and sensing strategy fix the prefactor. Occupancy averaging gives \(2/\overline N_{\mathrm{bind}}\); a resolved-path likelihood reaches \(1/\overline N_{\mathrm{bind}}\) asymptotically

Occupancy averaging; long record; many events; local error propagation; ideal disk receptor
Berg & Purcell, Biophys. J. 20, 193 (1977); Endres & Wingreen, PRL 103, 158101 (2009)

From fixed concentration to changing signals

From fixed concentration \(c\) to a changing signal \(H(t)\)

ENVIRONMENT

\[ H(t) \]

RECEPTOR

\[ R(t) \]

RESPONSE

\[ U(t) \]

Binding and
unbinding

\[ \tau_H \]

Signal persistence
\(\tau_H\)

\[ \tau_U \]

Response memory
\(\tau_U\)

Let’s consider a perfect receptor and focus on how well a noisy response can track a noisy environment

Longer \(\tau_U\) averages noise, but retains earlier values of \(H(t)\)

Mora & Nemenman, PRL 123, 198101 (2019); Malaguti & ten Wolde, eLife 10, e62574 (2021)

Illustrative H trajectory

Illustrative R trajectory

Illustrative U trajectory

How accurately does an exponential readout track its input?

Signal persistence
\(\tau_H\)

Response memory
\(\tau_U\)

\[ \frac{dH}{dt}=-\frac{H}{\tau_H}+\sqrt{\frac{2\sigma_H^2}{\tau_H}}\,\xi(t) \]

\[ \frac{dU}{dt}=\frac{H-U}{\tau_U}+\frac{\sqrt{2D_\eta}}{\tau_U}\,\eta(t) \]

Stationary variance \(\sigma_H^2\), \(\langle \xi(t)\xi(t')\rangle = \delta(t-t')\)

Noisy exponential averaging, \(\langle\eta(t)\eta(t')\rangle = \delta(t-t')\)

Unit white noise: \(\langle\eta(t)\eta(t^\prime)\rangle=\delta(t-t^\prime)\)

Choose \(\tau_U\) to minimize the tracking error \(\mathbb{E}[(U-H)^2]\)

Independent unit white noises \(\xi(t)\) and \(\eta(t)\); fixed \(\tau_H\), \(\sigma_H^2\), \(D_\eta\)

Mora & Nemenman, PRL 123, 198101 (2019)

\[ E[(U-H)^2]=\underbrace{\sigma_H^2\frac{\tau_U}{\tau_H+\tau_U}}_{\text{tracking error}}+\underbrace{\frac{D_\eta}{\tau_U}}_{\text{response noise}} \]

Illustrative H trajectory

Illustrative U trajectory

Averaging suppresses noise but increases tracking error

\[ E[(U-H)^2]=\underbrace{\sigma_H^2\frac{\tau_U}{\tau_H+\tau_U}}_{\text{tracking error}}+\underbrace{\frac{D_\eta}{\tau_U}}_{\text{response noise}} \]

Exact for the stationary linear model

When does their sum have a finite minimum?

\[ \tau_U\ll\tau_H:\quad E[(U-H)^2]\simeq\sigma_H^2\frac{\tau_U}{\tau_H}+\frac{D_\eta}{\tau_U} \]

\(\tau_U\to0\): white response noise diverges

\[ \begin{gathered}U\longrightarrow\langle H\rangle=0\\[3pt]E[(U-H)^2]\longrightarrow\sigma_H^2\end{gathered} \]

\(\tau_U\gg\tau_H\): average over many signal fluctuations

Vary \(\tau_U\) at fixed \(\sigma_H^2\), \(\tau_H\), \(D_\eta\) and unit gain

\[ \implies \qquad \epsilon=\sqrt{\frac{D_\eta}{\sigma_H^2\tau_H}} \]

\[ \tau_U^*=\frac{\epsilon}{1-\epsilon}\tau_H\qquad(0<\epsilon<1) \]

The optimum balances sampling and tracking

\[ \epsilon\ge1:\qquad\tau_U^*\to\infty \]

Long memory reduces the noise-induced error

Unit gain; white response noise
\(D_\eta=0\) gives the boundary optimum \(\tau_U\to0\)

\(x=\tau_U/\tau_H,\quad x^*=\epsilon/(1-\epsilon)\ (0<\epsilon<1);\qquad \epsilon\ge1:\ x^*\to\infty\)
Vary memory along a curve; changing noise defines another optimization. ε² = Dη/(σ_H²τ_H)
Derivation: tracking error and optimal memory

Signal persistence
\(\tau_H\)

Response memory
\(\tau_U\)

\[ \frac{dH}{dt}=-\frac{H}{\tau_H}+\sqrt{\frac{2\sigma_H^2}{\tau_H}}\,\xi(t) \]

\[ \frac{dU}{dt}=\frac{H-U}{\tau_U}+\frac{\sqrt{2D_\eta}}{\tau_U}\,\eta(t) \]

\(\langle \xi(t)\xi(t')\rangle = \delta(t-t')\)

Noisy exponential averaging, \(\langle\eta(t)\eta(t')\rangle = \delta(t-t')\)

\[ \frac{dH}{dt}=-\frac{H}{\tau_H}+\sqrt{\frac{2\sigma_H^2}{\tau_H}}\,\xi(t) \]

\[ \frac{dU}{dt}=\frac{H-U}{\tau_U}+\frac{\sqrt{2D_\eta}}{\tau_U}\,\eta(t) \]

Illustrative H trajectory

Illustrative U trajectory

Three stationary moments determine tracking error

\[ E[(U-H)^2]=V_H+V_U-2C \]

\(V_H=E[H^2]\): signal variance
\(V_U=E[U^2]\): readout variance
\(C=E[HU]\): present-state covariance

At fixed variances, a higher covariance \(C\) reduces the error because the environment and response are correlated

Signal variance: relaxation removes variance; noise adds it

\[ 0=-\frac{2V_H}{\tau_H}+\frac{2\sigma_H^2}{\tau_H} \]

\[ V_H=\sigma_H^2 \]

At stationarity, \(dE[H^2]/dt=0\)

Covariance: the response follows the signal while both relax

\[ 0=-\left(\frac1{\tau_H}+\frac1{\tau_U}\right)C+\frac{\sigma_H^2}{\tau_U} \]

\[ C=\sigma_H^2\frac{\tau_H}{\tau_H+\tau_U} \]

Independent noises: \(\langle\xi(t)\eta(t^\prime)\rangle=0\)
Longer memory reduces covariance with the present

Response variance: relaxation, signal and noise contributions

\[ 0=-\frac{2V_U}{\tau_U}+\frac{2C}{\tau_U}+\frac{2D_\eta}{\tau_U^2} \]

\[ V_U=C+\frac{D_\eta}{\tau_U} \]

\[ C=\sigma_H^2\frac{\tau_H}{\tau_H+\tau_U},\qquad V_H=\sigma_H^2 \]

Response noise adds variance at rate \(2D_\eta/\tau_U^2\)

Stationary zero means; independent noises; positive timescales; exact linear moment equations

Separate tracking error from observation noise

Insert \(V_H=\sigma_H^2\) and \(V_U=C+D_\eta/\tau_U\)

\[ E[(U-H)^2]=\sigma_H^2-C+\frac{D_\eta}{\tau_U} \]

Then use \(C=\sigma_H^2\tau_H/(\tau_H+\tau_U)\)

\[ E[(U-H)^2]=\underbrace{\sigma_H^2\frac{\tau_U}{\tau_H+\tau_U}}_{\text{tracking error}}+\underbrace{\frac{D_\eta}{\tau_U}}_{\text{response noise}} \]

Long memory: more error from signal changes. Short memory: more response noise

Minimize current-state error over the readout memory

Minimize \(E[(U-H)^2]\) over \(\tau_U\) at fixed \(\sigma_H^2\), \(\tau_H\), and \(D_\eta\)

\[ \frac{d}{d\tau_U}E[(U-H)^2]=\frac{\sigma_H^2\tau_H}{(\tau_H+\tau_U)^2}-\frac{D_\eta}{\tau_U^2}=0 \]

take the positive root:

\[ \frac{\tau_U}{\tau_H+\tau_U}=\epsilon,\qquad\epsilon=\sqrt{\frac{D_\eta}{\sigma_H^2\tau_H}} \]

\[ \tau_U^*=\frac{\epsilon}{1-\epsilon}\tau_H,\qquad0<\epsilon<1 \]

\[ \epsilon\ll1:\qquad\tau_U^*\simeq\sqrt{\frac{D_\eta\tau_H}{\sigma_H^2}} \]

For \(\epsilon\ge1\), increasing \(\tau_U\) keeps reducing the error
Long memory averages the signal to its mean; the error tends to \(\sigma_H^2\)

\(D_\eta=0\) gives the limiting case \(\tau_U\to0\)

Vary the response memory while keeping the signal and noise fixed

How much does the response tell us about the input?

Environmental signal \(H\)

\[ {\color{#0366c8}p_H(h)} \]

\[ p_{HU}(h,u) \]

\[ {\color{#800000}p_U(u)} \]

Signal \(h\)

\(u\)

Stochastic response \(U\)

\[ {\color{#0366c8}p_H(h)}=\int p_{HU}(h,u)\,du \]

\[ {\color{#800000}p_U(u)}=\int p_{HU}(h,u)\,dh \]

The statistical dependence in the joint distribution can be measured by the mutual information

\[ \begin{gathered}I(H;U)=\int p_{HU}(h,u)\log\frac{p_{HU}(h,u)}{{\color{#0366c8}p_H(h)}{\color{#800000}p_U(u)}}\,dh\,du\end{gathered} \]

Measures statistical dependence

\[ p_{HU}=p_Hp_U\quad\Longrightarrow\quad I(H;U)=0 \]

All responses

\[ {\color{#800000}p_U(u)} \]

Response \(u\)

The spread contains signal variation and response noise

At a known signal

\[ p_{U\mid H}(u\mid h) \]

Response \(u\)

Jointly Gaussian signals and responses

\[ I(H;U)={\color{#800000}\mathcal H(U)}-\mathcal H(U\mid H) \]

Mutual information measures the average reduction in uncertainty when knowing \(H\)

Two equally likely signals

Responses to each signal

\[ p(u\mid h_1) \]

\[ p(u\mid h_2) \]

Response \(u\)

Overlap makes the two signals harder to distinguish

All responses together

\[ {\color{#800000}p_U(u)}=\tfrac12p(u\mid h_1)+\tfrac12p(u\mid h_2) \]

Response \(u\)

\[ I(H;U)\le\mathcal H(H)=\log2\ \mathrm{nats}=1\ \mathrm{bit} \]

Two equally likely signals · Gaussian responses
Natural logarithms give nats · divide by \(\log 2\) for bits
Derivation: mutual information and uncertainty
Shannon, Bell Syst. Tech. J. 27, 379 (1948)

No one-to-one correspondence in general:

the same environmental condition can lead to different responses and the system is described by a
joint distribution \(p_{HU}(h,u)\)

\[ {\color{#800000}p_U(u)}=\int p_{HU}(h,u)\,dh \]

\[ {\color{#0366c8}p_H(h)}=\int p_{HU}(h,u)\,du \qquad {\color{#800000}p_U(u)}=\int p_{HU}(h,u)\,dh \]

\[ I(H;U)\le\mathcal H(H)=\log2=1\ \mathrm{bit} \]

With two equally likely signals \(I(H;U)\le\mathcal H(H)=\log2=1\ \mathrm{bit}\)

With two equally likely signals \({\color{#800000}p_U(u)}=\tfrac12p(u\mid h_1)+\tfrac12p(u\mid h_2)\)

Joint distribution of environmental signal and response

entropy marginal

entropy conditional

binary conditionals

binary mixture

signal marginal

response marginal

The entropy difference measures statistical dependence

Factor the joint density inside the logarithm

\[ \log\frac{p_{HU}(h,u)}{p_H(h)p_U(u)}=\log p_{U\mid H}(u\mid h)-\log p_U(u) \]

Integrating over \(h\) gives the marginal \(p_U(u)\)

\[ I(H;U)=\underbrace{-\int p_U(u)\log p_U(u)\,du}_{\mathcal H(U)}-\underbrace{\left[-\iint p_{HU}(h,u)\log p_{U\mid H}(u\mid h)\,dh\,du\right]}_{\mathcal H(U\mid H)} \]

\[ I(H;U)=\mathcal H(U)-\mathcal H(U\mid H) \]

Exchanging signal and response gives the same information

\[ I(H;U)=\mathcal H(H)-\mathcal H(H\mid U) \]

Discrete input: remaining uncertainty, averaged over responses

\[ \mathcal H(H\mid U)=-\int p_U(u)\sum_h P(h\mid u)\log P(h\mid u)\,du \]

For two equally likely input values

\[ \mathcal H(H)=-2\left(\tfrac12\log\tfrac12\right)=\log 2, \qquad I(H;U)=\log 2-\underbrace{\mathcal H(H\mid U)}_{\ge0}\le\log 2 \]

\[ I(H;U)=\log 2-\underbrace{\mathcal H(H\mid U)}_{\ge0}\le\log 2 \]

Identical conditional responses: \(I=0\)
Well-separated Gaussian responses: \(\mathcal H(H\mid U)\to0\)

Shannon, Bell Syst. Tech. J. 27, 379 (1948)

How much information can a noisy response carry?

E.g., same transcription-factor concentration (\(h\)) can lead to different gene expression levels (\(u\))

Expression \(u\)

Concentration \(h\)

\[ {\color{#800000}\bar u(h)} \]

\[ h_0 \]

Same concentration
different expression levels

\[ {\color{#0366c8}\sigma_U(h_0)} \]

Responses at \(h_0\)

\[ p_{U\mid H}(u\mid h_0) \]

\[ {\color{#800000}\bar u(h_0)} \]

Expression \(u\)

\[ {\color{#0366c8}\sigma_U(h_0)} \]

\[ p_{U\mid H}(u\mid h)=\mathcal N\!\left({\color{#800000}\bar u(h)},{\color{#0366c8}\sigma_U^2(h)}\right) \]

What is the maximum information a response can carry?

\[ C=\sup_{p_H} I(H;U) \]

Maximum information
this response can carry

Fix the response, noise and input range
Vary the input distribution \(p_H(h)\)

Derivation: small-noise information and its maximum
Tkačik, Callan & Bialek, Phys. Rev. E 78, 011910 (2008)

\[ p_{U\mid H}(u\mid h)=\mathcal N\!\left({\color{#800000}\bar u(h)},{\color{#0366c8}\sigma_U^2(h)}\right) \]

Concentration \(h\)

\[ \begin{aligned}I(H;U)\simeq&-\int p_H(h)\log\frac{p_H(h)}{|\bar u^{\prime}(h)|}\,dh+\\&-\int p_H(h)\log[\sqrt{2\pi e}\sigma_U(h)]\,dh\end{aligned} \]

For small noise, information depends on the response derivative

Channel capacity: \(C=\sup_{p_H} I(H;U)\)

What is the input distribution that maximizes mutual information?

\[ {\color{#0366c8}p_H^*(h)}=\frac{1}{Z}\,\frac{|{\color{#800000}\bar u^{\prime}(h)}|}{\sigma_U(h)} \]

\(Z\) makes the total probability equal to one

\[ \delta h(h)=\frac{\sigma_U(h)}{|\bar u^{\prime}(h)|}\qquad\Rightarrow\qquad p_H^*(h)\propto\frac{1}{\delta h(h)} \]

Larger optimal density where nearby inputs
can be distinguished more precisely

For constant noise, the optimal input density follows the response slope

The response is steeper at smaller \(h\), so the optimal input density is larger there

Concentration \(h\)

reg mean

reg selected noise overlay

reg conditional

Optimal input density for the saturating response with constant small noise

Average the noise entropy over the input distribution

At a fixed input \(h\), the response noise is Gaussian

\[ -\log p_{U\mid H}(u\mid h)=\log[\sqrt{2\pi}\sigma_U(h)]+\frac{[u-\bar u(h)]^2}{2\sigma_U^2(h)} \]

\[ \mathcal H(U\mid H=h)=-\mathbb{E}[\log p_{U\mid H}(U\mid h)\mid H=h] = \log[\sqrt{2\pi e}\sigma_U(h)] \]

The expectation over responses at fixed environment is the response noise

\[ \mathbb{E}[(U-\bar u(h))^2\mid H=h]=\sigma_U^2(h) \]

\[ \begin{aligned}\mathcal H(U\mid H=h)&=\log[\sqrt{2\pi}\sigma_U(h)]+\frac{\sigma_U^2(h)}{2\sigma_U^2(h)}\\&=\log[\sqrt{2\pi e}\sigma_U(h)]\end{aligned} \]

Average this noise entropy over the input density \(p_H(h)\)

\[ \mathcal H(U\mid H)=\int p_H(h)\log[\sqrt{2\pi e}\sigma_U(h)]\,dh \]

Tkačik, Callan & Bialek, Phys. Rev. E 78, 011910 (2008)

Average this noise entropy over the input density \(p_H(h)\)

\[ \mathcal H(U\mid H)=\int p_H(h)\log[\sqrt{2\pi e}\sigma_U(h)]\,dh \]

Small noise makes the response entropy tractable

The other term in mutual information is the entropy of all responses

\[ p_U(u)=\int p_H(h)p_{U\mid H}(u\mid h)\,dh \]

\[ \mathcal H(U)=-\int p_H(h)\,\mathbb{E}[\log p_U(U)\mid H=h]\,dh \]

Write each response as its mean plus a fluctuation

\[ U=\bar u(h)+\delta u,\quad \mathbb{E}[\delta u\mid h]=0,\quad \mathbb{E}[\delta u^2\mid h]=\sigma_U^2(h) \]

Expand across the narrow noise width

\[ \mathbb{E}[f(U)\mid h]=f(\bar u(h))+\tfrac12\sigma_U^2(h)f^{\prime\prime}(\bar u(h))+\cdots \]

Set \(f=\log p_U\) and keep the first term
\(p_U\) must vary little across the noise width

\[ \mathcal H(U)\simeq-\int p_H(h)\log p_U(\bar u(h))\,dh \]

\[ Y=\bar u(H),\qquad p_Y(\bar u(h))|\bar u^{\prime}(h)|\,dh=p_H(h)\,dh \]

Monotone mean and small noise away from endpoints: \(p_U\simeq p_Y\)

\[ p_Y(\bar u(h))=\frac{p_H(h)}{|\bar u^{\prime}(h)|} \]

\[ \mathcal H(U)\simeq-\int p_H(h)\log\frac{p_H(h)}{|\bar u^{\prime}(h)|}\,dh \]

Tkačik, Callan & Bialek, Phys. Rev. E 78, 011910 (2008)

The response slope and noise determine the optimal input distribution

Fix \(\bar u(h)\), \(\sigma_U(h)\) and \([h_{\min},h_{\max}]\); vary the input distribution \(p_H\)

\[ I_{\rm app}[p_H]=\int p_H(h)\log\frac{|\bar u^{\prime}(h)|}{\sqrt{2\pi e}\,\sigma_U(h)\,p_H(h)}\,dh \]

Impose \(\int p_H(h)\,dh=1\) while maximizing this mutual information

\[ \Phi[p_H]=I_{\rm app}[p_H]-\lambda\!\left(\int p_H(h)\,dh-1\right) \]

Set the variation of \(\Phi\) to zero and normalize the resulting distribution

\[ \frac{\delta\Phi}{\delta p_H(h)}=\log\frac{|\bar u^{\prime}(h)|}{\sqrt{2\pi e}\,\sigma_U(h)\,p_H(h)}-1-\lambda=0 \]

\[ p_H^*(h)=\frac{1}{Z}\frac{|\bar u^{\prime}(h)|}{\sigma_U(h)},\qquad Z=\int_{h_{\min}}^{h_{\max}}\frac{|\bar u^{\prime}(h)|}{\sigma_U(h)}\,dh \]

The second variation is negative, so this stationary distribution maximizes \(I_{\rm app}\)

\[ \delta^2 I_{\rm app}=-\int\frac{[\delta p_H(h)]^2}{p_H(h)}\,dh<0 \]

Convert response noise into the corresponding uncertainty in the input

\[ \Delta\bar u\simeq\bar u^{\prime}(h)\,\Delta h\quad\Rightarrow\quad\delta h(h)=\frac{\sigma_U(h)}{|\bar u^{\prime}(h)|},\qquad p_H^*(h)=\frac{1}{Z\,\delta h(h)} \]

For constant noise, the optimal input distribution makes the response density uniform

\[ p_U^*(\bar u(h))\simeq\frac{p_H^*(h)}{|\bar u^{\prime}(h)|}=\frac{1}{Z\,\sigma_U(h)} \]

Small noise, monotone mean response; endpoint smoothing neglected

Tkačik, Callan & Bialek, Phys. Rev. E 78, 011910 (2008)

Capacity grows with the number of resolvable input levels

Two representative inputs from a continuous range

\[ \sigma_U(h) \]

\[ p(u\mid h_1) \]

\[ {\color{#800000}p(u\mid h_2)} \]

Response \(u\)

\[ \sigma_U(h)/2 \]

\[ p(u\mid h_1) \]

\[ {\color{#800000}p(u\mid h_2)} \]

Response \(u\)

Less overlap makes nearby inputs easier to distinguish

Capacity depends on the distinguishable input intervals across the full input range

\[ C\simeq\log\!\left[\int_{h_{\min}}^{h_{\max}}\frac{dh}{\sqrt{2\pi e}\,\delta h(h)}\right], \quad \delta h(h)=\frac{\sigma_U(h)}{|\bar u^{\prime}(h)|} \]

Smooth monotone response · small noise · fixed input range

\[ \delta h(h)=\frac{\sigma_U(h)}{|\bar u^{\prime}(h)|} \]

\[ \sigma_U\longrightarrow\sigma_U/2\quad\Longrightarrow\quad C\longrightarrow C+\log2 \]

Halving the noise adds one bit of information capacity

Compare uniform and optimized input use

\[ I\!\left(H(t);U(t)\right) \]

Information available at one time

A response record can also carry information across times

Derivation: capacity from the optimal input distribution
Tkačik, Callan & Bialek, Phys. Rev. E 78, 011910 (2008)

\[ \sigma_U\longrightarrow\sigma_U/2\quad\Longrightarrow\quad C\longrightarrow C+\log2 \]

Halving the noise adds one bit of information capacity

capacity wide

capacity narrow

Substitute the optimal input density to obtain capacity

Finite input range · monotone smooth mean · small Gaussian noise

Negligible boundary effects

\[ p_H^*(h)=K\frac{|\bar u^{\prime}(h)|}{\sigma_U(h)} \]

\[ 1=K\int_{h_{\min}}^{h_{\max}}\frac{|\bar u^{\prime}(h)|}{\sigma_U(h)}\,dh=KZ\quad\Longrightarrow\quad K=1/Z \]

\[ I_{\rm app}[p_H^*]=\int p_H^*(h)\log\frac{|\bar u^{\prime}(h)|}{\sqrt{2\pi e}\sigma_U(h)[|\bar u^{\prime}(h)|/(Z\sigma_U(h))]}\,dh \]

The logarithm becomes constant and \(\int p_H^*(h)\,dh=1\)

\[ C_{\rm app}=I_{\rm app}[p_H^*]=\log\frac{Z}{\sqrt{2\pi e}} \]

Every other input density gives less information

\[ \begin{aligned}I_{\rm app}[p_H]&=C_{\rm app}-\int p_H(h)\log\frac{p_H(h)}{p_H^*(h)}\,dh\\&=C_{\rm app}-D_{\rm KL}(p_H\Vert p_H^*)\le C_{\rm app}\end{aligned} \]

Relative entropy is nonnegative, with equality at the normalized optimum

For a linear response with constant noise

\[ \bar u(h)=a h+b,\quad\sigma_U(h)=\sigma,\quad L=h_{\max}-h_{\min} \]

\[ Z=|a|L/\sigma,\qquad p_H^*=1/L,\qquad C_{\rm app}=\log\frac{|a|L}{\sqrt{2\pi e}\sigma} \]

A larger output range \(|a|L\) relative to noise gives higher capacity

Tkačik, Callan & Bialek, Phys. Rev. E 78, 011910 (2008)

When does an interaction carry information across timescales?

Two molecules switch between two states \(A\) and \(B\) on two different timescales \(\tau_1\) and \(\tau_2\)

Baseline switching rates: \(1/\tau_1\) and \(1/\tau_2\)

Interaction \(c\): first molecule in state \(B_1\) accelerates \(A_2\to B_2\) for the second

\[ \frac{1}{\tau_2}\quad\longrightarrow\quad\frac{1+c}{\tau_2}\quad\text{when }X_1=B_1 \]

How much does the current state of \(X_2\) tell us about \(X_1\)?

\[ I_{12}=I(X_1;X_2)=\sum_{x_1,x_2}p_{12}(x_1,x_2)\log_2\frac{p_{12}(x_1,x_2)}{p_1(x_1)p_2(x_2)} \]

Nicoletti & Busiello, Phys. Rev. X 14, 021007 (2024); J. Stat. Mech. 124004 (2025)

Two molecules, each switching between states A and B at its own timescale

State B1 increases the rate of the A2 to B2 transition in the second molecule

Changing the timescales changes the information

Keep the interaction fixed: \(X_1\longrightarrow X_2,\quad c=10\)

Direct interaction   \(\tau_1\ll\tau_2\)

\(X_1\) switches many times before \(X_2\) changes,
so \(X_2\) inherits an effective switching rate

\[ p_{12}\simeq p_1p_2,\qquad I_{12}\to0 \]

Feedback interaction   \(\tau_1\gg\tau_2\)

\(X_2\) relaxes while \(X_1\) stays fixed,
so its distribution becomes a mixture

\[ p_{12}\simeq p_1p_{2|1}^{\mathrm{st}},\qquad I_{12}>0 \]

\(X_1\longrightarrow X_2,\qquad c=10\)
Derivation: the two-layer calculation
Nicoletti & Busiello, Phys. Rev. X 14, 021007 (2024); J. Stat. Mech. 124004 (2025)

Fast regulator: the first three simulated information values approach zero

At fixed interaction strength, stationary mutual information rises as the regulator becomes slower

Fokker–Planck operators for continuous systems

For simplicity, take continuous degrees of freedom (dofs) \(x_1,x_2\)
The interaction remains \(X_1\to X_2\) and the derivation is the same for discrete systems

\[ \partial_t p_{12}=\left[\frac{\mathcal L_1}{\tau_1}+\frac{\mathcal L_2(x_1)}{\tau_2}\right]p_{12} \]

\[ \mathcal L_i p=-\partial_{x_i}\!\left[b_i p\right]+\partial_{x_i}^{,2}\!\left[D_i p\right] \]

\(b_i\): drift   \(D_i\): diffusion
\(\mathcal L_i\) acts only \(x_i\), while other states enter as parametric dependencies

For the two-state example, use transition operators and sums in the same calculation

Nicoletti & Busiello, Phys. Rev. X 14, 021007 (2024); J. Stat. Mech. 124004 (2025)

Fast-to-slow interaction: averaging defines the effective slow operator

Direct interaction: \(X_1\to X_2\), with \(\tau_1\ll\tau_2\)

\[ s=\frac{t}{\tau_2},\quad\epsilon=\frac{\tau_1}{\tau_2}\ll1,\qquad\partial_s p_{12}=\left[\epsilon^{-1}\mathcal L_1+\mathcal L_2(x_1)\right]p_{12} \]

\[ p_{12}=p_{12}^{\rm eff}+\epsilon p_{12}^{(1)}+\cdots\quad\Rightarrow\quad\mathcal L_1p_{12}^{\rm eff}=0 \]

\[ p_{12}^{\rm eff}=p_1^{\rm st}(x_1)\,p_2(x_2,s),\qquad \mathcal L_1p_1^{\rm st}=0 \]

The autonomous fast regulator reaches its stationary distribution: \(\int dx_1\,p_1^{\rm st}=1\)

Integrate the next order over \(x_1\); probability conservation removes the fast operator

\[ \int dx_1\,\mathcal L_1p_{12}^{(1)}=0 \]

\[ \partial_s p_2=\underbrace{\int dx_1\,\mathcal L_2(x_1)\!\left[p_1^{\rm st}(x_1)p_2\right]}_{\displaystyle \mathcal L_2^{\rm eff}p_2} \]

The effective operator averages over the full distribution of the fast input

\[ p_{12}^{\rm eff}=p_1^{\rm st}(x_1)\,p_2^{\rm eff}(x_2,s)\qquad\Rightarrow\qquad I_{12}=0 \]

Nicoletti & Busiello, Phys. Rev. X 14, 021007 (2024); J. Stat. Mech. 124004 (2025)

Slow-to-fast interaction: same operator at frozen slow state

Feedback interaction: \(X_1\to X_2\), with \(\tau_2\ll\tau_1\)

\[ s=\frac{t}{\tau_1},\quad\epsilon=\frac{\tau_2}{\tau_1}\ll1,\qquad\partial_s p_{12}=\left[\mathcal L_1+\epsilon^{-1}\mathcal L_2(x_1)\right]p_{12} \]

Hold \(x_1\) fixed while \(x_2\) relaxes under the original operator \(\mathcal L_2(x_1)\)

\[ \mathcal L_2(x_1)\,p_{2|1}^{\rm st}(x_2|x_1)=0,\qquad\int dx_2\,p_{2|1}^{\rm st}(x_2|x_1)=1 \]

Each frozen value of \(x_1\) can give a different stationary response

\[ p_{12}(x_1,x_2,s)=p_{2|1}^{\rm st}(x_2|x_1)\,p_1(x_1,s) \]

\[ \partial_s p_1=\int dx_2\,\mathcal L_1\!\left[p_{2|1}^{\rm st}p_1\right]=\mathcal L_1p_1 \]

The slow operator is also unchanged; the fast conditional density integrates to one

\[ \begin{aligned}p_2(x_2,s)&=\int dx_1\,p_1(x_1,s)p_{2|1}^{\rm st}(x_2|x_1)\\[4pt]I_{12}&=\int dx_1\,p_1(x_1,s)D_{\rm KL}\!\left[p_{2|1}^{\rm st}(\cdot|x_1)\,\Vert\,p_2\right]\end{aligned} \]

Information is positive when the conditional response varies across the populated input states

Nicoletti & Busiello, Phys. Rev. X 14, 021007 (2024); J. Stat. Mech. 124004 (2025)

Information generation by feedback interactions

\[ \tau_1\ll\tau_2\ll\tau_3 \]

Direct chain

Feedback chain

\(1\to2\to3\)
Average over
the faster input

\(3\to2\to1\)
Respond to
the slower state

\[ I_{12}=I_{13}=I_{23}=0 \]

\(I_{12},\ I_{13},\ I_{23} \ge 0\)

Derivation: direct and feedback chains
Nicoletti & Busiello, Phys. Rev. X 14, 021007 (2024); J. Stat. Mech. 124004 (2025)

Three-layer direct chain from fastest layer 1 to slowest layer 3

Three-layer feedback chain from slowest layer 3 to fastest layer 1

Path 2 to 3 to 1, through a intermediate layer slower than the source

Minimal propagation path 3 to 1 to 2, through a fast intermediate layer

preview-figure4a-topology

preview-figure4d-topology

Path 2 to 3 to 1, through a intermediate layer slower than the source

Minimal propagation path 3 to 1 to 2, through a fast intermediate layer

Direct chains averaging and feedback chains conditional dependencies

Direct chain: \(1\to2\to3\), with \(\tau_1\ll\tau_2\ll\tau_3\)

\[ \mathcal L_2^{\rm eff}p_2=\int dx_1\,\mathcal L_2(x_1)\!\left[p_1^{\rm st}(x_1)p_2\right] \]

\[ \mathcal L_3^{\rm eff}p_3=\int dx_2\,\mathcal L_3(x_2)\!\left[p_2^{\rm st,eff}(x_2)p_3\right] \]

Each layer receives an average over its faster input

\[ p_{123}^{\rm eff}=p_1^{\rm st}\,p_2^{\rm st,eff}\,p_3^{\rm eff}(t),\qquad I_{12}=I_{13}=I_{23}=0 \]

Feedback chain: \(3\to2\to1\), with \(\tau_1\ll\tau_2\ll\tau_3\)

\[ \mathcal L_1(x_2)p_{1|2}^{\rm st}=0,\qquad \mathcal L_2(x_3)p_{2|3}^{\rm st}=0 \]

Both operators keep their original form, with each slower input held fixed

\[ p_{123}=p_{1|2}^{\rm st}(x_1|x_2)\,p_{2|3}^{\rm st}(x_2|x_3)\,p_3(x_3,t) \]

\[ I(X_1;X_3|X_2)=0 \]

Nicoletti & Busiello, Phys. Rev. X 14, 021007 (2024); J. Stat. Mech. 124004 (2025)

Information propagation by direct interactions

\[ \tau_1\ll\tau_2\ll\tau_3 \]

\(2\to3\to1\)

\(3\to1\to2\)

The slow dof \(3\) averages
the fluctuations of \(2\)

The fast dof \(1\) retains
the slow state of \(3\)

Minimal propagation path

Only \(1\) and \(3\) share information

All three pairs can share information

\(f\): strength of the slow-to-fast interaction

Derivation: minimal propagation paths
Nicoletti & Busiello, Phys. Rev. X 14, 021007 (2024); J. Stat. Mech. 124004 (2025)

Propagation Path

Minimal Propagation Path

Path 2 to 3 to 1, through a intermediate layer slower than the source

Minimal propagation path 3 to 1 to 2, through a fast intermediate layer

Only the information entry between layers 1 and 3 is nonzero

All three pairwise information entries can be nonzero

Information I13 increases with the feedback strength

The three pairwise information curves increase with feedback strength

Path 2 to 3 to 1, through a intermediate layer slower than the source

Minimal propagation path 3 to 1 to 2, through a fast intermediate layer

Path 2 to 3 to 1, through a intermediate layer slower than the source

Minimal propagation path 3 to 1 to 2, through a fast intermediate layer

Minimal propagation path 3 to 1 to 2, through a fast intermediate layer

Three-layer direct chain from fastest layer 1 to slowest layer 3

Three-layer direct chain from fastest layer 1 to slowest layer 3

Three-layer feedback chain from slowest layer 3 to fastest layer 1

A direct interaction can preserve dependence on a slower dof

\(2\to3\to1\), with \(\tau_1\ll\tau_2\ll\tau_3\)

Layer 3 averages the input from 2; layer 1 responds to the frozen state of 3

\[ \mathcal L_3^{\rm eff}p_3=\int dx_2\,\mathcal L_3(x_2)\!\left[p_2^{\rm st}(x_2)p_3\right] \]

\[ p_{123}^{\rm eff}=p_{1|3}^{\rm st}(x_1|x_3)\,p_2^{\rm st}(x_2)\,p_3^{\rm eff}(x_3,t) \]

\[ I_{12}=I_{23}=0 \]

\(3\to1\to2\), with \(\tau_1\ll\tau_2\ll\tau_3\)

Average over the fast layer 1 while keeping the slow source \(x_3\) fixed

\[ \mathcal L_2^{\rm eff}(x_3)p_2=\int dx_1\,\mathcal L_2(x_1)\!\left[p_{1|3}^{\rm st}(x_1|x_3)p_2\right] \]

The effective operator for layer 2 still depends on \(x_3\)

\[ p_{123}^{\rm eff}=p_{1|3}^{\rm st}(x_1|x_3)\,p_{2|3}^{\rm st,eff}(x_2|x_3)\,p_3(x_3,t) \]

At fixed \(x_3\), the two faster layers are independent

\[ p_{12|3}^{\rm eff}=p_{1|3}^{\rm st}\,p_{2|3}^{\rm st,eff}\qquad\Rightarrow\qquad I(X_1;X_2|X_3)=0 \]

Averaging over their shared slow source can correlate the two layers

\[ p_{12}^{\rm eff}=\int dx_3\,p_{1|3}^{\rm st}(x_1|x_3)\,p_{2|3}^{\rm st,eff}(x_2|x_3)\,p_3(x_3,t) \]

Nicoletti & Busiello, Phys. Rev. X 14, 021007 (2024); J. Stat. Mech. 124004 (2025)

Timescales give information propagation an intrinsic directionality

\[ \tau_1\ll\tau_2\ll\tau_3 \]

Feedback interactions generate information

A slow source changes little
while the fast layer responds

Direct interactions can propagate it

The intermediate layer responds quickly
while the slow source persists

Change the ordering: which slow states still reach each layer?
Derivation: the general decomposition and path criterion
Nicoletti & Busiello, Phys. Rev. X 14, 021007 (2024); J. Stat. Mech. 124004 (2025)

A slow source sends information through a fast intermediate layer to another layer faster than the source

Information generated from the slow third layer propagates toward faster layers, with colored arrows showing the directed paths

Effective operators determine which slow dependencies survive

\(\tau_1\ll\tau_2\ll\cdots\ll\tau_N\): relax the fastest layer, average it, then continue

\[ \mathcal L_1\,p_{1|\rho(1)}^{\rm st}=0,\qquad\int dx_1\,p_{1|\rho(1)}^{\rm st}=1 \]

\[ \mathcal L_\mu^{\rm eff} f=\int dx_1\cdots dx_{\mu-1}\,\mathcal L_\mu\!\left[\left(\prod_{\nu<\mu}p_{\nu|\rho(\nu)}^{\rm st,eff}\right)f\right] \]

\(f\) is a density of the remaining layers; their slow states stay fixed in the average

At each stage, find the normalized stationary density of the effective operator

\[ \mathcal L_\mu^{\rm eff}\,p_{\mu|\rho(\mu)}^{\rm st,eff}=0,\qquad\int dx_\mu\,p_{\mu|\rho(\mu)}^{\rm st,eff}=1 \]

\[ p_{1,\ldots,N}^{\rm eff}=p_N^{\rm eff}(x_N,t)\prod_{\mu=1}^{N-1}p_{\mu|\rho(\mu)}^{\rm st,eff}\!\left(x_\mu\mid\{x_\nu\}_{\nu\in\rho(\mu)}\right) \]

\[ \partial_t p_N^{\rm eff}=\frac{1}{\tau_N}\mathcal L_N^{\rm eff}p_N^{\rm eff} \]

\(\rho(\mu)\) contains the slower layers retained as conditions

A propagation path starts at a slower source and contains a direct interaction

\[ \nu\to\cdots\to\mu:\quad\nu>\mu,\qquad \exists\ (\alpha\to\beta)\ \text{with }\alpha<\beta \]

It is minimal when every internal layer is faster than the target \(\mu\)

\[ \rho(\mu)=\left\{\nu>\mu:\ \nu\to\mu\text{ is an edge, or an mPP }\nu\to\mu\text{ exists}\right\} \]

These paths determine which slow states enter the effective stationary densities

Nicoletti & Busiello, Phys. Rev. X 14, 021007 (2024); J. Stat. Mech. 124004 (2025)

Why do some biological responses weaken when a stimulus repeats?

Habituation: a progressive decrease in the response to repeated stimulation

Black: mean response \(\langle U\rangle\) · Gray: storage buildup

A long pause allows the response to recover

Idea: habituation from slow negative feedback

Nicoletti et al., eLife 13, RP99767 (2025)

Original eLife figure: response train

Original eLife figure: recovery

Slow storage changes the response to the next stimulus

\(H\): input · \(R\): receptor · \(U\): readout · \(S\): storage

\[ \tau_U\ll\tau_R\ll\tau_S\approx\tau_H \]

From the previous considerations:
→ Fast readout
→ Fast but slow enough receptor
→ Storage plays the role of slow regulation

Average the fast response

\[ \begin{gathered}p^\mathrm{eff}(u,r,s,h;t)\\=p_{U|R}^\mathrm{eff,st}(u\mid r)\,p_{R|S,H}^\mathrm{eff,st}(r\mid s,h)\,F(s,h;t)\end{gathered} \]

Receptor and readout follow
the current input and storage

\(F\) changes as storage builds and decays

Derivation: fast response laws and slow storage
Nicoletti et al., eLife 13, RP99767 (2025)

\[ \begin{gathered}p^\mathrm{eff}(u,r,s,h;t)\\=p_{U|R}^\mathrm{eff,st}(u\mid r)\,p_{R|S,H}^\mathrm{eff,st}(r\mid s,h)\,p^\mathrm{eff}_{S,H}(s,h;t)\end{gathered} \]

The timescale-separated solution is controlled by the joint evolution of \(S\) and \(H\)

Original eLife figure: architecture

Original eLife figure: memory pause

Fast receptor activity determines the readout distribution

Freeze storage \(s\) and input \(h\) while the fast variables relax

\[ a(h)=e^{-\beta\Delta E}(1+e^{\beta h}),\qquad b(s)=1+e^{\beta\kappa\sigma s/N_S} \]

\[ \begin{gathered}a(h)(1-\theta)=b(s)\theta\\p_{R|S,H}^{\mathrm{eff,st}}(r=1\mid s,h)=\theta(s,h)=\frac{a(h)}{a(h)+b(s)}\end{gathered} \]

\(\beta\): inverse temperature; \(\Delta E\): receptor barrier; \(\kappa\): inhibition strength
\(\sigma\): cost per storage unit; \(N_S\): storage capacity

\[ \begin{gathered}\mu_r=e^{-\beta(V-gr)},\qquad\mu_rp_{U|R}^{\mathrm{eff,st}}(u\mid r)=(u+1)p_{U|R}^{\mathrm{eff,st}}(u+1\mid r)\\p_{U|R}^{\mathrm{eff,st}}(u\mid r)=e^{-\mu_r}\frac{\mu_r^u}{u!}\end{gathered} \]

Average the Poisson readout laws over receptor activity

\[ \begin{aligned}p_{U|S,H}^{\mathrm{eff,st}}(u\mid s,h)&=[1-\theta(s,h)]\operatorname{Pois}(u;\mu_0)+\theta(s,h)\operatorname{Pois}(u;\mu_1)\\\bar u(s,h)&=[1-\theta(s,h)]\mu_0+\theta(s,h)\mu_1\end{aligned} \]

Equal bare receptor-pathway rates, in units of \(\tau_R^{-1}\); frozen slow variables

\(V\): readout production cost; \(g\): receptor-induced reduction (paper’s \(c\))
Constant birth and linear death; unbounded readout counts; frozen slow variables

Nicoletti et al., eLife 13, RP99767 (2025)

Averaging fast activity gives the slow storage dynamics

Average the storage birth rate over fast readout activity

\[ \begin{gathered}w_{s\to s+1}(u)=\frac{u e^{-\beta\sigma}}{\tau_S}\\b_s(h)=\frac{\bar u(s,h)e^{-\beta\sigma}}{\tau_S},\qquad d_s=\frac{s}{\tau_S}\end{gathered} \]

The conditional mean sets storage production, and each stored unit decays spontaneously

Evolve the joint distribution of storage and input

\[ \begin{aligned}\partial_tp^{\mathrm{eff}}_{S,H}(s,h;t)&=\mathcal L_H p^{\mathrm{eff}}_{S,H}(s,h;t)+b_{s-1}(h)p^{\mathrm{eff}}_{S,H}(s-1,h;t)+d_{s+1}p^{\mathrm{eff}}_{S,H}(s+1,h;t)\\&\quad-[b_s(h)+d_s]p^{\mathrm{eff}}_{S,H}(s,h;t)\end{aligned} \]

\[ \frac{d\langle S\rangle}{dt}=\langle b_S(H)-d_S\rangle \]

Responses increase production, while surviving storage changes the next response, establishing a feedback loop

\(p^{\mathrm{eff}}_{S,H}(s,h;t)\); \(\mathcal L_H\) evolves the input
\(0\le s\le N_S\); reflecting boundaries \(d_0=0\), \(b_{N_S}=0\); evolve \(p^{\mathrm{eff}}_{S,H}\) numerically

Nicoletti et al., eLife 13, RP99767 (2025)

Does memory improve the information carried by the response?

At each time, compare \(H\) and \(U\) across repeated trials

Depending on the parameters,
slow storage can lead to strong habituation to repeated environmental inputs

But despite weaker responses,
their information with the environment increases

\[ I_{U,H}(t)=\mathcal{H}[U_t]-\mathcal{H}[U_t\mid H_t] \]

Derivation: information from the response distribution
Nicoletti et al., eLife 13, RP99767 (2025)

Original eLife figure: response dynamics

Original eLife figure: information dynamics

Marginalize the slow state to calculate readout information

Sum over internal states to obtain the input–response distribution

\[ p^{\mathrm{eff}}_{U,H}(u,h;t)=\sum_{s=0}^{N_S}\sum_{r=0}^1p_{U|R}^{\mathrm{eff,st}}(u\mid r)p_{R|S,H}^{\mathrm{eff,st}}(r\mid s,h)p^{\mathrm{eff}}_{S,H}(s,h;t) \]

\[ p^{\mathrm{eff}}_U(u;t)=\int_0^\infty p^{\mathrm{eff}}_{U,H}(u,h;t)\,dh,\qquad p_H(h;t)=\sum_{u=0}^\infty p^{\mathrm{eff}}_{U,H}(u,h;t) \]

Readout counts are discrete; input values are continuous

Calculate information, then compare matching stimulus phases

\[ \begin{aligned}I_{U,H}(t)&=\sum_{u=0}^{\infty}\int_0^\infty dh\,p^{\mathrm{eff}}_{U,H}(u,h;t)\log\frac{p^{\mathrm{eff}}_{U,H}(u,h;t)}{p^{\mathrm{eff}}_U(u;t)p_H(h;t)}\\&=\mathcal{H}[U_t]-\int_0^\infty dh\,p_H(h;t)\,\mathcal{H}[U_t\mid H_t=h]\end{aligned} \]

\[ \Delta I_{U,H}=I_{U,H}^{\mathrm{hab}}-I_{U,H}^{\mathrm{in}},\qquad\Delta\langle U\rangle=\langle U\rangle^{\mathrm{hab}}-\langle U\rangle^{\mathrm{in}} \]

\(\Delta I_{U,H}>0\): more information; \(\Delta\langle U\rangle<0\): weaker response

Same input statistics at the compared phases; natural logarithms; timescale-separation approximation

Nicoletti et al., eLife 13, RP99767 (2025)

Information gain peaks at intermediate habituation

Habituation depends on the thermal noise \(\beta\) and the storage cost \(\sigma\) at fixed input statistics

Gain in information

Response change: habituated − initial

The peak lies where the response reduction is intermediate

Receptor cycling dissipates heat

Pareto front determines optimal information-dissipation trade-off

\[ \delta Q_R^{\mathrm{st}}=\beta\!\left(H_{\mathrm{st}}+\kappa\sigma\langle S\rangle_{\mathrm{st}}/N_S\right) \]

Heat per receptor cycle, in thermal units

Derivation: the information–energy optimization
Nicoletti et al., eLife 13, RP99767 (2025)

Published Pareto front: stationary readout information versus receptor heat plus normalized storage energy consumption

Original eLife figure: information gain

Original eLife figure: habituation strength

Original eLife figure: optimal comparison

A stationary information–energy tradeoff selects operating parameters

Fix the input statistics: \(p_H^{\mathrm{st}}(h)=H_{\mathrm{st}}^{-1}e^{-h/H_{\mathrm{st}}}\)

\[ I^{\mathrm{st}}_{U,H}(\beta,\sigma),\qquad E_{\mathrm{tot}}^{\mathrm{st}}=\delta Q_R^{\mathrm{st}}+\frac{\tau_S J_{\mathrm{int}}^{\mathrm{st}}}{\sigma} \]

\[ \delta Q_R^{\mathrm{st}}=\beta\left(H_{\mathrm{st}}+\kappa\sigma\frac{\langle S\rangle_{\mathrm{st}}}{N_S}\right) \]

\(\delta Q_R\): receptor-cycle dissipation measure; \(J_{\mathrm{int}}\): net storage energy flux
\(\tau_S J_{\mathrm{int}}/\sigma\) is its normalized contribution

Choose the relative weight of information and energetic cost

\[ \mathcal L(\beta,\sigma)=\gamma I^{\mathrm{st}}_{U,H}(\beta,\sigma)-(1-\gamma)E_{\mathrm{tot}}^{\mathrm{st}}(\beta,\sigma) \]

\[ 0\le\gamma\le1,\qquad(\beta^*,\sigma^*)\in\operatorname*{arg\,max}_{\beta,\sigma}\mathcal L(\beta,\sigma) \]

Vary \(\gamma\) to find the best stationary information at different costs

Test the selected parameters under repeated stimulation

High information gain and intermediate response reduction
Storage cost varies at fixed \(\beta\); vertical axes are rescaled

Nicoletti et al., eLife 13, RP99767 (2025)

Published cross section comparing dynamic information gain and response change near the selected stationary parameter region

A slow internal state captures features of neural habituation

Repeated looming stimuli in a zebrafish larva

Top: recorded calcium activity · Bottom: model activity

Separate PCA projections of recorded and model activity

How do interactions within a neural circuit change encoding and response time?

Nicoletti et al., eLife 13, RP99767 (2025)

Original eLife figure: zebrafish traces

Original eLife figure: zebrafish pca

What makes neural responses distinguish their inputs?

Neural response \(\mathbf U=(U_E,U_I)\)

\[ \tau\dot U_\mu=-r_\mu U_\mu+\sum_{\nu=E,I}A_{\mu\nu}f(U_\nu)+H(t)\delta_{\mu E}+\sqrt{2D_\mu\tau}\,\eta_\mu(t) \]

\[ A=w\begin{pmatrix}1&-k\\1&-k\end{pmatrix} \]

\(w\): excitation strength
\(k\): relative inhibition

Input: a ground state \(h_0\) and \(M\) excited states \(h_i\)

\[ h_0\xrightleftharpoons[q_\downarrow]{q_\uparrow}h_i,\qquad i=1,\ldots,M \]

Timescale separation: input switching vs neural relaxation

\[ \tau_H=(q_\uparrow+q_\downarrow)^{-1}\qquad\text{and}\qquad\tau_U \]

Barzon, Busiello & Nicoletti, Phys. Rev. Lett. 134, 068403 (2025)

Excitatory and inhibitory populations with input to the excitatory population

Persistent inputs produce distinguishable neural states

\(p_i(\mathbf u)\): joint activity–input density   \(\pi_i\): stationary input probability

\[ \partial_t p_i=\tau^{-1}\mathcal L_i p_i+\tau_H^{-1}\sum_jQ_{ij}p_j \]

Fast input

\[ Q\mathbf p^{\mathrm{st}}=0,\qquad p_i^{\mathrm{st}}=\pi_i p_{\mathbf u}^{\mathrm{st,eff}} \]

\[ \mathcal L^{\mathrm{eff}}=\sum_i\pi_i\mathcal L_i=\mathcal L_{\bar h},\quad\bar h=\sum_i\pi_i h_i \]

Neurons respond to the average input \(\bar h\), so \(I(\mathbf U;H)\to0\)

Slow input

\[ \mathcal L_i p_i^{\mathrm{st}}=0,\qquad p_i^{\mathrm{st}}=\pi_i p_{\mathbf u\mid i}^{\mathrm{st}} \]

\[ p_{\mathbf u}^{\mathrm{st}}=\sum_i\pi_i p_{\mathbf u\mid i}^{\mathrm{st}} \]

Each input produces a neural distribution, so less overlap means larger \(I(\mathbf U;H)\)

Fast input: averaged drive \(\bar h\)

Slow input: distinct neural responses

\[ I(\mathbf U;H)=\sum_i\pi_iD_{\mathrm{KL}}\!\left[p_{\mathbf u\mid i}\Vert p_{\mathbf u}\right] \]

Barzon, Busiello & Nicoletti, Phys. Rev. Lett. 134, 068403 (2025)

Fast input traces and stationary neural activity

Slow input traces and separated conditional neural activity

Information increases with environmental persistence

E/I balance controls how well inputs can be distinguished

Approaching the edge of linear stability

\[ k_c=1-\frac rw,\qquad k>k_c \]

\[ k\to k_c:\quad I(\mathbf U;H)\longrightarrow\mathcal H(H) \]

Dotted curve: stability boundary \(k=k_c\)

Information depends on signal relative to noise

\[ p(\mathbf u\mid i)=\mathcal N(\mathbf m_i,\Sigma) \]

\[ d_{ij}^2=(\mathbf m_i-\mathbf m_j)^T\Sigma^{-1}(\mathbf m_i-\mathbf m_j) \]

With \(\delta=w(k-k_c)\),
mean separation grows as \(1/\delta\)
and noise width as \(1/\sqrt\delta\)

A nonlinear activation function turns the edge of stability into an information peak

Derivation: neural mixture bounds and E/I stability
Barzon, Busiello & Nicoletti, PRL 134, 068403 (2025)

the mutual information is maximized and tends to the environmental entropy

Information across excitation and relative inhibition

Linear information versus distance from the stability boundary

Nonlinear information peak near the linear stability boundary

Conditional overlap bounds neural information

Input \(i\) selects \(p_i(\mathbf u)=\mathcal N(\mathbf m_i,\Sigma)\) with probability \(\pi_i\)

\[ I(\mathbf U;H)=\mathcal{H}\!\left(\sum_i\pi_i p_i\right)-\sum_i\pi_i\mathcal{H}(p_i) \]

Compare mean separation with the common neural fluctuations

\[ \begin{gathered}d_{ij}^2=(\mathbf m_i-\mathbf m_j)^T\Sigma^{-1}(\mathbf m_i-\mathbf m_j)\\B_{ij}=\frac{d_{ij}^2}{8},\qquad K_{ij}=\frac{d_{ij}^2}{2}\end{gathered} \]

Bhattacharyya and KL distances give lower and upper bounds

\[ \begin{aligned}I_-&=-\sum_i\pi_i\log\!\sum_j\pi_j e^{-B_{ij}}\\I_+&=-\sum_i\pi_i\log\!\sum_j\pi_j e^{-K_{ij}}\end{aligned} \]

\[ I_-\le I(\mathbf U;H)\le I_+\le\mathcal{H}(\pi) \]

For equally spaced means, one parameter controls every pair distance

\[ \eta=\tfrac12\Delta\mathbf m^T\Sigma^{-1}\Delta\mathbf m,\qquad d_{ij}^2=2(i-j)^2\eta \]

\[ \eta\to0:\ I\to0,\qquad\eta\to\infty:\ I\to\mathcal{H}(\pi) \]

\(\pi=(1/2,1/4,1/4)\), so the ceiling is \(\tfrac32\log2\) nats = 1.5 bits

Stable linear circuit; slow Markov input; distinct finite labels; common positive-definite covariance

Barzon, Busiello & Nicoletti, PRL 134, 068403 (2025); Kolchinsky & Tracey, Entropy 19, 361 and correction 588 (2017)

The slow neural mode controls resolution and relaxation

\[ A=w\begin{pmatrix}1&-k\\1&-k\end{pmatrix},\qquad B=rI-A,\qquad r,w>0 \]

Linear activation: eigenvalues of \(B\) give decay rates in units of \(\tau^{-1}\)

\[ \lambda_1=r,\quad\lambda_2=r+w(k-1)=w(k-k_c),\qquad k_c=1-\frac rw,\quad k>k_c \]

\[ \mathbf m_i=B^{-1}h_i\mathbf e_E,\qquad B\Sigma+\Sigma B^T=2DI \]

Mean separation relative to noise controls distinguishability

\[ \eta=\tfrac12\Delta\mathbf m^T\Sigma^{-1}\Delta\mathbf m,\qquad k\downarrow k_c:\quad\eta\to\infty,\quad I(\mathbf U;H)\to\mathcal{H}(H) \]

\[ \tau_U=\frac{\tau}{\min\{r,w(k-k_c)\}}\longrightarrow\infty,\qquad\tau_H\gg\tau_U \]

Take the slow-input limit first; at \(k=k_c\) the stationary density is lost

Accessible interior edge: \(w>r\), \(k\ge0\); nonzero input spacing, isotropic \(D>0\)
Approach from the stable side, after taking the slow-input limit

Barzon, Busiello & Nicoletti, Phys. Rev. Lett. 134, 068403 (2025); Murphy & Miller, Neuron 61, 635 (2009)

Timescales shape how living systems sense and respond

Environment

\[ H(t) \]

\[ \tau_H \]

Receptor

\[ R(t) \]

\[ \tau_R \]

Response

\[ U(t) \]

\[ \tau_U \]

\[ \longrightarrow \]

\[ \longrightarrow \]

Average fluctuations

\[ \tau_R\ll T\ll\tau_H \]

Receptor memory shapes the
of independent measurements

Track a changing signal

\[ \tau_U\;\text{vs}\;\tau_H \]

Response memory balances
noise averaging and tracking

Information paths

Timescale ordering selects which dependencies survive

Habituation

Slow storage changes the response to later stimuli

Neural encoding

Interactions set the time needed to distinguish inputs

What does transduction of a hidden signal cost?

Hidden activity

\[ H(t) \]

Membrane fluctuations

\[ R(t) \]

Probe readout

\[ U(t) \]

\[ \longrightarrow \]

\[ \xrightarrow{\;a\;} \]

Tunable coupling

Information about the hidden signal

\[ I(H;U) \]

Information accessible to the probe

\[ I(R;U) \]

Tuning the coupling changes both the information acquired
and the energy dissipated

Can transduction adapt using only observed fluctuations?

Hidden chemical network produces a signaling molecule which binds a receptor and drives readout production

Hidden chemical network

Signaling molecule

Readout

Estimate information and changes in dissipation from observed records
Use them to adjust readout production

Observation time controls the accuracy of this adjustment

Thank you!

Giorgio Nicoletti

gnicolet@ictp.it
giorgionicoletti.github.io

Nicoletti & Busiello, Phys. Rev. Lett. 133, 158401 (2024); Nicoletti, Di Terlizzi & Busiello, Phys. Rev. Lett. 137, 127101 (2026)

Information propagation

\[ \tau_U\ll\tau_R\ll\tau_H, \, \tau_S \]

Timescale ordering selects
which dependencies survive

Hidden dof

Transducing dof

Readout

Information about the hidden signal \(I(H;U)\)

\[ I(H;U) \]

Information accessible to the probe \(I(R;U)\)
vs
Information about the hidden signal \(I(H;U)\)

\[ I(R;U) \]

Transducing information via non-reciprocal interactions requires energy
trade-offs, optimal strategies, and real-time balancing

Estimate information and changes in dissipation from observed records
Use them to adjust readout production

UNESCO

EPFL

Università degli Studi di Padova

Laboratory of Interdisciplinary Physics

Hidden chemical network produces a signaling molecule which binds a receptor and drives readout production