Standard graph neural networks rely on static topological structures to propagate information. Even when utilizing learned, task-dependent weights (a fixed geometry), the underlying mechanism remains a diffusion process. Over multiple layers, this fixed diffusion acts as a heat kernel, exponentially contracting the feature space and choking long-range signals a phenomenon widely known as oversquashing.
[!NOTE] Conformal Geometry Introduce the formulation in terms of conformal geometry, vector fields [[Lecture notes - Discrete Vector Bundles over Graphs]]
This note reformulates the problem by transitioning from algebraic graph theory (static limits, curvature) to the language of non-autonomous dynamical systems (Lyapunov exponents). We will demonstrate mathematically what happens when the geometry is no longer a static background object, but instead co-evolves dynamically with the node features.
We study the dynamics of GNNs under co-evolving geometry and show that their layer-to-layer Jacobian inherits a supplementary term imposed by geometry. We prove an impossibility theorem showing that there is no fixed geometry that can mimic this behavior. We further extend these findings through a Lyapunov Shift theorem: we prove that adaptive geometry generates a rank-1 geometric feedback channel that strictly counteracts diffusion-induced contraction. Finally, we prove that the same geometric feedback imposes mode delocalization, which allows the network to maintain persistent long-range sensitivity at the edge of chaos. On the architectural side, we propose a self-regulation mechanism designed to maintain the network in its persistent regime by controlling the Lyapunov shift.
[!NOTE] Existing approaches to non-dissipative deep dynamics typically enforce structural constraints on the propagation operator itself, such as orthogonality, unitarity, antisymmetry, or symplecticity, yielding approximately norm-preserving evolution by construction. In contrast, our architecture begins from a genuinely diffusive and contractive propagation operator. Criticality is not imposed through operator constraints but emerges through a closed-loop interaction between representation dynamics and an adaptive conformal geometry. The resulting feedback term continuously compensates diffusion-induced contraction, driving the dominant Lyapunov exponent toward zero without requiring orthogonal or energy-preserving parameterizations.
We first present the idea in the « linear » setting, i.e no nodewise non-linearity.
Consider a weighted, normalized Laplacian generated by a fixed node-wise conformal factor $\mu$: \(\mathcal{L}_\mu = I - D_\mu^{-1/2} A_\mu D_\mu^{-1/2}.\) Let us start with a simple linear GNN as testbed. A standard forward diffusion layer operates as: \(h^{(\ell+1)} = P_\mu h^{(\ell)}, \qquad P_\mu = I - \tau \mathcal{L}_\mu.\) For a deep network with a frozen geometry, the state at layer $K$ is simply the repeated application of this linear operator: \(h^{(K)} = P_\mu^K h^{(0)}.\) Because the eigenvalues of the normalized Laplacian are bounded by $0 \le \lambda_i(\mathcal{L}\mu) \le 2$, the eigenvalues of the propagation operator $P\mu$ (assuming $0 < \tau < 1$) are bounded strictly by 1: \(1 - 2\tau \le \eta_i \le 1 \implies \sigma_{\max}(P_\mu) \le 1.\) In the language of dynamical systems, the largest Lyapunov exponent $\Lambda$ dictates the exponential growth or decay rate of infinitesimal perturbations (gradients). For this fixed geometry, the maximum Lyapunov exponent is definitively non-positive: \(\Lambda_{\text{fixed}} = \log \sigma_{\max}(P_\mu) \le 0\) Key Insight: The informative subspace (orthogonal to the trivial constant mode) contracts at a rate governed by the spectral gap (second largest eigenvalue of $P_\mu$): $\Lambda_{\text{fixed}}^{(\perp)} = \log(1 - \tau \lambda_2)$. While a cleverly fixed geometry might slow this contraction, the system remains fundamentally diffusion-dominated. Long-range signals will inevitably die.
To break the diffusion limit, we allow the conformal factor to be recomputed at every layer using a 1-layer MLP parameterized by a shared weight vector $w^{(\ell)}$ and bias $b$: \(\mu_i^{(\ell)} = \sigma \left( w^{(\ell)\top} h_i^{(\ell)} + b \right).\) Now, the geometry and the node features co-evolve in a feedback loop: \(h^{(\ell)} \xrightarrow{\text{MLP}} \mu^{(\ell)} \xrightarrow{\text{Topology}} \mathcal{L}_{\mu^{(\ell)}} \xrightarrow{\text{Diffusion}} h^{(\ell+1)}.\) The propagation operator is now explicitly state-dependent, transforming the network into a nonlinear dynamical system: \(h^{(\ell+1)} = P(h^{(\ell)}) h^{(\ell)} \quad \text{where} \quad P(h) = I - \tau \mathcal{L}_{\mu(h)}.\)
To understand signal propagation in this nonlinear system, we must derive the exact layer-wise Jacobian $J_\ell = \frac{\partial h^{(\ell+1)}}{\partial h^{(\ell)}}$.
Let the symmetric transport weight and the local feature discrepancy be: \(W_{ij} = \frac{\mu_i + \mu_j}{2}, \qquad m_{ij} = (h_j - h_i)\Theta.\) The update rule for a single node $i$ is: \(h_i^{(\ell+1)} = h_i^{(\ell)} + \sum_{j \sim i} W_{ij} m_{ij}.\) We differentiate this update with respect to the input features. Note that differentiating the conformal factor yields the local gate vector g: \(g_j = \nabla_{h_j} \mu_j = \sigma'\left( z_j \right) w,\) where $z_j = w^\top h_j + b$ stands for the pre-activation.
Applying the product rule to $W_{ij} m_{ij}$, where $\mu_j$ depends on $h_j$ and $m_{ij}$ depends linearly on $h_j$: \(J_{ij} = \frac{\partial h_i^{(\ell+1)}}{\partial h_j^\ell} = W_{ij}\Theta + \frac{1}{2} m_{ij} g_j^\top.\)
For the self-derivative, $h_i$ appears in the residual connection, inside $\mu_i$ (affecting $W_{ij}$), and negatively inside $m_{ij}$: \(J_{ii} = \frac{\partial h_i^{(\ell+1)}}{\partial h_i} = I - d_i^W \Theta + \frac{1}{2} M_i g_i^\top,\) where $d_i^W = \sum_{j \sim i} W_{ij}$ is the weighted degree, and $M_i = \sum_{j \sim i} m_{ij}$ is the aggregate local discrepancy (the negative ordinary-Laplacian residual).
[!WARNING] Check sign Check sign between diffusion and geometric terms
For a single-step propagation $f(h) = \big(I - \tau L_{\mu(h)}\big) h$, the Jacobian $J(h) = \frac{\partial f}{\partial h}$ expands to: \(J(h) = \underbrace{\big(I - \tau L_{\mu(h)}\big)}_{D(h)} - \tau \underbrace{\sum_{i \sim j} \frac{\partial \mu_{ij}(h)}{\partial h} \otimes \nabla_{\mathcal{G}, ij} h}_{R(h)}\) This makes it immediately obvious that $\nabla_h \mu(h) \neq 0$ (the $g$ term above) creates a rank-one (or low-rank) directional feedback $R(h)$ aligned with the graph gradient of the signal $\nabla_{\mathcal{G}} h$ (the $m$ term above).
Stacking these blocks reveals a profound structural decomposition of the global layer Jacobian: \(J_\ell = D_\ell + R_\ell\) | | | | |—|—|—| |Component|Definition|Physical Meaning| |Diffusion $(D_\ell)$|$I - (L_W^{(\ell)} \otimes \Theta_\ell)$|The standard, fixed-geometry diffusion Jacobian. Strictly contractive.| |Feedback $(R_\ell)$|$\left( \frac{1}{2} s_{pq}^{(\ell)} g_q^{(\ell)\top} \right)_{pq}$|The dynamic geometric feedback mechanism. Vanishes if geometry is frozen.|
The term $R_\ell$ is entirely responsible for altering the Lyapunov spectrum. Because $g_q^{(\ell)} = \sigma’(z_q^{(\ell)}) w^{(\ell)}$, every non-zero block in column $q$ of $R_\ell$ shares the exact same right-factor: $w^{(\ell)}$.
This means the entire layer possesses a globally synchronized sensitivity direction: the adaptive geometry creates a structured low-rank feedback mechanism driven by the conformal direction $w^\ell$. This structured channel fundamentally cannot exist in a fixed-geometry model as we shall prove later.
To observe the asymptotic behavior of deep networks, we analyze the end-to-end multi-hop Jacobian across $K$ layers: \(\mathcal{J}_K = J_{K-1} J_{K-2} \cdots J_0\) The maximum Lyapunov exponent is formally defined as: \(\Lambda = \limsup_{K \to \infty} \frac{1}{K} \log \sigma_{\max}(\mathcal{J}_K)\) For a fixed geometry, $\Lambda = \Lambda_{\text{fixed}} \le 0$. For our adaptive geometry, we show that the exponent splits into the baseline diffusion plus a dynamic shift: $\Lambda = \Lambda_{\text{diff}} + \Delta \Lambda$.
[!NOTE] Jacobian cocycle for non-autonomous dynamics A complete treatment is given in [[A primer on non-autonomous dynamical systems with cocycles]].
The main object is now to study dynamics via the Jacobian cocycle. There are plenty of interesting results (Osseledets theorem etc …) and it is interesting to understand Lyapunov theory. Let’s start with the autonomous case to set things up, i.e we have the same dynamics $f$ at every time-step (layer). The Jacobian matrix $J_f(x)$ of a smooth function at $x \in \mathbb{R}^d$ satisfies the chain rule: \(J_{g \circ f}(x) = J_g(f(x)) J_f(x).\) This can in particular be applied to iterates $f^{(n)} = f \circ \cdots \circ f$: \(J_{f^{(2)}}(x) = J_f(f(x)) J_f(x), \, J_{f^{(3)}}(x) = J_{f} (f^{(2)}(x)) J_f (f(x)) J_f(x).\) By recursion, this leads to: \(\mathcal{J}_n (x)=J_{f^{(n)}}(x) = J_f(f^{(n-1)}(x)) J_f(f^{(n-2)}(x)) \cdots J_f(f(x)) J_f(x).\) We also see that $\mathcal{J}_n$ satisfies the following rule: \(J_{f^{(m+n)}}(x) = J_{f^{(m)}}(f^{(n)}(x)) J_{f^{(n)}}(x).\) This rule can be represented in an abstract way by introducing the matrix-valued function: $T: \mathbb{Z} \times \mathbb{R}^d \to \textrm{GL}(\mathbb{R},d)$. We observe that: \(T(m+n,x) = T(n,f^{(m)}(x)) \circ T(m,x),\) and this is the property defining a 1-cocycle. A general exponent used to study the trajectories of this discrete system can be defined as follows: \(\lambda_\textrm{max} = \limsup_{n \to \infty} \frac{1}{n} \log \| T(n,x) \| .\) Oseledets’ Multiplicative Ergodic Theorem guarantees that the following Lyapunov exponent exists almost everywhere: \(\lambda_k = \lim_{n \to \infty} \frac{1}{n} \sigma_k\bigl( (\mathcal{J}_n)\bigr),\) and this is the quantity we seek to control in this proof.
Assume the network discovers an aligned channel $v_\ell$ (parallel to $w^{(\ell)}$) since $R_\ell$ is essentially a projection onto $w^\ell$) such that it acts as an invariant direction for both operators independently: \(D_\ell v_\ell = d_\ell v_{\ell+1} \qquad \text{and} \qquad R_\ell v_\ell = r_\ell v_{\ell+1}\) where $d_\ell > 0$ (the diffusion survival rate) and $r_\ell \ge 0$ (the geometric feedback magnitude).
[!WARNING] Hypothesis Consider this proof mostly for pedagogical purposes. The perfect alignment hypothesis of diffusion and geometric feedback is very strong but it is relaxed below (see points 14 and 15) as a much weaker correlation condition.
Applying the full Jacobian to this vector: \(J_\ell v_\ell = (D_\ell + R_\ell) v_\ell = (d_\ell + r_\ell) v_{\ell+1}.\) Tracking this vector through $K$ layers yields the cumulative sensitivity: \(\Vert{}\mathcal{J}_K v_0\Vert{} = \prod_{\ell=0}^{K-1} (d_\ell + r_\ell).\) Substituting this into the Lyapunov definition provides a strict lower bound on the system’s maximum exponent: \(\Lambda \ge \liminf_{K \to \infty} \frac{1}{K} \sum_{\ell=0}^{K-1} \log(d_\ell + r_\ell).\) Subtracting the baseline diffusion exponent $\Lambda_{\text{diff}} = \liminf \frac{1}{K} \sum \log(d_\ell)$, we isolate the shift: \(\Delta \Lambda = \liminf_{K \to \infty} \frac{1}{K} \sum_{\ell=0}^{K-1} \log \left( 1 + \frac{r_\ell}{d_\ell} \right).\) Because the feedback term $r_\ell$ is non-negative, $\Delta \Lambda \ge 0$, the adaptive feedback mechanism acts strictly to counteract contraction.
Assuming the feedback is stable and smaller than the diffusion component $(r_\ell \ll d_\ell)$, we apply the Taylor expansion $\log(1+x) \approx x$ to yield the weak-feedback approximation: \(\Delta \Lambda \approx \frac{1}{K} \sum_{\ell=0}^{K-1} \frac{r_\ell}{d_\ell}.\)
We can explicitly connect this shift back to the network’s forward activations. Define the local alignment scalar be the dot product of the gate vector and the feature discrepancy (graph gradient of features): \(\kappa_{pq} = g_q^\top s_{pq}.\) Since $r_\ell$ is directly proportional to this alignment scalar $\kappa_\ell$ across the layer, the Lyapunov shift simplifies to: \(\Delta \Lambda \approx \frac{1}{2} \left\langle \frac{\kappa_\ell}{d_\ell} \right\rangle,\) where $\langle \cdot \rangle$ is the mean value operator.
This final equation reveals that the alignment field $\kappa$ acts as the infinitesimal source term responsible for shifting the Lyapunov exponents. Crucially, the objective of deep graph learning is not to push $\Lambda > 0$. A positive Lyapunov exponent dictates chaotic dynamic signals amplify exponentially, leading to gradient explosion and numerical collapse.
Instead, a trainable deep network requires the Edge of Chaos $(\Lambda \approx 0)$. In this regime:
By dynamically updating the geometry, the network utilizes the gradient descent of $w^{(\ell)}$ to tune $\kappa_\ell$. When diffusion $(d_\ell)$ threatens to crush a task-relevant signal, the network increases local alignment $(\kappa_\ell)$, injecting just enough $\Delta \Lambda$ to push the specific conformal channel back to $\Lambda \approx 0$.
The previous derivation considered the residual diffusion update
$$
h_i^{(\ell+1)}
=
h_i^{(\ell)}
Let $\phi:\mathbb R^k\rightarrow\mathbb R^k$ be a componentwise activation function (ReLU, GELU, SiLU, etc.). We define the pre-activation
$$
a_i^{(\ell)}
=
h_i^{(\ell)}
The adaptive conformal geometry remains unchanged:
$$
\mu_i^{(\ell)}
=
\sigma!\left(
w^{(\ell)\top}h_i^{(\ell)}+b^{(\ell)}
\right),
\quad
W_{ij}^{(\ell)}
=
\frac{
\mu_i^{(\ell)}
The geometry and the node representations therefore continue to co-evolve through the feedback loop \(h^{(\ell)} \rightarrow \mu^{(\ell)} \rightarrow \mathcal L_{\mu^{(\ell)}} \rightarrow a^{(\ell)} \rightarrow h^{(\ell+1)}.\)
Let \(J_\ell = \frac{\partial h^{(\ell+1)}} {\partial h^{(\ell)}}.\)
Applying the chain rule gives \(J_\ell = \frac{\partial h^{(\ell+1)}} {\partial a^{(\ell)}} \frac{\partial a^{(\ell)}} {\partial h^{(\ell)}}.\)
The second factor is exactly the Jacobian derived previously: \(\frac{\partial a^{(\ell)}} {\partial h^{(\ell)}} = D_\ell + R_\ell.\)
The first factor is the activation Jacobian \(\Sigma_\ell = \operatorname{diag} \!\left( \phi'(a_1^{(\ell)}), \dots, \phi'(a_N^{(\ell)}) \right).\)
Hence \(J_\ell = \Sigma_\ell \left( D_\ell + R_\ell \right).\)
This becomes the fundamental Jacobian decomposition of the adaptive conformal GNN.
The Jacobian now naturally decomposes into three distinct mechanisms:
$$
J_\ell
=
\underbrace{\Sigma_\ell D_\ell}_{\text{activated diffusion}}
$D_\ell$ is the standard graph propagation operator. For normalized-Laplacian propagation its spectrum lies inside the unit disk, producing diffusion-induced contraction.
$\Sigma_\ell$ acts as a coordinate-wise gating operator. For ReLU, $\phi(x)=\max(x,0)$, so \(\phi'(x) = \begin{cases} 1,&x>0,\\ 0,&x<0. \end{cases}\) Therefore $\Sigma_\ell$ is a binary diagonal mask. Active coordinates propagate normally while inactive coordinates are completely suppressed. The activation itself never creates expansion. Indeed, $|\Sigma_\ell|_2\le 1$. Thus activations contribute additional contraction or coordinate annihilation.
The feedback remains \(R_\ell = \left( \frac12 s_{pq}^{(\ell)} g_q^{(\ell)\top} \right).\) Because \(g_q^{(\ell)} = \sigma'(z_q^{(\ell)}) w^{(\ell)},\) every block-column shares the same right factor $w^{(\ell)}$. Consequently the adaptive geometry still generates a globally synchronized sensitivity channel. The activation gates this channel through $\Sigma_\ell$, but does not create it.
The end-to-end Jacobian becomes \(\mathcal J_K = J_{K-1} J_{K-2} \cdots J_0 = \prod_{\ell=0}^{K-1} \Sigma_\ell (D_\ell+R_\ell).\) The maximal Lyapunov exponent is \(\Lambda = \limsup_{K\to\infty} \frac1K \log \sigma_{\max} (\mathcal J_K).\) For a fixed geometry $(R_\ell=0)$: \(J_\ell^{\mathrm{fixed}} = \Sigma_\ell D_\ell.\) Both factors satisfy \(\|\Sigma_\ell\|_2\le 1, \qquad \|D_\ell\|_2\le1.\) Hence \(\Lambda_{\mathrm{fixed}} \le 0.\) The combination of graph diffusion and activation gating remains fundamentally contractive.
Assume an aligned conformal channel $v_\ell$ (again this hypothesis is relaxed in 14 and 15 below) satisfying
\(D_\ell v_\ell
=
d_\ell v_{\ell+1},
\quad
R_\ell v_\ell
=
r_\ell v_{\ell+1},
\textrm{ and }
\Sigma_\ell v_{\ell+1}
=
s_\ell v_{\ell+1},
\qquad
0\le s_\ell\le 1.\)
Then
\(J_\ell v_\ell
=
s_\ell
(d_\ell+r_\ell)
v_{\ell+1}.\)
After $K$ layers,
\(\|
\mathcal J_K v_0
\|
=
\prod_{\ell=0}^{K-1}
s_\ell(d_\ell+r_\ell).\)
Therefore
$$
\Lambda_{\mathrm{channel}}
=
\left\langle
\log s_\ell
\right\rangle
Similarly, for fixed geometry,
$$
\Lambda_{\mathrm{fixed}}
=
\left\langle
\log s_\ell
\right\rangle
Theorem (Lyapunov Shift with Pointwise Activations)
Assume:
\(J_\ell
=
\Sigma_\ell(D_\ell+R_\ell),\)
with $|\Sigma_\ell|2\le1$, and suppose an aligned conformal channel exists such that
\(D_\ell v_\ell=d_\ell v_{\ell+1},
\qquad
R_\ell v_\ell=r_\ell v_{\ell+1},
\qquad
r_\ell\ge0.\)
Then the Lyapunov exponent along that channel satisfies
$$
\Lambda{\mathrm{adaptive}}
=
\Lambda_{\mathrm{fixed}}
The key insight becomes stronger than before. A fixed-geometry GNN contains two contraction mechanisms: graph diffusion $D_\ell$ and activation gating $\Sigma_\ell$. Both push Lyapunov exponents downward. On the other hand, adaptive conformal geometry introduces a third term: $R_\ell$. This term is the only component capable of shifting Lyapunov exponents upward. Pointwise nonlinearities may attenuate or mask this effect, but they do not generate it. The upward Lyapunov shift remains entirely attributable to the evolving geometry itself. Consequently, the adaptive conformal mechanism should be viewed as a state-dependent geometric feedback controller whose role is to compensate the combined contraction generated by diffusion and activation gating, moving selected propagation channels toward the critical regime $\Lambda \approx 0$, where long-range sensitivity can persist without destabilizing the network.
Or Conformal Channel Normalization and Self-Regulation of the Lyapunov Shift
The adaptive conformal mechanism introduces a new feedback term in the layer Jacobian $J_\ell = D_\ell + R_\ell$, whose effect on long-range sensitivity is quantified by the Lyapunov shift \(\Delta\Lambda = \left\langle \log\!\left( 1+\frac{r_\ell}{d_\ell} \right) \right\rangle.\) In the weak-feedback regime, \(\Delta\Lambda \approx \left\langle \frac{r_\ell}{d_\ell} \right\rangle ,\) with $r_\ell$ controlled by the alignment between the geometric feedback direction and the local feature discrepancies.
The challenge is that the conformal generator \(\mu_i^{(\ell)} = \sigma\!\left( w_\ell^\top h_i^{(\ell)} + b_\ell \right)\)
may enter two undesirable regimes:
| Geometry freezing: $ | w^\top h_i+b | \gg 1$ causes sigmoid saturation and therefore $\sigma’(w^\top h_i+b)\approx 0$, implying $R_\ell \approx 0$. The adaptive geometry degenerates into a fixed geometry. |
The desired regime is neither strong contraction nor instability, but rather $\Lambda \approx 0$, namely the edge-of-chaos regime where sensitivity persists across depth while the dynamics remain bounded.
The adaptive geometry only depends on the scalar projection \(a_i^{(\ell)} = w_\ell^\top h_i^{(\ell)} .\) This quantity is the true latent variable driving the geometry: \(\mu_i^{(\ell)} = \sigma(a_i^{(\ell)}+b_\ell).\) It also controls the geometric Jacobian through \(g_i^{(\ell)} = \sigma'(a_i^{(\ell)}+b_\ell)\,w_\ell.\) Consequently:
We therefore refer to $a_i=w^\top h_i$ as the conformal channel.
To prevent collapse or explosion of the geometric feedback, we normalize the scalar field $a_i$ before generating the conformal factor.
Compute $a_i=w^\top h_i$ . Let \(\bar a = \frac1N\sum_{i=1}^{N} a_i ,\) and \(s_a^2 = \frac1N \sum_{i=1}^{N} (a_i-\bar a)^2 .\) Define the normalized conformal coordinate \(\widetilde a_i = \frac{a_i-\bar a} {\sqrt{s_a^2+\varepsilon}} .\) The geometry is then generated as \(\mu_i = \sigma(\gamma \widetilde a_i+\beta),\) where $\gamma$ and $\beta$ are trainable scalar parameters.
This operation is analogous to Batch Normalization, but applied only to the scalar field that generates the conformal geometry. The normalization enforces approximately $\operatorname{Var}(\widetilde a) \approx 1$.
As a result:
Instead of learning an unconstrained geometry, the model learns a geometry whose driving coordinate remains near a critical operating point.
The geometric feedback term can be written as \(R_\ell = \left( \frac12 s_{pq}^{(\ell)} g_q^{(\ell)\top} \right)_{pq},\)
with \(g_q^{(\ell)} = \sigma'(\gamma\widetilde a_q+\beta)\,w_\ell .\)
Because $\widetilde a_q$ is normalized, the derivative \(\sigma'(\gamma\widetilde a_q+\beta)\)
remains bounded away from pathological saturation and $R_\ell$ remains active even in very deep networks. The adaptive geometry therefore preserves its ability to modify the Jacobian spectrum over the entire training trajectory.
Recall the approximate shift formula \(\Delta\Lambda \approx \frac12 \left\langle \frac{\kappa_\ell}{d_\ell} \right\rangle ,\)
where \(\kappa_{ij}^{(\ell)} = g_j^{(\ell)\top} s_{ij}^{(\ell)}\)
is the alignment field. Without normalization: $\kappa_{ij} \rightarrow 0$ when the sigmoid saturates, causing $\Delta\Lambda \rightarrow 0$: The system then reverts to diffusion-dominated contraction.
Conversely, excessively large conformal gains may produce $\kappa_{ij} \gg d_\ell$, leading to large positive Lyapunov shifts and unstable dynamics. Conformal channel normalization prevents both situations by maintaining the feedback term in an intermediate operating range. Normalization acts as a homeostatic regulator of the Lyapunov shift.
The adaptive conformal mechanism can now be viewed as a feedback controller acting on the dominant Lyapunov exponent:
$$
\Lambda
=
\Lambda_{\mathrm{diff}}
\Delta\Lambda .
$$
The role of conformal channel normalization is not to maximize $\Delta\Lambda$, but to keep it within a range where $\Lambda \approx 0$. This yields the desirable critical regime:
In this perspective, the normalized conformal channel becomes a self-regulating geometric mechanism that continuously adjusts the feedback strength while preventing both geometric collapse and geometric instability. This provides a natural architectural route toward maintaining adaptive graph diffusion near the edge of chaos.
In adaptive conformal graph neural networks, the geometry is no longer a static object. Instead, the graph metric co-evolves with the node representations throughout the forward pass. This transforms graph propagation from a fixed diffusion process into a nonlinear dynamical system. The asymptotic propagation of infinitesimal perturbations is governed by the dominant Lyapunov exponent $\Lambda$. The analysis developed in the previous sections yields a natural decomposition
$$
\Lambda
=
\Lambda_{\mathrm{diff}}
A fundamental observation follows immediately:
| If $\Delta\Lambda \ll | \Lambda_{\mathrm{diff}} | $, diffusion dominates, resulting in vanishing gradients, oversmoothing, and oversquashing. |
| If $\Delta\Lambda > | \Lambda_{\mathrm{diff}} | $, the dominant Lyapunov exponent becomes positive and the system enters an unstable regime characterized by exponentially growing perturbations. |
At this critical point:
Consequently, adaptive conformal geometry should not be viewed merely as a learnable graph metric. It naturally induces a feedback controller whose role is to compensate diffusion-induced contraction and keep the propagation dynamics close to criticality.
The adaptive geometry is generated from the scalar projection \(a_i^{(\ell)} = w_\ell^\top h_i^{(\ell)}.\) We call \(a^{(\ell)} = (a_1^{(\ell)},\ldots,a_N^{(\ell)})\) the conformal channel. This scalar field (graph signal) is the sole quantity through which the node representations influence the graph geometry. To avoid contamination from the spatially constant mode that don’t contract under diffusion, we first project onto the zero-mean graph subspace $\mathbf 1_N^\perp$.
Let
\(\bar a^{(\ell)}
=
\frac1N
\sum_{i=1}^{N}
a_i^{(\ell)}.\)
Define
\(\widetilde a_i^{(\ell)}
=
\frac{
a_i^{(\ell)}-\bar a^{(\ell)}
}
{
\operatorname{std}(a^{(\ell)})+\varepsilon
}.\)
The conformal factor is then generated through
$$
\mu_i^{(\ell)}
=
\sigma
!\left(
\gamma_\ell
\widetilde a_i^{(\ell)}
where $\gamma_\ell$ is a trainable gain parameter and $\beta_\ell$ is a trainable bias.
The normalized conformal channel plays a role analogous to Batch Normalization, but only for the geometry-generating coordinate. The gain parameter $\gamma_\ell$ directly controls the strength of geometric feedback:
Consequently, $\gamma_\ell$acts as a natural control parameter for the Lyapunov spectrum.
The layer output is \(h^{(\ell+1)} = F_\ell(h^{(\ell)}).\) The Jacobian splits as $J_\ell = D_\ell + R_\ell$. Here \(D_\ell = I-\tau L_{\mu^{(\ell)}}\) is the diffusion operator, while \(R_\ell = \frac{\partial D_\ell}{\partial h^{(\ell)}}h^{(\ell)}\) is the geometry-feedback term. This component exists only because the metric itself depends on the state and it is exactly responsible for the Lyapunov shift derived previously.
To control the dynamics, the network requires an estimate of its current sensitivity growth rate. Rather than explicitly forming the huge Jacobian product $J_{0:K-1} = J_{K-1}\cdots J_0$, we estimate the dominant Lyapunov exponent through a layer-wise power iteration using Jacobian-vector products (JVPs).
Every graph Laplacian satisfies $L_\mu \mathbf 1N=0$. Therefore probe directions containing a spatially constant component do not experience meaningful diffusion contraction. To remove this degeneracy, every probe tensor $q\ell \in \mathbb R^{N\times d}$ is constrained to $\mathbf 1N^\perp\otimes\mathbb R^d$. That is,
$$
\sum{i=1}^{N}
(q_\ell)_{i,:}
=
Choose $q_0 \in \mathbf 1_N^\perp\otimes\mathbb R^d$ such that $|q_0|_F=1$.
For each layer $\ell=0,\ldots,K-1$, compute $\hat q_{\ell+1} = J_\ell q_\ell$ using a single Jacobian-vector product.
Project back to $\mathbf 1N^\perp$ and compute $s\ell = |\hat q_{\ell+1}|F$. Then set $q{\ell+1} = \frac{\hat q_{\ell+1}} {s_\ell}$.
The resulting empirical exponent is \(\widehat\Lambda_K = \frac1K \sum_{\ell=0}^{K-1} \log s_\ell.\) This quantity is the finite-depth analogue of the dominant Lyapunov exponent.
Importantly:
| The complexity remains $O( | E | d)$. |
The estimate $\widehat\Lambda_K$ can now be used to actively regulate the geometry. The target state is $\Lambda_{\mathrm{target}}=0$.
Add the penalty $\mathcal L_{\mathrm{crit}} = \eta \, \widehat\Lambda_K^2$. The optimizer is encouraged to maintain $\widehat\Lambda_K \approx 0$. To avoid expensive higher-order differentiation through the power iteration, the probe vectors $q_\ell$ are detached during accumulation.
Instead of adding a loss term, one can directly adapt the conformal gain.
Update \(\gamma_{\ell+1} = \gamma_\ell \exp\!\big( -\alpha\widehat\Lambda_\ell \big),\) for some $\alpha>0$. This creates an autonomous feedback loop: If $\widehat\Lambda_\ell<0$, the dynamics are overly contractive and $\gamma$ increases. If $\widehat\Lambda_\ell>0$, the dynamics are expansive and $\gamma$ decreases. Thus the geometry automatically compensates contraction while suppressing instability.
The gain controller admits a particularly simple characterization.
Proposition (Critical Fixed Point)
Assume \(\gamma_{\ell+1} = \gamma_\ell \exp(-\alpha\widehat\Lambda_\ell)\) with $\alpha>0$. Any equilibrium of the gain dynamics satisfies $\widehat\Lambda_\ell=0$.
Proof
At equilibrium, $\gamma_{\ell+1} = \gamma_\ell$. Therefore \(\exp(-\alpha\widehat\Lambda_\ell)=1.\) Since $\alpha>0$, $\widehat\Lambda_\ell=0$. $\square$
The edge of chaos is not imposed externally. It emerges as the fixed point of the geometric feedback dynamics. The adaptive geometry continuously monitors the propagation stability and injects precisely the amount of conformal feedback needed to offset diffusion-induced contraction.
[!WARNING] Fixed Point This proposition only guarantees the existence of a fixed point. It does not guarantee that it is reached.
The complete architecture forms a closed-loop control system: \(q_\ell \;\xrightarrow{J_\ell}\; q_{\ell+1} \;\xrightarrow{}\; \widehat\Lambda \;\xrightarrow{}\; \gamma \;\xrightarrow{}\; R \;\xrightarrow{}\; J.\)
The loop can be summarized as \(\Lambda_{\mathrm{diff}} \longrightarrow \gamma \longrightarrow R \longrightarrow \Delta\Lambda \longrightarrow \Lambda.\)
Whenever diffusion becomes excessively contractive, the geometry increases the feedback gain. Whenever the network approaches instability, the gain is suppressed. The equilibrium condition is $\Lambda_{\mathrm{diff}} + \Delta\Lambda = 0$.
The adaptive conformal geometry induces a natural homeostatic mechanism for regulating deep graph propagation.
Viewed in this way, adaptive conformal geometry is not merely a learnable metric. It is a self-regulating dynamical mechanism that continuously tunes long-range information propagation toward criticality.
The previous aligned-channel hypothesis assumed that the diffusion action and the conformal feedback action remain perfectly aligned. While analytically convenient, this assumption is quite strong and hard to satisfy in practice. The following result replaces exact alignment by a weaker and empirically measurable correlation condition.
Let $J_\ell = D_\ell + R_\ell$ denote the Jacobian of the adaptive conformal layer at depth $\ell$, where $D_\ell = I-\tau L_{\mu^{(\ell)}}$ is the diffusion component and $R_\ell = \frac{\partial D_\ell}{\partial h^{(\ell)}}h^{(\ell)}$ is the conformal feedback component.
Let $\mathcal J_K = J_{K-1}\cdots J_0$ be the depth-$K$ Jacobian cocycle, and define the top Lyapunov exponent \(\Lambda = \limsup_{K\to\infty} \frac1K \log \sigma_{\max}(\mathcal J_K).\) Consider an arbitrary sequence of nonzero probe vectors $v_\ell \in \mathbb R^{Nd}$ and define \(u_\ell := D_\ell v_\ell, \qquad r_\ell := R_\ell v_\ell.\) Introduce the diffusion and feedback amplitudes \(d_\ell := \|u_\ell\|, \qquad c_\ell := \|r_\ell\|,\) together with their cosine correlation: \(\rho_\ell = \frac{ \langle u_\ell,r_\ell\rangle }{ \|u_\ell\|\,\|r_\ell\| }, \qquad \rho_\ell\in[-1,1].\)
Assume there exists a constant $\rho_0>0$ such that $\rho_\ell \ge \rho_0$ for all sufficiently large $\ell$, or more generally in Cesàro average, \(\liminf_{K\to\infty} \frac1K \sum_{\ell=0}^{K-1} \rho_\ell \ge \rho_0.\) This assumption requires only that diffusion and conformal feedback remain positively correlated on average; no exact alignment is required.
Under the above assumptions, \(\Lambda \ge \Lambda_{\mathrm{diff}} + \Delta\Lambda_{\rho},\) where \(\Lambda_{\mathrm{diff}} = \liminf_{K\to\infty} \frac1K \sum_{\ell=0}^{K-1} \log d_\ell\) is the diffusion Lyapunov exponent and \(\Delta\Lambda_{\rho} = \liminf_{K\to\infty} \frac1{2K} \sum_{\ell=0}^{K-1} \log \left( 1+ \left(\frac{c_\ell}{d_\ell}\right)^2 + 2\rho_0 \frac{c_\ell}{d_\ell} \right).\) Consequently, $\Delta\Lambda_{\rho}\ge0$ whenever $\rho_0>0$.
Since $J_\ell v_\ell = u_\ell+r_\ell$ $|J_\ell v_\ell|^2 = |u_\ell+r_\ell|^2$, we have \(\|J_\ell v_\ell\|^2 = \|u_\ell\|^2 + \|r_\ell\|^2 + 2\langle u_\ell,r_\ell\rangle.\) Using the definitions of $d_\ell$, $c_\ell$, and $\rho_\ell$, \(\|J_\ell v_\ell\|^2 = d_\ell^2 + c_\ell^2 + 2d_\ell c_\ell \rho_\ell.\) By the positive-correlation assumption, $\rho_\ell \ge \rho_0$, hence \(\|J_\ell v_\ell\| \ge d_\ell \sqrt{ 1+ \left(\frac{c_\ell}{d_\ell}\right)^2 + 2\rho_0\frac{c_\ell}{d_\ell} }.\) Iterating along the Jacobian cocycle yields \(\|\mathcal J_K v_0\| \ge \prod_{\ell=0}^{K-1} d_\ell \sqrt{ 1+ \left(\frac{c_\ell}{d_\ell}\right)^2 + 2\rho_0\frac{c_\ell}{d_\ell} }.\) Taking the logarithm, dividing by $K$, and passing to the limit gives \(\Lambda \ge \Lambda_{\mathrm{diff}} + \Delta\Lambda_{\rho}.\) $\square$
If diffusion and feedback are perfectly aligned, $\rho_0=1$, then \(1+ \left(\frac{c}{d}\right)^2 + 2\frac{c}{d} = \left( 1+\frac{c}{d} \right)^2.\) Therefore \(\Delta\Lambda_{\rho} = \liminf_{K\to\infty} \frac1K \sum_{\ell=0}^{K-1} \log \left( 1+\frac{c_\ell}{d_\ell} \right),\) which is exactly the previously derived aligned-channel theorem.
Suppose $c_\ell \ll d_\ell$, i.e the strength of geometric feedback is much smaller than the strength of diffusion. Using \(\log(1+x) = x+O(x^2),\) we obtain \(\Delta\Lambda_{\rho} = \rho_0 \left\langle \frac{c_\ell}{d_\ell} \right\rangle + O\!\left( \left\langle \left(\frac{c_\ell}{d_\ell}\right)^2 \right\rangle \right).\) Where $\langle \cdot \rangle$ is the Césaro expectation. To first order, \(\Delta\Lambda_{\rho} \approx \rho_0 \left\langle \frac{c_\ell}{d_\ell} \right\rangle.\) This reveals that $\rho_0$ acts as an efficiency factor quantifying how effectively geometric feedback compensates diffusion-induced contraction.
This theorem replaces exact channel alignment by a far weaker and more realistic condition: the adaptive geometry need not create a propagation mode identical to the diffusion mode, it is sufficient that the geometric feedback remains positively correlated with diffusion on average.
The resulting Lyapunov correction $\Delta\Lambda_{\rho}$ quantifies how much of the diffusion-induced contraction is compensated by the evolving geometry.
Interestingly, \(\rho_\ell = \frac{ \langle D_\ell q_\ell,\; R_\ell q_\ell\rangle }{ \|D_\ell q_\ell\| \, \|R_\ell q_\ell\| }\) can be measured directly using the same JVP probe vectors employed for the online Lyapunov estimator, making the theorem fully testable experimentally. In particular, the theory predicts that successful training should drive the system toward positive average correlation and therefore toward a reduced contraction rate, with the critical regime corresponding to $\Lambda_{\mathrm{diff}} + \Delta\Lambda_{\rho} \approx 0$.
The previous theorem was stated for the linear case (no nodewise non-linearity), its extension to non-linear layers is provided below.
Recall that in the full architecture with nodewise non-linearity $\phi$, the layerwise Jacobian can be written as $J_\ell = \Sigma_\ell (D_\ell+R_\ell)$, where $\Sigma_\ell = \operatorname{diag}\big(\phi’(a^{(\ell)})\big)$. For ReLU, $\Sigma_\ell$ is a binary diagonal mask.
Now the vectors actually entering propagation are $u_\ell = \Sigma_\ell D_\ell v_\ell$, and $r_\ell = \Sigma_\ell R_\ell v_\ell$. Thus \(J_\ell v_\ell = u_\ell+ r_\ell.\) The natural correlation coefficient becomes \(\rho_\ell = \frac{ \langle \Sigma_\ell D_\ell v_\ell, \Sigma_\ell R_\ell v_\ell \rangle } { \| \Sigma_\ell D_\ell v_\ell \| \, \| \Sigma_\ell R_\ell v_\ell \| }.\) This is simply the previous correlation measured after activation gating. We formalize this as a new version of the Persistent Correlation Assumption.
There exists $\rho_0>0$ such that \(\rho_\ell = \frac{ \langle \Sigma_\ell D_\ell v_\ell, \Sigma_\ell R_\ell v_\ell \rangle } { \| \Sigma_\ell D_\ell v_\ell \| \, \| \Sigma_\ell R_\ell v_\ell \| } \ge \rho_0\) on average across depth.
Conceptually this is much nicer than the original aligned-channel hypothesis: it says only that the geometry keeps helping the directions that actually survive the activation function. That is both weaker and more realistic since a feedback direction that is immediately killed by ReLU cannot contribute to long-range sensitivity.
Under the Persistent Correlation Assumption,
\(\Lambda \ge \Lambda_{\mathrm{act-diff}} + \Delta\Lambda_{\rho},\) where \(\Lambda_{\mathrm{diff}} = \liminf_{K\to\infty} \frac1K \sum_{\ell=0}^{K-1} \log d_\ell\) is the diffusion Lyapunov exponent and \(\Delta\Lambda_{\rho} = \liminf_{K\to\infty} \frac1{2K} \sum_{\ell=0}^{K-1} \log \left( 1+ \left(\frac{c_\ell}{d_\ell}\right)^2 + 2\rho_0 \frac{c_\ell}{d_\ell} \right).\) Consequently, $\Delta\Lambda_{\rho}\ge0$ whenever $\rho_0>0$.
Define
\(d_\ell = \|\Sigma_\ell D_\ell v_\ell\|, \quad c_\ell = \|\Sigma_\ell R_\ell v_\ell\|.\)
Then
\(\|J_\ell v_\ell\|^2 = d_\ell^2 + c_\ell^2 + 2 d_\ell c_\ell \rho_\ell.\)
By the positive-correlation assumption,
\(\|J_\ell v_\ell\| \ge d_\ell \sqrt{ 1+ \left(\frac{c_\ell}{d_\ell}\right)^2 + 2\rho_0\frac{c_\ell}{d_\ell} }.\)
Iterating along the Jacobian cocycle yields
\(\|\mathcal J_K v_0\| \ge \prod_{\ell=0}^{K-1} d_\ell \sqrt{ 1+ \left(\frac{c_\ell}{d_\ell}\right)^2 + 2\rho_0\frac{c_\ell}{d_\ell} }.\)
Taking the logarithm, dividing by $K$, and passing to the limit gives
\(\Lambda \ge \Lambda_{\mathrm{act+diff}} + \Delta\Lambda_{\rho},\)
with
\(\Lambda_{\mathrm{act+diff}} = \liminf \frac1K \sum_{\ell} \log d_\ell,\)
and
\(\Delta\Lambda_{\rho} = \frac12 \Bigg\langle \log \Big( 1+ ( c/ d)^2 + 2\rho ( c/ d) \Big) \Bigg\rangle.\)
$\square$