Abstract
Historically, sequence modeling forces a binary design fork: continuous flow or discrete steps. Continuous models (like Closed-Form Continuous-Time neural networks [CfCs] and Neural ODEs) treat state as smooth, uninterrupted drift—tracking long-term dynamical trajectories, but failing when faced with sudden discontinuities and phase jumps where gradients explode. Discrete architectures (such as Transformers) treat dynamical systems as frozen frames, absorbing shocks but losing continuous sub-step drift and generating runaway quantization drift over long horizons.
The Dual-Representation Siamese Thalamocortical World Engine (DR-S-TWE) resolves this dichotomy. DR-S-TWE processes non-stationary time series in representation superposition: executing parallel continuous and discrete encoders synchronized via a surprise-sensitive neuromodulatory gate inspired by thalamocortical reticular synchronization. On non-linear benchmarks, DR-S-TWE achieves +0.9890 correlation on the Lorenz chaotic attractor and +0.9553 correlation (52% MAE reduction) on non-stationary weather shocks, outperforming zero-shot foundation models.
1. The Continuous-Discrete Dilemma
Why does next-token prediction fail physical and financial dynamics? Real-world systems operate in continuous time—governed by differential equations, physical conservation invariants, and high-frequency market liquidity regimes. Autoregressive token guessing over frozen bins creates two severe failure modes:
- Continuous Solvers at Shocks: When a market crashes or a thermodynamic valve trips, continuous differential equations attempt to resolve infinite derivatives ($\frac{dx}{dt} \to \infty$), breaking numerical integration.
- Discrete Models at Drift: Quantized token models lose sub-step geometric drift, accumulating compound error that hallucinates physically impossible trajectories over extended rollouts.
2. Mathematical Formulation
2.1 Siamese Projection
For an incoming observation $x_t \in \mathbb{R}^d$, the continuous stream projects directly into continuous representation space:
Simultaneously, the discrete stream soft-quantizes each dimension across $B$ bin centers $C \in \mathbb{R}^B$ using a Radial Basis Function (RBF) softmax distribution:
2.2 Neuromodulatory Surprise Gating
A gating network calculates a dynamic routing coefficient $g_t$ based on the prior state and current observation. Under low surprise (predictable secular drift), $g_t \to 1.0$, routing compute through the continuous stream. Under high surprise (discontinuities, gap openings, sudden liquidity shocks), $g_t \to 0.0$, routing focus through the discrete path:
2.3 Closed-Form Continuous-Time Integration
The superposed embedding $e_t$ updates the hidden state $h_t$ via closed-form analytical decay, eliminating numerical instability without requiring iterative ODE solvers:
3. Empirical Benchmarks
| Domain / Benchmark | Model Architecture | Correlation | MAE |
|---|---|---|---|
| Lorenz Chaotic Attractor | DR-S-TWE (Task-Tuned) | +0.9890 | 0.0899 |
| Lorenz Chaotic Attractor | Amazon Chronos 8M (Zero-Shot) | +0.9847 | 0.0413 |
| Melbourne Weather Shock | DR-S-TWE (Task-Tuned) | +0.9553 | 0.2299 |
| Melbourne Weather Shock | Amazon Chronos 8M (Zero-Shot) | +0.7593 | 0.4820 |
4. Collaboration & Patent Notice
Patent Notice: The core dual-representation continuous-discrete architecture and neuromodulatory surprise routing mechanisms described herein are subject to U.S. Provisional Patent Application No. 64/108,066, filed July 9, 2026, assigned to Conner Kupferberg.