Mamba-3-Lite

From-scratch PyTorch reproduction of Mamba-3 at ~434M params — complex-valued SSD state spaces (N=64, complex64), a fully-connected MIMO head mixer, and zero causal convolution. Pure PyTorch, one sanctioned opt-in Triton kernel. Read it like a field notebook: a name, a wiring sketch, then the measurements.

STATE COMPRESSION
50% SMALLER N
N=64 complex64 packs two real sub-states into one complex state (vs Mamba-2 N=128)
MIMO MIXER
16×16 CROSS-HEAD
Fully-connected mixer replaces SISO constraint · zero sequence cost overhead
CAUSAL CONV
0 ELIMINATED
Zero causal_conv1d passes · pure chunked linear projection saves memory bandwidth
COMPLEX EIGEN
α+iω DECAY+ROT
Complex exponential captures both decay (α) and oscillation (ω) per state
FIG. A0 COMPLEX STATE-SPACE DYNAMICS N=64 · 32 CONJUGATE PAIRS
01 · params
~434M
02 · layers
28
03 · state
N=64 complex64
04 · hidden
1024
05 · ffn
2048 SwiGLU
06 · vocab
50 K
07 · ctx
2048 chunk 64
08 · tokens
8.0 B
§ MAMBA-3 CORE MECHANISMS · INTERACTIVE BENCHMARK LABS
01 · COMPLEX SSD Eigenvalue Spectrum Explorer

Visualize how complex eigenvalues λ = α + iω govern state dynamics. Decay rate α controls memory retention; rotation ω enables oscillatory patterns impossible in real SSMs.

Decay Range (α): [-3.0, -0.05]
Active Eigen Pairs: 32 conjugate pairs
Memory Retention (midpoint α half-life): ~0.4 steps
State Capacity vs Real SSM: 2.0× per parameter
02 · MIMO MIXER Cross-Head Information Flow

Simulate the 16×16 fully-connected mixer across SSM heads. Identity-initialized weights gradually learn cross-head correlations during training, replacing SISO isolation.

Mixer Dimensions: 1024 → 1024 (16×64)
Init Strategy: Pure identity (eye_) · init_std skipped
Active Cross-Head Links: 0 of 240 off-diagonal
03 · CHUNKWISE SSD Linear Scan vs Chunked Parity

Compare naive O(T) sequential scan against the chunkwise algorithm (Q=64). The chunkwise approach decomposes into parallelizable intra-chunk and inter-chunk passes.

NAIVE s_t = Ā·s_{t-1} + B_t·x_t, Ā=exp(softplus(dt)·A) (O(T) loop)
CHUNK L[l,s]=exp(A_cs[l]-A_cs[s])·1[l≥s]; Y_diag = einsum(C,B,L,X) (intra-chunk matmul)
STATE s_chunk = einsum(B,decay,X); cross-chunk via decay_chunk matrix (exp of chunk-decay differences)
Chunk Size (Q): 64 tokens
Sequence Length: 2,048 (32 chunks)
Numerical Match (vs naive): < 1e-5 max |Δ|
FIG. A1 FULL TRAINING STEP PIPELINE FORWARD · AUTOGRAD BACKWARD · ADAMW
§ core

Core Architecture

§ concepts

Architecture & Concepts

§ guides

Guides & Playbooks

§ refs

API References