Read the workspaceLLM/AGENTS.mdand the parentCoreProjects/AGENTS.md
(+self.md) first. Higher-level rules are authoritative; this file adds
project-specific rules only (and wins on Mamba-3-Lite-specific conflicts).
Quick checks (run before claiming work is done)#
cd LLM/Mamba-3-Lite
python3 -m pytest tests/ -v # CPU-friendly; triton tests auto-skip
python3 tests/test_doc_refs.py # doc↔code symbol anchors
# GPU E2E (CUDA + triton required):
# ENABLE_TRITON_KERNELS=1 python3 tests/e2e_gpu_smoke.pyProject:LLM/Mamba-3-Lite/· Type: State-Space Model (Mamba-3)
Scale: ~434M params · 8.0B Chinchilla-optimal tokens · 12–15h on A100 80GB
Stack: PyTorch ≥2.1, nomamba-ssm, no custom CUDA (see rule #1 —
one sanctioned opt-in Triton kernel; everything else stays pure PyTorch).
Hardware: A100 80GB (no offloading).
Architecture detail: seeREADME.mdin this folder,docs/concepts/ssd-theory.md
(authoritative chunkwise-complex-SSD derivation) and the fulldocs/tree
(concepts/ + references/ + guides/, seedocs/README.md). Cross-project helper: rootAGENTS.md §2.12(mamba2-ssd-engineer).
1. Subagent: mamba2-ssd-engineer (also covers Mamba-3)#
Triggers: "Explain complex-valued SSD", "Why is N halved with complex
states?", "How does MIMO head mixing replace SISO?", "Why no causal conv
in Mamba-3?", "Tune chunk_size for throughput."
Knows cold:
- Faithful Mamba-3 (Dao & Gu, 2025) reproduction succeeding Mamba-2 with
three architectural breakthroughs — all in pure PyTorch today:- Complex-valued SSD state spaces. N=64 complex64 — half the state
dimension of Mamba-2 (N=128) for parity perplexity. Two real
sub-states packed into one complex state. - MIMO (Multi-Input Multi-Output) head mixing. Fully-connected mixer
across SSM heads replaces the classical SISO constraint. - Zero causal convolution. The memory-bound
causal_conv1dpass is
eliminated; replaced by a purely chunked linear projection.
- Complex-valued SSD state spaces. N=64 complex64 — half the state
- 28 layers · vocab 50,257 (GPT-2 BPE tokenizer) · ~434M params · 8.0B-token
Chinchilla run planned. - Training: BF16 +
torch.compile+ TF32. FA2 disabled (we use the
chunkwise recurrence). NaN guard with checkpoint rollback. - Pure SSM — no attention layers, no MoE, no MTP. Deliberate
separation from HyMo (the portfolio's hybrid attention/SSM project).
Triton kernel contract:
- Sanctioned Triton paths:
models/ssd_triton.py→per_chunk_ssd_triton(fused per-chunk
pass for the complex-SSD chunkwise recurrence; fusesLmaterialisation,
Y_diag, and per-chunkstateinto a single Triton kernel; opt-in
viassd_dispatch='triton'+ENABLE_TRITON_KERNELS=1).
- When a kernel is added: place it in
models/<name>_triton.py, gate
onimport tritonwithtry/except ImportErrorsetting
HAS_TRITON = False, wrap in atorch.autograd.Function, add
tests/test_<name>_triton.pywith a CPU-runnable pure-PyTorch
reference, and add the new path to the sanctioned list in rule #1.
2. Hard rules#
- **Raw PyTorch by default; custom Triton kernels are first-party for
sanctioned hot paths.** Bulk of the codebase (complex SSD, MIMO
mixer, chunkwise projection, RMSNorm, embeddings) stays raw
PyTorch. No HuggingFace Trainer, no Lightning, no high-level
wrappers. The sanctioned Triton paths are listed in §1 above. No
new component gets a custom kernel without updating this file and
adding adocs/references/<name>.mddoc.- Conflict-resolution note: the previous "no Triton" hard rule
(rule #5 in earlier versions of this file) is superseded by
this rule. The current project ships one sanctioned Triton path
(per_chunk_ssd_triton); the carve-out is in use. ssd_dispatchopt-in is two-layered. Assd_dispatch='triton'
on the model config ANDENABLE_TRITON_KERNELS=1in the
environment are both required. Missing either one forces the
dispatch back to'pytorch'with a one-line warning (per-block
warn-and-fallback inMamba3Block, and a process-level guard in
training/pretrain.py:_enforce_triton_env_var).
- Conflict-resolution note: the previous "no Triton" hard rule
- Always read
docs/concepts/ssd-theory.md(and, for
context,docs/concepts/state-space-foundations.mdanddocs/concepts/mimo.md) before answering complex-SSD algorithm
questions — it is the authoritative derivation (covers the chunkwise
projection einsum-by-einsum, the complex recurrence, and MIMO mixer). - Always verify the regression tests pass after any change to
models/ssd_complex.py— the chunkwise linear projection must match the
naive O(T) scan oracle exactly. - Never pack a state size that doesn't decompose cleanly into real
pairs (N must be even for complex packing to be exact). - Never suggest adding MoE, MTP, or attention layers. Pure SSM
repo (successor to Mamba-2). - Never let a Triton kernel silently fall back to the raw-PyTorch
path during a default-config training run. The opt-in is explicit
(per-kernel config key +ENABLE_TRITON_KERNELS=1env-var). If
the kernel fails to compile or throws at runtime, the run must
surface a clear error, not a silent fallback. - Always add a unit test in
tests/for any new Triton kernel
path. The test must run on CPU (using the pure-PyTorch reference)
withouttritoninstalled. GPU-only behaviour is gated behind
@pytest.mark.gpuand is auto-skipped on CPU-only machines. - Concise comments only. Comments justify non-obvious code; they never restate the next line. The canonical, verifiable targets are in
../AGENTS.md§Cross-project engineering invariants — do not restate them here.
Violations are reviewable onwc -l <file>andgrep -c '^[[:space:]]*#' <file>.
3. Numerical-stability rules#
- NaN guard with rollback (mirrors DeepSeek-v3-Lite pattern).
- Recurrent state stays in
complex64; logits in FP32. - Selective scan's gating sigmoid kept in FP32.
- Complex64 multiplications monitored for NaN;
torch.view_as_complexon
raw real pairs can fail silently if the imaginary stride is wrong.
4. Files#
models/ssd_complex.py— complex SSD block + MIMO mixer + chunkwise projection.models/ssd_triton.py— sanctioned Triton kernel (per_chunk_ssd_triton).models/mamba_block.py— residual block (RMSNorm → SSD → MIMO → SwiGLU).models/transformer.py—Mamba3Transformer+ModelConfig.training/,data/,scripts/,tests/,utils/.docs/— concepts/ (concept-building), references/ (symbol-anchored API), guides/,
andtraining.md(data pipeline).docs/README.mdis the doc map.
tests/test_doc_refs.pyis the alignment checker.README.md,requirements.txt,pytest.ini,LICENSE.
5. Known caveats#
- Full 8B-token run not yet started.
- FA2 is disabled — the chunkwise linear projection is used instead.
- Complex states mean a downstream user can't trivially swap in a
real-only recurrence without N→2N resize and a re-init.
6. Docs rule#
Docs ship with code; stale docs fail CI. Any change that adds, renames,
or removes a public symbol must update the docs that cite it. tests/test_doc_refs.py
parses every file.py:Symbol anchor in docs/ and fails on unknown files or
symbols — run it (and python3 -m pytest tests/) before committing doc or
code changes. New components (e.g. a sanctioned Triton kernel) must add a
docs/references/<name>.md doc following the style in the existing docs.
7. Obsidian Vault Rule#
Strict Whitelist Mirroring. All agents and skills must ensure the Obsidian
vault (~/Documents/obsidian/LLM/Mamba-3-Lite/) mirrors ONLY README.md,
AGENTS.md, SKILLS.md, and curated documentation under docs/**.
Generated HTML (docs_html/), tooling dumps, and non-doc artifacts are
strictly excluded and must never enter the vault. Always run bash scripts/sync_to_vault.sh
from workspace root after modifying docs.