AGENTS & System Architecture

§ AGENTS.md 1,075 words τ ~5 min read
Read the workspace LLM/AGENTS.md and the parent CoreProjects/AGENTS.md
(+ self.md) first. Higher-level rules are authoritative; this file adds
project-specific rules only (and wins on Mamba-3-Lite-specific conflicts).

Quick checks (run before claiming work is done)#

bash
cd LLM/Mamba-3-Lite
python3 -m pytest tests/ -v          # CPU-friendly; triton tests auto-skip
python3 tests/test_doc_refs.py       # doc↔code symbol anchors
# GPU E2E (CUDA + triton required):
#   ENABLE_TRITON_KERNELS=1 python3 tests/e2e_gpu_smoke.py
Project: LLM/Mamba-3-Lite/ · Type: State-Space Model (Mamba-3)
Scale: ~434M params · 8.0B Chinchilla-optimal tokens · 12–15h on A100 80GB
Stack: PyTorch ≥2.1, no mamba-ssm, no custom CUDA (see rule #1 —
one sanctioned opt-in Triton kernel; everything else stays pure PyTorch).
Hardware: A100 80GB (no offloading).
Architecture detail: see README.md in this folder, docs/concepts/ssd-theory.md
(authoritative chunkwise-complex-SSD derivation) and the full docs/ tree
(concepts/ + references/ + guides/, see docs/README.md). Cross-project helper: root
AGENTS.md §2.12 (mamba2-ssd-engineer).

1. Subagent: mamba2-ssd-engineer (also covers Mamba-3)#

Triggers: "Explain complex-valued SSD", "Why is N halved with complex

states?", "How does MIMO head mixing replace SISO?", "Why no causal conv

in Mamba-3?", "Tune chunk_size for throughput."

Knows cold:

  • Faithful Mamba-3 (Dao & Gu, 2025) reproduction succeeding Mamba-2 with
    three architectural breakthroughs — all in pure PyTorch today:
    1. Complex-valued SSD state spaces. N=64 complex64 — half the state
      dimension of Mamba-2 (N=128) for parity perplexity. Two real
      sub-states packed into one complex state.
    2. MIMO (Multi-Input Multi-Output) head mixing. Fully-connected mixer
      across SSM heads replaces the classical SISO constraint.
    3. Zero causal convolution. The memory-bound causal_conv1d pass is
      eliminated; replaced by a purely chunked linear projection.
  • 28 layers · vocab 50,257 (GPT-2 BPE tokenizer) · ~434M params · 8.0B-token
    Chinchilla run planned.
  • Training: BF16 + torch.compile + TF32. FA2 disabled (we use the
    chunkwise recurrence). NaN guard with checkpoint rollback.
  • Pure SSM — no attention layers, no MoE, no MTP. Deliberate
    separation from HyMo (the portfolio's hybrid attention/SSM project).

Triton kernel contract:

  • Sanctioned Triton paths:
    • models/ssd_triton.py → per_chunk_ssd_triton (fused per-chunk
      pass for the complex-SSD chunkwise recurrence; fuses L materialisation,
      Y_diag, and per-chunk state into a single Triton kernel; opt-in
      via ssd_dispatch='triton' + ENABLE_TRITON_KERNELS=1).
  • When a kernel is added: place it in models/<name>_triton.py, gate
    on import triton with try/except ImportError setting
    HAS_TRITON = False, wrap in a torch.autograd.Function, add
    tests/test_<name>_triton.py with a CPU-runnable pure-PyTorch
    reference, and add the new path to the sanctioned list in rule #1.

2. Hard rules#

  1. **Raw PyTorch by default; custom Triton kernels are first-party for
    sanctioned hot paths.** Bulk of the codebase (complex SSD, MIMO
    mixer, chunkwise projection, RMSNorm, embeddings) stays raw
    PyTorch. No HuggingFace Trainer, no Lightning, no high-level
    wrappers. The sanctioned Triton paths are listed in §1 above. No
    new component gets a custom kernel without updating this file and
    adding a docs/references/<name>.md doc.
    • Conflict-resolution note: the previous "no Triton" hard rule
      (rule #5 in earlier versions of this file) is superseded by
      this rule. The current project ships one sanctioned Triton path
      (per_chunk_ssd_triton); the carve-out is in use.
    • ssd_dispatch opt-in is two-layered. A ssd_dispatch='triton'
      on the model config AND ENABLE_TRITON_KERNELS=1 in the
      environment are both required. Missing either one forces the
      dispatch back to 'pytorch' with a one-line warning (per-block
      warn-and-fallback in Mamba3Block, and a process-level guard in
      training/pretrain.py:_enforce_triton_env_var).
  2. Always read docs/concepts/ssd-theory.md (and, for
    context, docs/concepts/state-space-foundations.md and docs/concepts/mimo.md) before answering complex-SSD algorithm
    questions — it is the authoritative derivation (covers the chunkwise
    projection einsum-by-einsum, the complex recurrence, and MIMO mixer).
  3. Always verify the regression tests pass after any change to
    models/ssd_complex.py — the chunkwise linear projection must match the
    naive O(T) scan oracle exactly.
  4. Never pack a state size that doesn't decompose cleanly into real
    pairs (N must be even for complex packing to be exact).
  5. Never suggest adding MoE, MTP, or attention layers. Pure SSM
    repo (successor to Mamba-2).
  6. Never let a Triton kernel silently fall back to the raw-PyTorch
    path during a default-config training run. The opt-in is explicit
    (per-kernel config key + ENABLE_TRITON_KERNELS=1 env-var). If
    the kernel fails to compile or throws at runtime, the run must
    surface a clear error, not a silent fallback.
  7. Always add a unit test in tests/ for any new Triton kernel
    path. The test must run on CPU (using the pure-PyTorch reference)
    without triton installed. GPU-only behaviour is gated behind
    @pytest.mark.gpu and is auto-skipped on CPU-only machines.
  8. Concise comments only. Comments justify non-obvious code; they never restate the next line. The canonical, verifiable targets are in ../AGENTS.md §Cross-project engineering invariants — do not restate them here.
    Violations are reviewable on wc -l <file> and grep -c '^[[:space:]]*#' <file>.

3. Numerical-stability rules#

  • NaN guard with rollback (mirrors DeepSeek-v3-Lite pattern).
  • Recurrent state stays in complex64; logits in FP32.
  • Selective scan's gating sigmoid kept in FP32.
  • Complex64 multiplications monitored for NaN; torch.view_as_complex on
    raw real pairs can fail silently if the imaginary stride is wrong.

4. Files#

  • models/ssd_complex.py — complex SSD block + MIMO mixer + chunkwise projection.
  • models/ssd_triton.py — sanctioned Triton kernel (per_chunk_ssd_triton).
  • models/mamba_block.py — residual block (RMSNorm → SSD → MIMO → SwiGLU).
  • models/transformer.py — Mamba3Transformer + ModelConfig.
  • training/, data/, scripts/, tests/, utils/.
  • docs/ — concepts/ (concept-building), references/ (symbol-anchored API), guides/,
    and training.md (data pipeline). docs/README.md is the doc map.
    tests/test_doc_refs.py is the alignment checker.
  • README.md, requirements.txt, pytest.ini, LICENSE.

5. Known caveats#

  • Full 8B-token run not yet started.
  • FA2 is disabled — the chunkwise linear projection is used instead.
  • Complex states mean a downstream user can't trivially swap in a
    real-only recurrence without N→2N resize and a re-init.

6. Docs rule#

Docs ship with code; stale docs fail CI. Any change that adds, renames,

or removes a public symbol must update the docs that cite it. tests/test_doc_refs.py

parses every file.py:Symbol anchor in docs/ and fails on unknown files or

symbols — run it (and python3 -m pytest tests/) before committing doc or

code changes. New components (e.g. a sanctioned Triton kernel) must add a

docs/references/<name>.md doc following the style in the existing docs.

7. Obsidian Vault Rule#

Strict Whitelist Mirroring. All agents and skills must ensure the Obsidian

vault (~/Documents/obsidian/LLM/Mamba-3-Lite/) mirrors ONLY README.md,

AGENTS.md, SKILLS.md, and curated documentation under docs/**.

Generated HTML (docs_html/), tooling dumps, and non-doc artifacts are

strictly excluded and must never enter the vault. Always run bash scripts/sync_to_vault.sh

from workspace root after modifying docs.