A Rust-based small language model using a recurrent state-space layer, episodic memory bank and plastic weights instead of a transformer.
At a Glance
PSSA is free to use, modify, and distribute under the GPL-3.0 license.
Engagement
Available On
Alternatives
Listed Oct 2026
About PSSA
PSSA (plastic state-space architecture) is an open-source research prototype published on GitHub by Sparticle62ops under the GPL-3.0 license. It is a small language model written in Rust from scratch, with no PyTorch, TensorFlow or other ML framework underneath. The README describes a 1.5M-parameter model trained on WikiText-103 and presents it as a research prototype rather than a competitor to established models.
What It Is
PSSA is a non-transformer language model that reads text one token at a time through a recurrent state-space layer. According to the README, it keeps a bank of episodic memories it can look up and rewrites part of its own weights while it runs. The project ships as a Rust command-line program, oxide_ai_pssa, with commands to train, generate, chat, score checkpoints, download datasets and clean WikiText corpora. It also includes a transformer baseline for comparison.
How the architecture works
The README describes each token passing through one layer made of a selective state-space recurrence, a bounded read from an episodic memory bank in hyperbolic space, a learned gate, and a SiLU MLP. The recurrence is described as the same family as S4 and Mamba and claims no novelty. The README lists the project's own claims as:
- A hyperbolic bounded read, limited to four slots, conditioned on the recurrent state.
- Novelty and refractory rules on memory writes.
- A closed-form ridge-regression step that folds fast plastic updates into the base transition matrix.
The defaults are 256 channels, 16 states per channel, a rank-16 adapter and a 512-slot memory bank.
Reported results
The README reports a comparison against a parameter-matched transformer on the same corpus, tokenizer, optimizer schedule and seed, over 12.7M tokens of cleaned WikiText-103. It states that PSSA reached a training cross-entropy of 3.98 versus 4.43 for the transformer. On a 198,939-token held-out slice it reports 3.997 versus 4.429 cross-entropy and 24.1% versus 18.0% next-token accuracy. It also reports CPU throughput of 1,716 versus 415 tokens per second on a 2-vCPU machine, and 200-token generation in 226 ms versus 2,735 ms.
Limits the author states
The README is explicit about scale. Text quality is poor for both models, and the comparison is about learning efficiency rather than fluency. The author notes that part of the gap could reflect an undertrained baseline, and that two experiments, retention after a corpus switch and ablation of the memory bank, are not yet measured. The limitations section also describes it as a CPU-oriented prototype with a minimal CLI parser and project-specific binary checkpoint files.
Setup path
Building requires a Rust toolchain with Edition 2024 support. The README says network access is needed only for HTTP or Hugging Face datasets. An optional CUDA path, enabled with the cuda feature and using cuBLAS, is available for training, and a WebGPU probe exists; everything falls back to the CPU implementation. Datasets can be local files, directories, URLs or hf:owner/dataset sources, and long corpora can be trained as chained, resumable runs. Tests run with cargo test --release.
Where the project seeks help
The author says the project most needs GPU compute for larger-scale tests, plus contributions on kernel performance, a modern recurrent baseline and evaluation beyond next-token loss. A Discord community is linked from the README.
Community Discussions
Be the first to start a conversation about PSSA
Share your experience with PSSA, ask questions, or help others learn from your insights.
Pricing
Open Source
PSSA is free to use, modify, and distribute under the GPL-3.0 license.
- Written in Rust from scratch, with no ML framework dependency
- GPL-3.0 licensed source code
- CPU reference path always available
- Optional CUDA device for the GPU training path
Capabilities
Key Features
- Recurrent selective state-space core with fixed-size state
- Episodic memory bank with 512 slots and bounded top-4 hyperbolic retrieval
- Plastic weights with novelty-driven writes and refractory rate limiting
- Closed-form ridge-regression consolidation into the transition matrix
- Hand-written Rust linear algebra with a scalar reference path for gradient checks
- Byte-level BPE tokenizer
- Optional CUDA (cuBLAS) and WebGPU acceleration
- Chained resumable training over long corpora
- CLI for train, generate, chat, score, throughput, download and WikiText cleaning
- Parameter-matched transformer baseline for comparison
