The Sparsr architecture
The machine, the instruction set it covers today, and what it is designed to grow into.
The wide side and the scalar side
Sparsr is not a coprocessor bolted onto a host CPU. The wide side is the point of the design. The scalar side exists to drive it.
-
A 32-bit scalar core
A conventional five-stage pipelined 32-bit RISC core with the usual 32-bit register file. It handles addressing, control and the ordinary integer work around a kernel. It is deliberately unremarkable — nothing about Sparsr is a bet on the scalar side.
-
An 8,192-bit wide register file
A separate file of 32 registers of 8,192 bits each, with its own instructions. This is the width the machine is designed around, from the load through the operation to the store. A 512-bit vector extension beside a 64-bit core is a different arrangement.
The wide instruction set today
What is implemented and running. The Sparsr VM runs all of it at 8,192 bits. The FPGA hardware runs the same instructions at 4,096 bits today and is being widened. Stated plainly, because a processor company that is vague about its own opcode list is not worth evaluating.
| Instruction | What it does |
|---|---|
| WAND, WOR, WXOR | Bitwise logic across the full 8,192 bits, one instruction per operation. |
| WL, WS | Wide load and store between on-chip wide memory and a wide register. The row number is a base register plus a displacement, the same form the scalar core uses for its own loads and stores, so a loop written by a compiler can walk the rows. A row is 8,192 bits, whatever its density. |
| ADD, SUB, AND, OR, XOR, NOR, LW, SW | The scalar side does ordinary 32-bit integer and memory operations. |
Memory, as planned
Wide instructions work on whole 8,192-bit rows, sparse or dense alike. Every row a kernel can load is held whole, so density is not a limit on what Sparsr can hold or compute on.
On-chip wide memory holds 2 MB, which is 2,048 rows. Wide loads and stores address it by row number. The scalar core can also read and write those rows byte by byte, so a kernel can keep its results next to its data.
Larger data sets go in off-card memory: 16 GiB beside the chip. Wide instructions cannot address it: a copy engine brings rows from it into on-chip wide memory before the kernel needs them.
Compression happens in one place only, on the link from your machine to the card. The host runtime sends a sparse row as a compressed list and a dense row whole, and your kernel reads the same row either way.
The Sparsr VM models this design today. The FPGA hardware is being built to it, and hardware with this memory is not available yet.
The programming model
A Sparsr program is a host application plus one or more kernels — the same split as a GPU, and for the same reason.
-
The host runs the outer loop
Control flow, iteration and orchestration live in the host application, written in C, C++, C# or Python. The runtime moves data and dispatches kernels.
-
The kernel is the wide part
A kernel is the piece that touches the wide registers. Kernels are written in Sparsr assembly or in C/C++ through sparsr_intrinsics.h, and are straight-line wide-data code today.
-
The backend is a runtime choice
libsparsr_host.so selects between the software emulator and FPGA hardware at runtime, so the same host application and the same kernel run against either without recompiling the host.
-
Most users should not have to write one
For an established high-level domain the kernels come out the same every time, so they belong in a library rather than in your code. The first is torchhd-sparsr, a PyTorch device that runs standard Torchhd programs on Sparsr through an ordinary .to("sparsr") call. It is on PyPI (pip install torchhd-sparsr) and runs on the free Sparsr VM today. For every other domain, using Sparsr still means writing a kernel.
On the roadmap
Sparsr's instruction set is explicitly not fixed. Adding a primitive at 8,192 bits is ordinary roadmap work, so a missing instruction is a question of sequencing. These are the ones being worked on, and they are published because knowing what is not there yet is what lets you judge whether Sparsr fits your kernel.
| Planned primitive | What it unlocks |
|---|---|
| A wide reduction unit | Population count, parity, prefix-rank and threshold-accumulate over a wide register, in segmented per-lane form rather than only as a reduction to one scalar. This is the most requested primitive across every workload we have surveyed. |
| A per-block wide rotate | Rotation and permutation within a wide register — the binding operation for hyperdimensional computing, and for quasi-cyclic codes. |
| Off-card wide memory | 16 GiB beside the chip, filled from your machine and copied into on-chip wide memory by a dedicated engine. A data set much larger than the chip can then stay on the card between batches instead of crossing PCIe again. |
| Wide arithmetic | Carry-propagating addition across the wide register, for the bit-parallel dynamic-programming kernels that need it. |
Where it runs
Three destinations. The third is a partnership rather than a product, and it is the one people tend not to ask about.
-
On your own machine
The software emulator runs the full instruction set on an ordinary CPU. It is the target you develop and test against, and it is free. You do not need hardware or an account.
-
On FPGA
Sparsr's FPGA target is AWS EC2 F2, on a Virtex UltraScale+ HBM part. F2 is the development and validation vehicle for the design rather than a finished product baseline — it is where the RTL is proven. You can run Sparsr on an F2 instance in your own AWS account today.
-
On your own silicon
Sparsr is a processor design, so it does not have to run on hardware we operate. If you already build silicon or sell your own FPGA-based products, the wide register file and its instruction set can be embedded in your device as licensed IP. We can specify instructions around your workload rather than ours. We are open to that, and would rather hear from you early in your design cycle than late.
Does your kernel fit?
A short test. Your kernel is worth a conversation if its inner loop is dominated by bitwise operations over long bit vectors. The data can be sparse or dense. That holds even when the primitive you need is one of the roadmap items above.