The Sparsr architecture

The machine, the instruction set it covers today, and what it is designed to grow into.

The wide side and the scalar side

Sparsr is not a coprocessor bolted onto a host CPU. The wide side is the point of the design. The scalar side exists to drive it.

  • A 32-bit scalar core

    A conventional five-stage pipelined 32-bit RISC core with the usual 32-bit register file. It handles addressing, control and the ordinary integer work around a kernel. It is deliberately unremarkable — nothing about Sparsr is a bet on the scalar side.

  • An 8,192-bit wide register file

    A separate file of 32 registers of 8,192 bits each, with its own instructions. This is the width the machine is designed around, from the load through the operation to the store. A 512-bit vector extension beside a 64-bit core is a different arrangement.

The wide instruction set today

What is implemented and running. The Sparsr VM runs all of it at 8,192 bits. The FPGA hardware runs the same instructions at 4,096 bits today and is being widened. Stated plainly, because a processor company that is vague about its own opcode list is not worth evaluating.

InstructionWhat it does
WAND, WOR, WXORBitwise logic across the full 8,192 bits, one instruction per operation.
WL, WSWide load and store between on-chip wide memory and a wide register. The row number is a base register plus a displacement, the same form the scalar core uses for its own loads and stores, so a loop written by a compiler can walk the rows. A row is 8,192 bits, whatever its density.
ADD, SUB, AND, OR, XOR, NOR, LW, SWThe scalar side does ordinary 32-bit integer and memory operations.

Memory, as planned

Wide instructions work on whole 8,192-bit rows, sparse or dense alike. Every row a kernel can load is held whole, so density is not a limit on what Sparsr can hold or compute on.

On-chip wide memory holds 2 MB, which is 2,048 rows. Wide loads and stores address it by row number. The scalar core can also read and write those rows byte by byte, so a kernel can keep its results next to its data.

Larger data sets go in off-card memory: 16 GiB beside the chip. Wide instructions cannot address it: a copy engine brings rows from it into on-chip wide memory before the kernel needs them.

Compression happens in one place only, on the link from your machine to the card. The host runtime sends a sparse row as a compressed list and a dense row whole, and your kernel reads the same row either way.

The Sparsr VM models this design today. The FPGA hardware is being built to it, and hardware with this memory is not available yet.

In the planned memory path, your machine sends sparse rows compressed and dense rows whole; 16 GiB of off-card memory keeps them; a copy engine brings rows into 2 MB of on-chip wide memory, where every row is whole; wide instructions load them into 8,192-bit registers

The programming model

A Sparsr program is a host application plus one or more kernels — the same split as a GPU, and for the same reason.

  • The host runs the outer loop

    Control flow, iteration and orchestration live in the host application, written in C, C++, C# or Python. The runtime moves data and dispatches kernels.

  • The kernel is the wide part

    A kernel is the piece that touches the wide registers. Kernels are written in Sparsr assembly or in C/C++ through sparsr_intrinsics.h, and are straight-line wide-data code today.

  • The backend is a runtime choice

    libsparsr_host.so selects between the software emulator and FPGA hardware at runtime, so the same host application and the same kernel run against either without recompiling the host.

  • Most users should not have to write one

    For an established high-level domain the kernels come out the same every time, so they belong in a library rather than in your code. The first is torchhd-sparsr, a PyTorch device that runs standard Torchhd programs on Sparsr through an ordinary .to("sparsr") call. It is on PyPI (pip install torchhd-sparsr) and runs on the free Sparsr VM today. For every other domain, using Sparsr still means writing a kernel.

On the roadmap

Sparsr's instruction set is explicitly not fixed. Adding a primitive at 8,192 bits is ordinary roadmap work, so a missing instruction is a question of sequencing. These are the ones being worked on, and they are published because knowing what is not there yet is what lets you judge whether Sparsr fits your kernel.

Planned primitiveWhat it unlocks
A wide reduction unitPopulation count, parity, prefix-rank and threshold-accumulate over a wide register, in segmented per-lane form rather than only as a reduction to one scalar. This is the most requested primitive across every workload we have surveyed.
A per-block wide rotateRotation and permutation within a wide register — the binding operation for hyperdimensional computing, and for quasi-cyclic codes.
Off-card wide memory16 GiB beside the chip, filled from your machine and copied into on-chip wide memory by a dedicated engine. A data set much larger than the chip can then stay on the card between batches instead of crossing PCIe again.
Wide arithmeticCarry-propagating addition across the wide register, for the bit-parallel dynamic-programming kernels that need it.

Where it runs

Three destinations. The third is a partnership rather than a product, and it is the one people tend not to ask about.

  • On your own machine

    The software emulator runs the full instruction set on an ordinary CPU. It is the target you develop and test against, and it is free. You do not need hardware or an account.

  • On FPGA

    Sparsr's FPGA target is AWS EC2 F2, on a Virtex UltraScale+ HBM part. F2 is the development and validation vehicle for the design rather than a finished product baseline — it is where the RTL is proven. You can run Sparsr on an F2 instance in your own AWS account today.

  • On your own silicon

    Sparsr is a processor design, so it does not have to run on hardware we operate. If you already build silicon or sell your own FPGA-based products, the wide register file and its instruction set can be embedded in your device as licensed IP. We can specify instructions around your workload rather than ours. We are open to that, and would rather hear from you early in your design cycle than late.

Does your kernel fit?

A short test. Your kernel is worth a conversation if its inner loop is dominated by bitwise operations over long bit vectors. The data can be sparse or dense. That holds even when the primitive you need is one of the roadmap items above.

See the application areas Talk to us