Where Sparsr fits

The application areas we have surveyed in depth, what each one needs from the machine, and what a product built in it looks like. Just as usefully, what Sparsr is not for.

The shape of a problem that suits Sparsr

These are the domains we have taken apart kernel by kernel: what each one would actually run on, and what a product in that field looks like.

  • Cheminformatics

    All-versus-all similarity screening over molecular fingerprints. A fingerprint of up to 8,192 bits is one wide register, and against a transposed library the kernel needs only the wide logic and load/store instructions Sparsr already has. Our reference kernel matches RDKit exactly across every pair we have checked. In practice that is virtual screening: a drug discovery team narrowing a library of millions of compounds down to the few hundred worth testing at the bench.

  • Hyperdimensional computing

    A hypervector is one register rather than a tiled array, which removes most of the code that makes VSA implementations awkward on a CPU. Binding runs today. Bundling and similarity search want the reduction unit and the rotate from the roadmap. The products built on it are the ones that pick up a new class from one or two examples, on the device, with no retraining run. Seizure detection tuned to a single patient is one. A prosthetic hand recalibrated to the person wearing it is another.

  • Quantum error correction

    Stabilizer simulation and Pauli-frame sampling are wide XOR plus a population count. The working set for frame sampling is small enough to sit in on-chip memory, which is exactly the case a wide machine handles well. That is the decoding layer every quantum computer needs between its qubits and an answer, and it has to keep pace with the machine in real time.

  • Logic and fault simulation

    Bit-parallel fault simulation puts one test pattern in each bit of a wide word, which is the cleanest zero-gap fit we have found: the kernel is pure wide bitwise logic on the instruction set as it stands. Chips pass this before they are manufactured. It produces the test patterns that prove a fabricated die is free of defects.

  • Genomics and bioinformatics

    The kernels here have three different costs. Bitsliced k-mer search runs on today's instruction set. Population-count-based genotype work and bit-parallel sequence alignment each wait on a roadmap primitive. The products here are variant calling and sequence search, the step between a sequencing machine and a clinical or research answer.

  • Error-correcting codes

    The purest arithmetic fit in the survey: addition over GF(2) is WXOR and multiplication is WAND, with no translation in between. Syndrome computation wants a parity reduction. Quasi-cyclic codes want the wide rotate. Every storage device and every radio link runs this. It is the layer that recovers the original data when some of the bits arrive wrong.

  • Statistical physics

    Monte Carlo on spin glasses with one disorder sample per bit. An 8,192-bit word holds 8,192 independent copies of a lattice site, and one Metropolis update is a few wide logic instructions over all of them at once. No bit ever moves between lanes, so the kernel needs neither a population count nor a rotate. What it needs is memory: a lattice does not fit the on-chip rows, and that is the measured limit today. The planned off-card memory is what lets a lattice stay on the card. That is materials and disordered-systems research, where a result comes from running the same simulation over thousands of independent samples.

What runs today, and what is still being built

The table asks two separate questions, so it has two columns. Hardware says whether the kernel is expressible on the instruction set as it stands. Access says what using it looks like from your side. For all but one of these, that still means a kernel someone has to write.

Application areaHardwareAccess
Cheminformatics — fingerprint screening✅ Runs today⌨️ Custom kernel
Logic and fault simulation✅ Runs today⌨️ Custom kernel
Genomics — bitsliced k-mer search✅ Runs today⌨️ Custom kernel
Statistical physics — multi-spin Monte Carlo✅ Update kernel runs today; bit-parallel RNG needs the wide rotate 🚧⌨️ Custom kernel
Information retrieval — binary nearest-neighbour search🚧 Needs the reduction unit📦 Library in development
Hyperdimensional computing and VSA🚧 Needs the reduction unit and the rotate📦 torchhd-sparsr, on PyPI
Quantum error correction and stabilizer simulation🚧 Needs the reduction unit⌨️ Custom kernel
Error-correcting codes🚧 Needs a parity reduction mode; quasi-cyclic codes also need the rotate⌨️ Custom kernel
Binary neural networks🚧 Needs the reduction unit and a wide XNOR⌨️ Custom kernel
Bitmap-index analytics and sparse query indexing🚧 Needs the reduction unit and a bit-level container format⌨️ Custom kernel
Robotics — offline place-recognition and sensor fusion🚧 Downstream of the hyperdimensional computing work, and asks for no hardware of its own📦 Follows the HDC library

Reading the two columns

The distinction that decides whether you evaluate Sparsr now or later.

  • ✅ Runs today

    The kernel is expressible on the instruction set as it stands — nothing in the hardware roadmap is in the way. An evaluation is cheapest to start here, because the only thing between you and a result is the kernel itself.

  • 🚧 Being built

    The kernel needs a wide instruction that does not exist yet. Each one is a tracked piece of work rather than an aspiration, and the architecture page lists them. Our own build order is the only thing holding these back.

  • ⌨️ Custom kernel

    Using Sparsr for this today means writing the wide-data inner loop yourself, in assembly or in C/C++ through the intrinsics header. The free emulator means you can do that now, without hardware.

  • 📦 Library

    A library with the kernels already in it, so you write ordinary domain code and never touch the wide registers. The first, torchhd-sparsr, is a PyTorch device for Torchhd. It is on PyPI (pip install torchhd-sparsr) and runs on the free Sparsr VM today. The others in this column are what is coming, not what you can install today.

  • Not on the list?

    The pattern generalises. If your domain has a standard library with a stable set of inner loops, that library is exactly the kind of thing worth putting a Sparsr backend behind. Which one we build next is open to discussion.

What Sparsr is not for

Every one of these is a real architectural mismatch rather than a gap we intend to close, and saying so is more useful to you than another row in the table above.

  • Floating-point numerics

    Sparsr has no floating-point unit and is not getting one. If the values in your matrix are real numbers rather than bits, this is the wrong machine — including for sparse linear solvers and for most scientific simulation. Sparsity does not change that answer.

  • Dense integer matrix multiply

    Once operands are 8 bits or wider, a wide bitwise datapath is the wrong shape for the multiply-accumulate involved. Dense GEMM, and the neural network inference built on it, belong on a GPU or a tensor unit.

  • Pointer-chasing traversal

    Graph traversal and index structures whose cost is fine-grained gather and scatter are bound by memory latency, not by how wide a register is. Width does not help, and we do not claim it does.

A missing instruction is not a rejection

Sparsr's instruction set is designed to grow, and adding a wide primitive is normal roadmap work rather than an exception. So if your kernel is bitwise and wide, sparse or dense, the interesting question is not whether Sparsr can run it today — it is what it would cost to make it run well.

That conversation is the most useful thing we can offer a research group right now, and it is free.

Research partnerships See the instruction set