1bit.MONSTER
One engine, any model, zero Python. A model-agnostic, hardware-agnostic pure-C++ inference engine. This is the index to the whole site.
latest Every model sits in one shared pool — SharedBO pages that NPU, GPU and host all alias, zero copies between lanes. Zaya 8B Q4NX runs entirely on the Strix Halo NPU via the HRX lane (decode 8.4 t/s, served over HTTP), Qwen3-0.6B prefill logits are byte-identical to the real FastFlowLM runtime (corr 1.00000), and the Windows-NPU memory model — one heap, one chained API call per layer — is matched on Linux. read the post →AMD's XDNA 2 NPU shipped closed — 22 proprietary .so files, zero docs. Reverse-engineered from scratch in the open.
One engine, every model →99.89% of HuggingFace's arch-bearing checkpoints map to a token this engine runs. Same binary on NPU, GPU, or CPU.
JARVIS, out of the box →Mic → VAD → STT → LLM → TTS → speaker. One process, zero cloud, ships with the engine.
The claim, the census, the console. One engine, any model, zero Python.
engine Models →The families the engine runs, with real sizes and backend availability.
models 1bit JARVIS →The in-process voice assistant that ships with the engine.
jarvis Blog →The build log: the four-day NPU reversal, the gates, the journey.
blog Store →Merch with real inventory and a working cart. Checkout on the live store.
store Docs →From scratch to NPU. The docs index, routed to the real sources.
docs Benchmarks →Measured, not projected. Kernel-level and end-to-end numbers.
benchmarks