Foundation in formation Interim stewardship: Celiums Solutions LLC Read the status note

Hyphae BitNet Exact ternary CPU inference

Use CPU memory and arithmetic deliberately.

Hyphae BitNet is a CPU runtime for supported ternary models. Its current executable is hyphae-bitnet, with celiums-bitnet and historical ABI names retained for compatibility.

Make supported model inference inspectable through explicit numerical, model, memory, and hardware contracts.

01 / Why it matters

Why this work matters

Exact numerical rules

Use declared integer accumulation contracts with reference checks.

Bounded working memory

Refuse model or session allocations that exceed the configured memory budget.

Supported-model gates

Admit explicit model families and formats instead of implying arbitrary model support.

Measurable behavior

Report prefill and decode separately, with model, CPU, and workload identities.

02 / How it works

How it works

The main stages make the project’s boundaries visible.

  1. 01

    Validate the model

    Check the supported family, format, and operator-pinned model identity

  2. 02

    Budget memory

    Admit the selected compute layout within the working-set cap

  3. 03

    Run CPU kernels

    Current ARM and x86 Q1 paths use packed panels; prefill and decode differ

  4. 04

    Expose results

    CLI, experimental C ABI, HTTP subset, and benchmark records

03 / Capabilities

Current capabilities

Model families

Supported I2_S BitNet and Q1_0 Bonsai CPU text paths, with explicit family selection.

Runtime interfaces

Generation, native HTTP serving, benchmarks, validation, and an experimental C ABI.

CPU paths

Linux x86-64 archives and measured ARM64 Q1 work; NEON DOTPROD remains the measured Graviton4 default.

Optional gateway

A separate process connects Hyphae retrieval, memory, and receipts to the runtime.

04 / Evidence

Evidence and its scope

The v0.3.3 release publishes source/runtime version 0.3.2 and records Graviton4 validation and A/B measurements.

0.3.2

runtime version

Distributed under tag v0.3.3

CPU

current product target

SVE2

experimental packed kernel

Earlier expand-to-int8 speedups describe historical ARM code. The current ARM path uses packed Q1 panels; results vary by phase, model, layout, and thread count.

05 / What it does now

What it does now

  • Runs supported BitNet I2_S and Bonsai Q1_0 CPU text models.
  • Provides run, serve, bench, validate, and version commands.
  • Retains the published v0.3.0 C ABI and adds sized extension APIs for new options.
06 / What it intends to do

Next directions

  • Measure further exact CPU kernel changes against pinned references.
  • Improve scheduling and request ownership before broader serving claims.
  • Keep experimental accelerator work behind explicit validation boundaries.
07 / Explicit boundaries

Limits to keep in view

  • The source version is 0.3.2 even though the public release tag is v0.3.3.
  • SVE2 is experimental; historical expanded ARM layouts are not the current default.
  • Binaries compiled against the unreleased broken v0.3.1 option layouts must be rebuilt.
  • No supported GPU inference, arbitrary model loading, continuous batching, or complete HTTP API compatibility.
  • Third-party engine and model rights remain with their respective owners.

Hyphae BitNet

Make supported model inference inspectable through explicit numerical, model, memory, and hardware contracts.