Expand description
GGML quantization block layout and zero-heap row dequantization.
Byte strides match ggml_row_size() in llama.cpp / ggml. Embedding lookup slices
raw mmap bytes via fetch_token_embedding; this module dequantizes into
caller-supplied &mut [f32] buffers (no Vec in the hot path).
Structs§
- Block
Q6K - GGML
block_q6_K— 210 bytes, 256 weights. Mirrors WGSLBlockQ6Klayout. - Ggml
Block Layout - Elements per quantization block and packed byte size (from ggml).
Enums§
- Execution
Error - Errors from zero-copy mmap tensor slicing.
- Ggml
Dequant Error
Constants§
- BLOCK_
Q4K_ SOA_ BYTES - Bytes per SoA superblock (256 weights).
- BLOCK_
Q4K_ SOA_ ELEMS - BLOCK_
Q6K_ BYTES - BLOCK_
Q6K_ ELEMS - GGML_
TYPE_ BF16 - Brain float16 (1 sign / 8 exp / 7 mantissa) — used by Gemma-4 and other modern GGUFs
for norms / residual scales alongside Q4_K weights (
ggml_typeenum value 30). - GGML_
TYPE_ F16 - GGML_
TYPE_ F32 - GGML element-type identifiers used in GGUF tensor-info headers.
- GGML_
TYPE_ Q4_ 0 - GGML_
TYPE_ Q4_ K - GGML_
TYPE_ Q4_ K_ SOA - Qualia conversion-time SoA Q4_K (not a stock GGML type).
- GGML_
TYPE_ Q5_ 0 - GGML_
TYPE_ Q6_ K - GGML_
TYPE_ Q8_ 0
Functions§
- dequant_
matrix_ row_ into - Dequantize one matrix row (
rowindex alongdims[1]) intoout. - dequantize_
row_ into - Dequantize one embedding row from raw mmap bytes into
out. Returns the number off32elements written (≤out.len()). - expand_
q4k_ tensor_ to_ soa - Expand a full Q4_K tensor blob (row-major superblocks) into SoA layout.
n_row_elems= dims[0] (weights per row).n_rows= dims[1]. - fetch_
tensor_ bytes - Zero-copy slice of an entire tensor payload from the mmap.
- fetch_
tensor_ row_ range_ bytes - Zero-copy slice covering vocabulary rows
[row_start, row_start + row_count). - fetch_
token_ embedding - Return a zero-copy
&[u8]slice of the packed embedding row fortoken_id. - ggml_
block_ layout - Return block layout for a GGML type, or
Noneif unsupported. - ggml_
row_ bytes - Packed byte length of one logical row (
n_elemsweights) for the given GGML type. - q4k_
block_ to_ soa - Convert one stock Q4_K superblock (144 B) → SoA superblock (160 B).
- quantize_
f32_ to_ q4_ k_ soa_ tensor - Quantize a full f32 weight matrix to Q4_K_SOA layout.
- tensor_
byte_ len - Total packed byte length of a GGUF tensor from its shape and
ggml_type. - tensor_
row_ byte_ len - Packed byte width of one logical matrix row (
dims[0]elements).