Expand description
End-to-end Webizen-gated chat inference with retrieval, streaming, and provenance.
Structs§
Functions§
- clear_
cancel_ inference - is_
inference_ cancelled - request_
cancel_ inference - run_
chat_ inference_ for_ agent - Run a local turn using a named roster agent. A pinned model is activated on demand before inference; the lifecycle implementation owns any resident mapping replacement. This is a cold control-path operation and never runs inside the decode/evaluator hot path.
- run_
chat_ inference_ full - run_
chat_ inference_ with_ options - stream_
event_ done - stream_
event_ error - stream_
event_ token - NDJSON stream events for Flutter FRB.
- validate_
axiom_ preflight - Pre-flight axiom bounds check before KV prefill / orchestrator dispatch.