Repository navigation
Releases: basecompute/baseRT
Release list
BaseRT engine 0.3.0
BaseRT 0.3.0
basert serve is a new server. It keeps the command line and the
OpenAI-compatible API it has always had, and runs them on
superfluid, our serving daemon:
requests are decoded together through one shared forward pass,
a prompt a conversation repeats is reused instead of read again, and hybrid
models (Qwen 3.5, 3.6 and 3.8) get context windows up to four times larger
from the same memory. This release also carries the fixes prepared for 0.2.7,
which was never published.
New
- Serving. Continuous batching and prefix reuse are always on; the server
runs one request at a time unless--continuous-batching [N]or
--max-batch-size Nasks for more. A server is ready sooner (Qwen3-8B on an
M5 Pro: 5.3 s to 1.0 s). A streamed tool call arrives as partial deltas —
the function's name first, then its arguments as they are produced — and
tool_choicenaming a function, orrequired, is honoured./v1/completions
acceptssuffixfor fill-in-the-middle./v1/modelslists each model's
capabilities, including whether its template takes a thinking switch and
which reasoning-effort levels it defines. - Larger windows for hybrid models. The paged KV cache charges only the
attention layers, so on a 48 GB M5 Pro the automatically sized window grows
from 72,704 to 262,144 tokens on Qwen3.6-35B-A3B and Qwen3.5-35B-A3B, from
31,744 to 130,048 on Qwen3.6-27B, and from 27,648 to 111,616 on Qwen3.8-27B.
Dense models keep the window they had. - GLM 5.2 on Metal serves with continuous batching, and speculates with
the multi-token-prediction head its bundle carries. - Speculation. Block drafters with a reduced draft vocabulary (the DFlash
heads published for Qwen3-30B-A3B and Nemotron-3-Super) convert and serve,
token-for-token identical to plain decoding under greedy sampling. A sampled
round verifies all of its rows on the GPU in one pass (an 8-lane round on
Qwen3.8-27B: 109 ms to 93 ms). A strategy that loses to plain decoding is
measured against a real plain run and retired, then rechecked later rather
than written off.basert complete --specturns speculation on for a bundle
with its own head. - Throughput under agent load. Sampled lanes no longer sort the whole
vocabulary per token, served lanes sample on the GPU, and prompts are read
prefill-first. Eight agents across six turns on Qwen3.6-35B-A3B (M5 Max):
wall time 345 s to 136 s, first token at the median 46 s to 6 s. - Automatic context sizing on Metal plans within the memory the engine can
keep pinned, and counts the recurrent state a hybrid model allocates per
lane, so a derived window no longer pushes the weights out during a long
prompt.
Changed
- Flags the old server took but this one has no use for are accepted and
noted on stderr:--paged-kvand--prefix-cache(always on),
--prefix-cache-fileand--prefix-cache-save-interval(the cache is not
kept on disk),--files-dirand--media-dir(files live in the server's
state directory; images and audio arrive asdata:URIs or uploads), and
the engine settings--metallib,--prefill-chunk,--paged-weights,
--no-paged-weights-retry,--gpu-wait-timeout-ms,--decode-replayand
--no-baked-decode.--request-timeoutbounds fill-in-the-middle
completions only. A flag value that is not a number is refused rather than
read as 0.--model-diralone loads its first model at startup.--verbose
prints the load and GPU diagnostics without the per-kernel dispatch list. - State — the request log and uploaded files — lives under
$BASERT_STATE_DIR
(default:serve/beside the model cache), one directory per port, or
wherever--sessions <dir>says; the request log starts fresh at each start. /propsno longer reportstrained_context;/v1/modelsreports the
serving window asmeta.n_ctx; replies saysystem_fingerprint: superfluid;
/metricsseries are namedsuperfluid_*. Error messages are worded
differently. Without--api-keyon a non-loopback address, the server
answers only requests addressed to an IP or to this machine's own name.- Requests that omit sampling parameters keep the old server's defaults
(temperature 0, repeat penalty 1.05, and top-p 0.9 / top-k 40 unless the
model publishes its own truncation); an explicit value always wins.
presence_penaltyandfrequency_penaltyoutside [-2, 2] are refused, as
before. - A model cache without
hub.jsonis served under its repository name, and
a second variant of the same model asname:variant. - Text that looks like a chat marker —
<|im_end|>,<|im_start|>,
<think>— inside a message, a tool result or a document is content, never
framing: a quoted marker can no longer end a turn, forge a system turn, or
put the rest of a session into reasoning. - A tokenizer-only handle opens the bundle's metadata only. On NVIDIA that is
tens of gigabytes less resident for a server that keeps several.
Fixed
- Gemma 4 26B-A4B misread every image it was shown, on Mac and NVIDIA: a step
of its vision encoder was skipped. On NVIDIA its per-head normalization also
ran on a partial warp. Gemma 4 E2B and E4B were not affected. - Hybrid-model prefix reuse now survives a replayed or branched conversation.
A conversation replayed after the snapshot store had cycled, or a second
turn branched off a shared first turn, read its whole prompt again, so
identical loads ran two to three times slower than distinct ones. - Hybrid models serve at one lane (
--max-batch 1), as the Local app starts
large ones; the recurrent state the single request needs is now reserved. - The unfused attention path for 512-wide heads with Q8 KV on Metal misread
part of its input on every odd-length prompt. - GLM-DSA served a stale key row on every single-sequence paged step.
- Llama-family models produced garbage in batch-invariant mode
(BASERT_BATCH_INVARIANT=1); Gemma 4 and Qwen 3.5 now use the same kernels
in that mode whether a request runs alone or in a batch. - A LoRA adapter on Qwen 3.5's shared-expert gate is applied during
generation on NVIDIA, not only while reading the prompt. - Two mixture-of-experts paths on NVIDIA no longer misread weights when a
model mixes scale formats or keeps routing weights in full precision. - The GPU sampler treated
top_k: 0as greedy and applied top-p and min-p
without the requested temperature. - A small GPU buffer is freed when a model is unloaded.
BaseRT engine 0.2.6
BaseRT 0.2.6
Fixes garbage output from Qwen3.x 27B and 35B-A3B models on M1-family
Macs. These models answered their first token correctly and then produced
repeated nonsense on M1, M1 Pro, M1 Max and M1 Ultra, in 0.2.4 and 0.2.5.
Other chips and models were not affected.
Changed
- A GPU dispatch that the hardware would silently skip now stops with an
error naming the kernel, instead of producing wrong output.
BaseRT engine 0.2.5
BaseRT 0.2.5
Tool calls stream as they are written, and convert-on-pull no longer fails
after a complete download. An agent UI used to sit blank while a tool call
was generated, then receive the whole thing at once; the call now arrives the
way OpenAI clients reassemble it. And a convert-on-pull of a repo whose
weights are 32 MB or larger — which is all of them — failed at the end with
no .safetensors shards despite having downloaded every byte. (basert pull
still accepts only supported architectures backed by safetensors.)
Scope. 0.2.5 is a patch release cut from
main: streaming tool calls,
the convert-on-pull fix, and benchmark-harness telemetry.
Added
Streaming tool calls arrive in parts
/v1/chat/completions with stream: true now delivers a tool call as an
opening delta.tool_calls entry carrying index, id and function.name,
then further entries on the same index carrying only function.arguments
fragments, then finish_reason: "tool_calls".
{"index":0,"id":"call_…","type":"function","function":{"name":"get_weather","arguments":""}}
{"index":0,"function":{"arguments":"{\"city\":\"Paris\""}}
{"index":0,"function":{"arguments":",\"unit\":\"celsius\""}}
{"index":0,"function":{"arguments":"}"}}
The function name comes first and cannot be streamed before it is known — the
server learns it by parsing the call body — so the opening entry waits for it
and everything after that streams. How much streams depends on how the model
writes calls: a JSON body streams its raw arguments bytes, and an XML-dialect
body (Qwen3.5/3.6, Qwen3-Coder) streams one argument per parameter, as each
parameter's value is settled. Gemma's native call:fn{k:v} form and the ATEM
dialect are not JSON and still arrive whole when the call closes.
A call cut short by the token budget keeps finish_reason: "length" — including
when earlier calls in the same response completed, so a client never sees
"tool_calls" alongside a half-written call it cannot run. The
fragments already sent stand, but nothing is invented to close them.
Benchmark harness: telemetry from the measured run
basert-benchmark-harness now reports power, energy, temperature and memory
observed during the timed repetitions themselves rather than from a separate
pass, and --headline-first runs PP512 and TG128 first in one freshly loaded
model with a fixed 4,096-token reservation, before the larger prefill sweep —
so the headline numbers are not read off a machine already heated by them. The
report records the order and capacities it actually used. basert-bench is
unchanged, as are ordinary harness invocations.
Also new: --isolated-workloads (each workload in its own process at its own
minimum context capacity), --workload {all|prefill|decode|headline} and
--idle-baseline. -c/--ctx is rejected with --headline-first and with
--isolated-workloads, since those profiles choose the capacity themselves.
Fixed
basert pull <org/model>from a Hugging Face repo failed with
no .safetensors shardsafter downloading the whole model. The resumable
downloader used for files of 32 MB or more left the weights out of the
snapshot the converter reads. A pull that already downloaded the model
resumes and converts without fetching it again.- Energy figures from the benchmark harness counted the probes around a
workload, not just the workload: on CUDA and ROCm the vendor telemetry
commands landed inside the measured interval while average power divided by
the shorter one. Both counters are now read next to the timed repetitions.
Changed
- Benchmark report schema (breaking for report consumers). Same-run
telemetry isbasert-telemetry/4; the 0.2.4/2and/3sections are
gone, and the harness now fails a run whose workload does not carry/4
rather than degrading. The new throughput profiles add
basert-throughput-protocol/1,/2andbasert-bench-capacity/1. Anything
parsing a 0.2.4 harness report needs updating; numbers from the new
--headline-first/--isolated-workloadsprofiles are not comparable to
shared-context 0.2.4 numbers, and--telemetrynow samples inside the timed
window, so its throughput is not comparable to 0.2.4--telemetry
throughput either. - The benchmark harness executable is rebuilt whenever it is missing, rather
than only when its sources change. A packaged build could otherwise omit it.
BaseRT engine 0.2.4
BaseRT 0.2.4
Three new model families, and a context window that sizes itself. gpt-oss,
Nemotron 3 Nano and GLM 5.2 all run now, with continuous batching rather than
one request at a time. And basert serve no longer starts every model at 4096
tokens — it measures the machine and gives the model the window the hardware
can hold, up to whatever it was trained for.
Added
The context window sizes itself
--max-context defaulted to 4096, which threw away most of a 32k model on a
workstation and overran the budget on a laptop when raised by hand. It is now
derived from the GPU memory available, the weights, and what a token of KV
cache costs — capped only by the model's trained window.
Context window: 131072 tokens (auto-sized for 96 GB of GPU memory at 8 lanes).
Pass --max-context N to pin a different one.
--max-tokens follows: unset, a request may fill the rest of the window
instead of stopping at 2048. basert chat and basert complete size
themselves the same way.
On a Mac sharing its GPU with another program, pin --max-context — the
measurement can't see another process's allocation.
gpt-oss
basecompute/gpt-oss-20b in bf16, Q8 and Q4. The MXFP4 experts are carried
across verbatim, never requantized. Harmony tool calls are parsed and emitted
as tool calls.
Nemotron 3 Nano
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B in bf16, Q8 and Q4 — a Mamba-2
hybrid, where most layers carry recurrent state instead of a KV cache.
GLM 5.2
Multi-head latent attention, with the sparse-indexer path for long contexts.
Continuous batching for it is Metal-only in this release; on CUDA it runs one
sequence at a time.
Fixed
- gpt-oss returned NaN on roughly one prompt position in twenty.
- Large models could emit a stray end-of-turn token before answering, on Metal
under continuous batching. - A loaded LoRA adapter ignored repetition, presence and frequency penalties,
logit_bias, andseed. - gpt-oss could keep generating past a finished tool call.
- A failed conversion left a model-sized temporary file behind.
- Image requests overreported prompt token usage when prompt caching was on.
--validatecould not pass on a current MLX checkpoint.
Changed
BaseRTModelConfigchanged layout — recompile anything built against
0.2.3 that reads it.- The shared library reports its real version again, and the macOS archive
now ships the symlink its install name needs. Linking against the library
from 0.2.1–0.2.3 failed to load; the bundled tools were unaffected.
BaseRT engine 0.2.3
BaseRT 0.2.3
Model pulls that finish. basert pull downloads over many connections at
once and resumes where it stopped — a 20 GB bundle interrupted at 90% picks
up at 90%, including after the process has exited. The 19.8 GB
Qwen3.5-35B-A3B Q4 bundle measured 17.6 MiB/s before and 54–84 MiB/s
after on the same link, roughly 19 minutes down to 5. A transfer that stalls
without closing is cut off and retried instead of hanging, and a short transfer
is caught rather than installed. Tool calling is fixed on two fronts as well: Muse Glimmer's
dialect is parsed rather than printed, and templates that use is undefined
(Qwen 3.8 among them) render instead of silently falling back.
Added
Xet transfers
-
Xet is supported, behind
BASERT_HF_XET=1. Its advantage is
deduplication across models: chunks shared with something already fetched
never cross the wire. It cannot resume, though — an interrupted transfer
starts over — so it is not whatbasert pullreaches for by default.basert pull basecompute/Qwen3.5-35B-A3B # resumable (default) BASERT_HF_XET=1 basert pull basecompute/Qwen3.5-35B-A3B # dedup, no resume
Parallel, resumable downloads
-
Files are fetched over many connections at once, in fixed chunks,
instead of through one stream. -
An interrupted pull resumes where it stopped, per chunk: completion is
recorded next to the partial file as each chunk lands and fsynced, so a
crash, a^C, or a link that drops overnight continues from what is on disk,
across process restarts. Just re-run the samebasert pull. -
A stalled transfer fails instead of hanging. Connections that stop
delivering bytes without closing are now cut off and retried, and a transfer
shorter than the size the Hub advertised is reported as an error rather than
installed as a truncated model that fails later at load time. -
Tunable when the defaults do not suit a link:
Variable Default Meaning BASERT_HF_CONNECTIONS24 Concurrent range requests BASERT_HF_CHUNK_MB16 Bytes per range request BASERT_HF_READ_TIMEOUT_SECS60 Stall timeout for a single read BASERT_HF_MAX_RETRIES5 Retries per chunk ( 0disables)BASERT_HF_XETunset Use Xet's CAS path for large files (dedup, no resume) Peak memory is
connections × chunk— 384 MB at the defaults, since each
worker holds one chunk in flight. Lower either knob on a constrained
machine.
Fixed
-
Muse Glimmer's tool calls are parsed, not printed. The model frames a
call as a message addressed to the tool:<|start|>assistant to=fn<|message|><atem:function_calls> <atem:invoke name="fn">…</atem:invoke></atem:function_calls><|eot|>which matched none of the dialects the parser knew (Gemma-native, ChatML
JSON, Qwen XML), so it fell through to content and the raw markup was
streamed to the user — the reported symptom was a reply that read
<atem:invoke name="bash">. The dialect is now recognized and emitted as a
proper tool call. -
--kquant-passthroughembeddings on non-Glimmer models. Every model but
Muse Glimmer handed k-quant embedding lookups the slab parameter struct,
whose second field is the vocabulary size where the k-quant kernels expect
the row width. Rows were read at the wrong stride, and the output was
incoherent — Qwen3-0.6B was the reported case. -
developerandfunctionroles no longer break the native template.
A model's chat template accepts the roles it was written for — Qwen 3.8's
takessystem,user,assistant,tooland raises on anything else — and
a raise drops the whole request into the generic ChatML fallback, quietly
costing the model its think prefill and tool-call dialect.developer(which
supersededsystemin the newer OpenAI API, and is what agent clients send)
andfunction(the deprecated spelling oftool) are now resolved before
rendering. When a template does reject a role, the warning names the roles
the request carried instead of suggesting a re-convert that would not help. -
Chat templates using
is undefinedno longer fall back to ChatML.
Qwen 3.8's template gates its reasoning-effort block on
enable_thinking is undefined, a standard Jinja2 test the bundled template
engine did not implement. Rendering threw on every request and the server
quietly fell back to generic ChatML — which still looked right for
ChatML-shaped models while silently dropping the reasoning-effort system
block, the think prefill, and the model's entire tool-call dialect. Tool
calling on affected models works again.
Changed
- Building
tools/base-convertnow requires Rust 1.85+ (was 1.80).
BaseRT engine 0.2.2
BaseRT 0.2.2
Muse Glimmer, and native GGUF k-quants. BaseRT now runs Muse Glimmer 30B —
text and vision — on Apple Silicon, and reads GGUF k-quantized weights
(Q4_K / Q5_K / Q6_K) directly, without a dequantize-and-repack step. Short
prompts prefill 1.4–1.55× faster on M5.
Added
Muse Glimmer
-
Muse Glimmer 30B, text and vision, including the perception tower — so
the model answers about images, not just text. -
Pre-converted bundles at
basecompute/Muse-Glimmer-30B:basert pull basecompute/Muse-Glimmer-30B # 21 GB, recommended basert pull basecompute/Muse-Glimmer-30B:q4k-17gb # 18 GB, smaller
Both are the model's own published k-quants, repackaged — the super-block
bytes are copied through unchanged, so the weights are bit-identical to the
upstream GGUFs rather than a re-quantization of them. Both carry the vision
tower; there is no separate projector file to manage. -
Images use the model's own
<|patch|>placeholder:basert complete basecompute/Muse-Glimmer-30B --chat \ --image photo.png --prompt "<|patch|>Describe this image."
Native GGUF k-quants
Q4_K/Q5_K/Q6_Krun natively — prefill, decode and embedding
lookup all read the quantized weights directly.Q2_KandQ3_Kare
supported as well.--kquant-passthroughcopies GGUF super-blocks into a.basebundle
verbatim, so the weights stay bit-identical to the source. Re-quantizing an
already-quantized model compounds its error; copying the blocks through
introduces none, so this is not a quality trade and needs no override flag.--mmprojfolds a companionmmproj-*.ggufperception tower into the
same bundle. Without it a bundle converted from GGUF is text-only, because
the tower ships as a separate file. A bundle built this way is
interchangeable with one converted from the original safetensors.
Changed
- Faster short-prompt prefill on M5 — 1.4–1.55× measured on M5 Max. Short
prompts previously fell back to a slower path than the hardware supports;
they now use the same accelerated one longer prompts already did. Long
prefill is unchanged, because it was already there. - Faster k-quant prefill — 3.4×, from running k-quantized weights on the
accelerated path rather than the general one. - Wider accelerated coverage for the 5-, 2- and 3-bit quantizations,
alongside the 4-, 6- and 8-bit ones. - Decode speed is unchanged. At these model shapes it is limited by memory
bandwidth rather than compute, so the changes above do not move it, and the
decode-side ideas tried for this release measured slower and were dropped.
Fixed
- M3 Ultra was identified as a base M3, so it was tuned for a far smaller
chip than it is. Reachable only when the system does not report a GPU core
count, but it would have affected every future per-chip decision. - Language bindings could corrupt memory. The model config is returned by
value, and the Python, Rust and Node definitions of that struct had fallen
behind the C header — a mirror even four bytes short lets the engine write
past the caller's allocation. All four bindings are back in step, and a test
now holds them there. - Models converted from GGUF ignored the end-of-turn token and kept
generating past the end of a reply. baseRT_format_chatis available without a chat-capable build —
it no longer depends on the chat layer being present.
Notes
- k-quant support is Apple Silicon (Metal) only in this release.
- Folding a perception tower from a companion
mmprojGGUF is exact for still
images, which is what these models take today. Such a file does not carry
everything a video-frame path would need.
BaseRT engine 0.2.1
BaseRT 0.2.1
Speech-to-text. BaseRT now runs the full OpenAI Whisper family — all twelve
variants from tiny to large-v3-turbo, including the English-only builds — as
.base bundles on both Apple Silicon (Metal) and NVIDIA GB10 (CUDA), served
through an OpenAI-compatible transcription API with live streaming.
Added
Whisper speech-to-text
- All twelve Whisper variants convert to
.baseand run end-to-end:
tiny,base,small,medium,large,large-v2,large-v3,
large-v3-turbo, and the.enEnglish-only builds. Special-token IDs are
read from each model's own tokenizer rather than inferred from vocabulary
size, solarge-v3andlarge-v3-turbo(100 languages, 128 mel bins) map
correctly. - Quantized bundles —
whisper-q8andwhisper-q4profiles quantize the
linear projections while the convolutional front end, positional and token
embeddings, norms, and biases stay f16. Q8 meets the same word-error-rate
parity gates as f16 against the reference implementation. - CUDA support — Whisper runs on NVIDIA GB10 with transcripts byte-identical
to the Metal path. POST /v1/audio/transcriptionsandPOST /v1/audio/translations,
OpenAI-compatible, includingverbose_jsonwith per-segment
avg_logprob,no_speech_prob,compression_ratioandtemperature, plus
the detected language.- Streaming transcription —
stream: trueemits segments from the decode
loop as each window completes, in the documentedtranscript.text.delta/
transcript.text.donewire format, so official OpenAI SDKs parse it directly.
Time-to-first-event is one window's decode rather than the whole file. - Decode parity with the reference implementation —
task=translate,
language auto-detection (language: "auto"),initial_promptwith
condition-on-previous-text, and the temperature-fallback ladder.
Word-level timestamps remain future work.
C API
baseRT_load_model_exwithBaseRTLoadOptions— per-load KV width, paged-KV,
batch size, prefix cache, prefill chunk and paged-weight mode, replacing the
process-wide setters. The options struct is ABI-guarded bystruct_size.- Transcription accessors:
baseRT_set_task,baseRT_set_initial_prompt,
baseRT_set_condition_on_previous_text, per-segment results, detected
language, and source-audio duration. - Batched logits:
baseRT_batch_logits_stride,baseRT_argmax_logits_row,
baseRT_sample_logits_row,baseRT_logits_row_logprobs,
baseRT_mask_logits_row. baseRT_capabilitiesreports what the running backend supports.
All of the above are bound in the Python, Node, Swift and Rust bindings.
Changed
- Faster shared-prefix serving — concurrent requests that share a prompt
prefix now detect and reuse it through a cascade dispatch path. - Faster MoE decode on CUDA — fused SiLU gate/up for 4-bit experts.
- Metal GEMM accuracy — f32 accumulation on the paths where f16
accumulation was measurably lossy.
Fixed
context_length_exceededreported a capped token count. The number came
from a fixed tokenizer buffer, so any over-long prompt reported the same
figure no matter how far over it was — a client could not trim to a number
that was itself the ceiling. It now reports the real count.- Image requests under-reported usage and never returned
finish_reason: "length": the first token produced during image prefill was
streamed to the client but not counted, so a truncated answer looked complete. - Gemma 4 audio produced unusable embeddings. The channel-last convolution
indexed its weights as[out, kh, kw, in]where the bundle stores
[out, in, kh, kw]. For a single input channel the two are identical, so the
first subsampling convolution was correct and masked the fault while the
second convolved with transposed weights. - Repetition penalty silently did nothing on some Apple GPUs — a sampling
kernel read a scratch buffer that was never bound, which is undefined
behaviour and therefore hardware-dependent. - A KV-cache shift kernel had a data race — an in-place shift whose source
and destination overlap, spread across threads, so results depended on
scheduling order. - Rate limiting could grow without bound. Idle buckets expired on time
alone, so a source rotating addresses added an entry per request and nothing
became evictable within the window. Bucket count is now capped. - Absolute
--media-dirpaths are honoured as written and must resolve
inside the configured root, instead of being re-anchored under it — which
made a legitimate absolute path fail as an image-decoding error. - Rejections now appear in the access log, so an operator can tell a
refused request from a hung one. - Second images and audio parts in a chat request are rejected with a clear
400 instead of a 500 or, for audio, a confident answer about audio the model
never received. - Image-decoding failures no longer echo the server's resolved filesystem path
back to the caller. basert inspectreported0tensors for working vision and audio towers: it
counted name prefixes that converted bundles do not use.- Chat-template cache eviction and a Qwen3 tool round-trip fault.
Notes
BASERT_VERSIONnow advertises the patch component, and the release pipeline
verifies the tag, the compiled-in version, and every package manifest agree
before publishing.
BaseRT engine 0.2.0
BaseRT 0.2.0
A second hardware backend. BaseRT now runs on NVIDIA GB10 (DGX Spark, Linux/arm64) alongside Apple Silicon — the same .base bundles, the same basert CLI and OpenAI-compatible server, now on CUDA. This release also brings the Qwen3.5 / Qwen3.6 hybrid model family and a continuous-batching overhaul of basert serve.
Added
- CUDA backend — NVIDIA GB10 / DGX Spark — full inference on Linux/arm64 + CUDA: dense, MoE, and hybrid (Gated-DeltaNet) architectures, across prefill, single-stream decode, and continuous-batching serving.
install.shauto-detects the platform (macOS/arm64 → Metal, Linux/arm64 → CUDA) andbasert pullresolves the right bundle for the host. - Qwen3.5 & Qwen3.6 hybrid models — support for the Gated-DeltaNet hybrid architecture (dense and MoE), including vision (Qwen3.5-VL), on both Metal and CUDA. GB10 continuous batching for the dense hybrids.
- Continuous-batching serving —
basert servenow drives tool calls (streaming + non-streaming), grammar-constrained decode, per-token logprobs, andn>1choices through the continuous-batching engine, with a decode-priority scheduler and radix prefix-cache reuse (dense and hybrid, via a keyed GDN snapshot store). - CUDA-native model bundles —
cuda-q8andcuda-q4mixvariants for 11 models (Qwen3-0.6B/30B-A3B, Qwen3.5-2B/35B-A3B, Qwen3.6-27B/35B, Llama-3.2-1B/3B, Gemma-3-1B, Gemma-4-E2B/26B), on Hugging Face and the catalog. - Faster Metal GEMMs — large-tile prefill GEMM kernels for native bf16 / f16 weights and Q8, plus batched Gated-DeltaNet decode/prefill kernels for the M1–M5 families.
- Tokenizer conformance — an HF-exact conformance harness and a Unicode-category pretokenizer; Phi-3-mini un-quarantined after revalidation.
Fixed
- Chunked-decode correctness on CUDA — free generation past the decode chunk size no longer repeats an earlier block (an eager-dispatch write-after-read hazard); prefill and teacher-forced decode were always correct.
- Large quantized tensors — weight tensors with more than ~512M elements no longer decode to garbage (64-bit byte-offset fix).
- Gemma normalization precision — the canonical post-attention / feed-forward norm weights are kept in floating point in the shipped bundles (were quantized in some profiles).
- Hybrid prefix reuse — correct KV/GDN dtype handling and prefix-cache reuse for hybrid-GDN models under serving.
- MoE quantization parity — Qwen3.5 MoE quant rules matched to canonical tensor names so experts quantize as intended.
Changed
- Backend-aware model resolution — CUDA hosts prefer CUDA-native bundles, Apple Silicon prefers Metal, with a universal fallback; catalog entries are backend-tagged and can be derived from a
.baseheader. - Performance — Metal decode/prefill tuning across the M1 Max / M4 Pro / M5 Pro families (large-tile GEMM, higher-occupancy Gemma-4 MoE gate/up). Benchmarks vs llama.cpp and vLLM on GB10 and Apple Silicon are in the repo.
Install
# macOS (Apple Silicon, Metal) or Linux (arm64, CUDA) — auto-detected
curl -fsSL https://raw.githubusercontent.com/basecompute/baseRT/main/install.sh | sh
basert pull <model> # resolves the right bundle for your host
basert serve <model>.base # OpenAI-compatible serverBaseRT engine 0.1.7
Maintenance release focused on the three most-reported serve/CLI issues, plus converter and chat quality-of-life fixes.
Fixed
- Tool calls for Qwen 3.5 / 3.6 (#21) —
basert servenow emits structuredtool_callsfor the Qwen XML tool-call dialect, in both streaming and non-streaming modes. Truncated generations never execute partial arguments, and logprobs stay aligned across tool-call segments. - Deterministic HTTP 500 on some short CJK prompts (#22) — fixed a byte-alphabet hole (soft hyphen
0xAD) in GPT-2-style byte tokenization; tokenization is now HF-exact for the affected models. Server errors are structured and logged instead of opaque 500s. --no-thinkand chat-template handling (#23) —--no-thinknow works across all thinking-capable architectures; the CLI renders model-embedded chat templates; fixed a Gemma 4 channel-marker leak and KV-reuse correctness on divergent history.- MLX 3/5/6-bit conversion — MLX quantized checkpoints at 3/5/6-bit widths now unpack correctly (little-endian bitstream); previously these converted to wrong weights.
basert pullHF-cache duplication — pulled models are no longer duplicated in the global Hugging Face cache (staged download with resume preserved; setBASERT_KEEP_HF_SOURCES=1to keep sources).
Added
- Chat line editing — cursor movement, kill ops, and word delete (
Ctrl+W/Alt+Backspace/Ctrl+Backspace) inbasert chat, with UTF-8/CJK support. - q6 MoE kernels — Q6 expert GEMM (plus q8/q6 scale-dtype siblings) and a bundled
default-q6conversion profile (--target base-q6). baseRT_decode_token_raw()— length-preserving raw-byte token decode in the C API (byte-level BPE tokens can legally contain NUL; the C-string variants truncate there). No struct-layout changes.
Changed
- Chat / complete / bench banners show the model id (
org/repo · variant) instead ofmodel.base. baserthelp output regrouped runtime-first.
Performance is unchanged vs 0.1.6 (validated dense + MoE on M4 Pro and M5 Pro, within noise).
BaseRT engine 0.1.6
BaseRT 0.1.6
- Qwen 3.5 / 3.6 support: dense, MoE, and vision-language models (hybrid Gated-DeltaNet architecture), in both the runtime and
basert convert. - M5 tensor-core kernels included: release binaries now ship the Metal 4 kernel library (
baseRT_tensor.metallib) — earlier releases were missing it, so M5 GPUs fell back to slower kernels. Loads automatically; older Macs are unaffected. - Big-model support: models larger than the GPU wiring limit now load and decode at full speed (e.g. a 28 GB model on a 48 GB Mac) instead of failing.
- Qwen3.5-VL accuracy: image preprocessing and multimodal RoPE now match the reference implementation token-for-token.
- CLI:
basert-completeaccepts-nfor max tokens; unknown flags are now errors. - Requires macOS 15+ (Metal 4 kernels engage on macOS 26 / M5).
Contents: libbaseRT.dylib (embedded kernels) + basert CLI + basert-* runtime tools + headers.