Skip to content

Releases: basecompute/baseRT

BaseRT engine 0.3.0

Choose a tag to compare

@prabod prabod released this 08 Oct 14:02
128bdd7

BaseRT 0.3.0

basert serve is a new server. It keeps the command line and the
OpenAI-compatible API it has always had, and runs them on
superfluid, our serving daemon:
requests are decoded together through one shared forward pass,
a prompt a conversation repeats is reused instead of read again, and hybrid
models (Qwen 3.5, 3.6 and 3.8) get context windows up to four times larger
from the same memory. This release also carries the fixes prepared for 0.2.7,
which was never published.

New

  • Serving. Continuous batching and prefix reuse are always on; the server
    runs one request at a time unless --continuous-batching [N] or
    --max-batch-size N asks for more. A server is ready sooner (Qwen3-8B on an
    M5 Pro: 5.3 s to 1.0 s). A streamed tool call arrives as partial deltas —
    the function's name first, then its arguments as they are produced — and
    tool_choice naming a function, or required, is honoured. /v1/completions
    accepts suffix for fill-in-the-middle. /v1/models lists each model's
    capabilities, including whether its template takes a thinking switch and
    which reasoning-effort levels it defines.
  • Larger windows for hybrid models. The paged KV cache charges only the
    attention layers, so on a 48 GB M5 Pro the automatically sized window grows
    from 72,704 to 262,144 tokens on Qwen3.6-35B-A3B and Qwen3.5-35B-A3B, from
    31,744 to 130,048 on Qwen3.6-27B, and from 27,648 to 111,616 on Qwen3.8-27B.
    Dense models keep the window they had.
  • GLM 5.2 on Metal serves with continuous batching, and speculates with
    the multi-token-prediction head its bundle carries.
  • Speculation. Block drafters with a reduced draft vocabulary (the DFlash
    heads published for Qwen3-30B-A3B and Nemotron-3-Super) convert and serve,
    token-for-token identical to plain decoding under greedy sampling. A sampled
    round verifies all of its rows on the GPU in one pass (an 8-lane round on
    Qwen3.8-27B: 109 ms to 93 ms). A strategy that loses to plain decoding is
    measured against a real plain run and retired, then rechecked later rather
    than written off. basert complete --spec turns speculation on for a bundle
    with its own head.
  • Throughput under agent load. Sampled lanes no longer sort the whole
    vocabulary per token, served lanes sample on the GPU, and prompts are read
    prefill-first. Eight agents across six turns on Qwen3.6-35B-A3B (M5 Max):
    wall time 345 s to 136 s, first token at the median 46 s to 6 s.
  • Automatic context sizing on Metal plans within the memory the engine can
    keep pinned, and counts the recurrent state a hybrid model allocates per
    lane, so a derived window no longer pushes the weights out during a long
    prompt.

Changed

  • Flags the old server took but this one has no use for are accepted and
    noted on stderr: --paged-kv and --prefix-cache (always on),
    --prefix-cache-file and --prefix-cache-save-interval (the cache is not
    kept on disk), --files-dir and --media-dir (files live in the server's
    state directory; images and audio arrive as data: URIs or uploads), and
    the engine settings --metallib, --prefill-chunk, --paged-weights,
    --no-paged-weights-retry, --gpu-wait-timeout-ms, --decode-replay and
    --no-baked-decode. --request-timeout bounds fill-in-the-middle
    completions only. A flag value that is not a number is refused rather than
    read as 0. --model-dir alone loads its first model at startup. --verbose
    prints the load and GPU diagnostics without the per-kernel dispatch list.
  • State — the request log and uploaded files — lives under $BASERT_STATE_DIR
    (default: serve/ beside the model cache), one directory per port, or
    wherever --sessions <dir> says; the request log starts fresh at each start.
  • /props no longer reports trained_context; /v1/models reports the
    serving window as meta.n_ctx; replies say system_fingerprint: superfluid;
    /metrics series are named superfluid_*. Error messages are worded
    differently. Without --api-key on a non-loopback address, the server
    answers only requests addressed to an IP or to this machine's own name.
  • Requests that omit sampling parameters keep the old server's defaults
    (temperature 0, repeat penalty 1.05, and top-p 0.9 / top-k 40 unless the
    model publishes its own truncation); an explicit value always wins.
    presence_penalty and frequency_penalty outside [-2, 2] are refused, as
    before.
  • A model cache without hub.json is served under its repository name, and
    a second variant of the same model as name:variant.
  • Text that looks like a chat marker — <|im_end|>, <|im_start|>,
    <think> — inside a message, a tool result or a document is content, never
    framing: a quoted marker can no longer end a turn, forge a system turn, or
    put the rest of a session into reasoning.
  • A tokenizer-only handle opens the bundle's metadata only. On NVIDIA that is
    tens of gigabytes less resident for a server that keeps several.

Fixed

  • Gemma 4 26B-A4B misread every image it was shown, on Mac and NVIDIA: a step
    of its vision encoder was skipped. On NVIDIA its per-head normalization also
    ran on a partial warp. Gemma 4 E2B and E4B were not affected.
  • Hybrid-model prefix reuse now survives a replayed or branched conversation.
    A conversation replayed after the snapshot store had cycled, or a second
    turn branched off a shared first turn, read its whole prompt again, so
    identical loads ran two to three times slower than distinct ones.
  • Hybrid models serve at one lane (--max-batch 1), as the Local app starts
    large ones; the recurrent state the single request needs is now reserved.
  • The unfused attention path for 512-wide heads with Q8 KV on Metal misread
    part of its input on every odd-length prompt.
  • GLM-DSA served a stale key row on every single-sequence paged step.
  • Llama-family models produced garbage in batch-invariant mode
    (BASERT_BATCH_INVARIANT=1); Gemma 4 and Qwen 3.5 now use the same kernels
    in that mode whether a request runs alone or in a batch.
  • A LoRA adapter on Qwen 3.5's shared-expert gate is applied during
    generation on NVIDIA, not only while reading the prompt.
  • Two mixture-of-experts paths on NVIDIA no longer misread weights when a
    model mixes scale formats or keeps routing weights in full precision.
  • The GPU sampler treated top_k: 0 as greedy and applied top-p and min-p
    without the requested temperature.
  • A small GPU buffer is freed when a model is unloaded.

BaseRT engine 0.2.6

Choose a tag to compare

@prabod prabod released this 18 Sep 01:51
3b31090

BaseRT 0.2.6

Fixes garbage output from Qwen3.x 27B and 35B-A3B models on M1-family
Macs.
These models answered their first token correctly and then produced
repeated nonsense on M1, M1 Pro, M1 Max and M1 Ultra, in 0.2.4 and 0.2.5.
Other chips and models were not affected.

Changed

  • A GPU dispatch that the hardware would silently skip now stops with an
    error naming the kernel, instead of producing wrong output.

BaseRT engine 0.2.5

Choose a tag to compare

@prabod prabod released this 17 Sep 05:21
3b31090

BaseRT 0.2.5

Tool calls stream as they are written, and convert-on-pull no longer fails
after a complete download.
An agent UI used to sit blank while a tool call
was generated, then receive the whole thing at once; the call now arrives the
way OpenAI clients reassemble it. And a convert-on-pull of a repo whose
weights are 32 MB or larger — which is all of them — failed at the end with
no .safetensors shards despite having downloaded every byte. (basert pull
still accepts only supported architectures backed by safetensors.)

Scope. 0.2.5 is a patch release cut from main: streaming tool calls,
the convert-on-pull fix, and benchmark-harness telemetry.

Added

Streaming tool calls arrive in parts

/v1/chat/completions with stream: true now delivers a tool call as an
opening delta.tool_calls entry carrying index, id and function.name,
then further entries on the same index carrying only function.arguments
fragments, then finish_reason: "tool_calls".

{"index":0,"id":"call_…","type":"function","function":{"name":"get_weather","arguments":""}}
{"index":0,"function":{"arguments":"{\"city\":\"Paris\""}}
{"index":0,"function":{"arguments":",\"unit\":\"celsius\""}}
{"index":0,"function":{"arguments":"}"}}

The function name comes first and cannot be streamed before it is known — the
server learns it by parsing the call body — so the opening entry waits for it
and everything after that streams. How much streams depends on how the model
writes calls: a JSON body streams its raw arguments bytes, and an XML-dialect
body (Qwen3.5/3.6, Qwen3-Coder) streams one argument per parameter, as each
parameter's value is settled. Gemma's native call:fn{k:v} form and the ATEM
dialect are not JSON and still arrive whole when the call closes.

A call cut short by the token budget keeps finish_reason: "length" — including
when earlier calls in the same response completed, so a client never sees
"tool_calls" alongside a half-written call it cannot run. The
fragments already sent stand, but nothing is invented to close them.

Benchmark harness: telemetry from the measured run

basert-benchmark-harness now reports power, energy, temperature and memory
observed during the timed repetitions themselves rather than from a separate
pass, and --headline-first runs PP512 and TG128 first in one freshly loaded
model with a fixed 4,096-token reservation, before the larger prefill sweep —
so the headline numbers are not read off a machine already heated by them. The
report records the order and capacities it actually used. basert-bench is
unchanged, as are ordinary harness invocations.

Also new: --isolated-workloads (each workload in its own process at its own
minimum context capacity), --workload {all|prefill|decode|headline} and
--idle-baseline. -c/--ctx is rejected with --headline-first and with
--isolated-workloads, since those profiles choose the capacity themselves.

Fixed

  • basert pull <org/model> from a Hugging Face repo failed with
    no .safetensors shards after downloading the whole model. The resumable
    downloader used for files of 32 MB or more left the weights out of the
    snapshot the converter reads. A pull that already downloaded the model
    resumes and converts without fetching it again.
  • Energy figures from the benchmark harness counted the probes around a
    workload, not just the workload: on CUDA and ROCm the vendor telemetry
    commands landed inside the measured interval while average power divided by
    the shorter one. Both counters are now read next to the timed repetitions.

Changed

  • Benchmark report schema (breaking for report consumers). Same-run
    telemetry is basert-telemetry/4; the 0.2.4 /2 and /3 sections are
    gone, and the harness now fails a run whose workload does not carry /4
    rather than degrading. The new throughput profiles add
    basert-throughput-protocol/1, /2 and basert-bench-capacity/1. Anything
    parsing a 0.2.4 harness report needs updating; numbers from the new
    --headline-first / --isolated-workloads profiles are not comparable to
    shared-context 0.2.4 numbers, and --telemetry now samples inside the timed
    window, so its throughput is not comparable to 0.2.4 --telemetry
    throughput either.
  • The benchmark harness executable is rebuilt whenever it is missing, rather
    than only when its sources change. A packaged build could otherwise omit it.

BaseRT engine 0.2.4

Choose a tag to compare

@prabod prabod released this 09 Sep 13:30
0644388

BaseRT 0.2.4

Three new model families, and a context window that sizes itself. gpt-oss,
Nemotron 3 Nano and GLM 5.2 all run now, with continuous batching rather than
one request at a time. And basert serve no longer starts every model at 4096
tokens — it measures the machine and gives the model the window the hardware
can hold, up to whatever it was trained for.

Added

The context window sizes itself

--max-context defaulted to 4096, which threw away most of a 32k model on a
workstation and overran the budget on a laptop when raised by hand. It is now
derived from the GPU memory available, the weights, and what a token of KV
cache costs — capped only by the model's trained window.

Context window: 131072 tokens (auto-sized for 96 GB of GPU memory at 8 lanes).
Pass --max-context N to pin a different one.

--max-tokens follows: unset, a request may fill the rest of the window
instead of stopping at 2048. basert chat and basert complete size
themselves the same way.

On a Mac sharing its GPU with another program, pin --max-context — the
measurement can't see another process's allocation.

gpt-oss

basecompute/gpt-oss-20b in bf16, Q8 and Q4. The MXFP4 experts are carried
across verbatim, never requantized. Harmony tool calls are parsed and emitted
as tool calls.

Nemotron 3 Nano

basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B in bf16, Q8 and Q4 — a Mamba-2
hybrid, where most layers carry recurrent state instead of a KV cache.

GLM 5.2

Multi-head latent attention, with the sparse-indexer path for long contexts.
Continuous batching for it is Metal-only in this release; on CUDA it runs one
sequence at a time.

Fixed

  • gpt-oss returned NaN on roughly one prompt position in twenty.
  • Large models could emit a stray end-of-turn token before answering, on Metal
    under continuous batching.
  • A loaded LoRA adapter ignored repetition, presence and frequency penalties,
    logit_bias, and seed.
  • gpt-oss could keep generating past a finished tool call.
  • A failed conversion left a model-sized temporary file behind.
  • Image requests overreported prompt token usage when prompt caching was on.
  • --validate could not pass on a current MLX checkpoint.

Changed

  • BaseRTModelConfig changed layout — recompile anything built against
    0.2.3 that reads it.
  • The shared library reports its real version again, and the macOS archive
    now ships the symlink its install name needs. Linking against the library
    from 0.2.1–0.2.3 failed to load; the bundled tools were unaffected.

BaseRT engine 0.2.3

Choose a tag to compare

@prabod prabod released this 19 Aug 04:08
0644388

BaseRT 0.2.3

Model pulls that finish. basert pull downloads over many connections at
once and resumes where it stopped — a 20 GB bundle interrupted at 90% picks
up at 90%, including after the process has exited. The 19.8 GB
Qwen3.5-35B-A3B Q4 bundle measured 17.6 MiB/s before and 54–84 MiB/s
after
on the same link, roughly 19 minutes down to 5. A transfer that stalls
without closing is cut off and retried instead of hanging, and a short transfer
is caught rather than installed. Tool calling is fixed on two fronts as well: Muse Glimmer's
dialect is parsed rather than printed, and templates that use is undefined
(Qwen 3.8 among them) render instead of silently falling back.

Added

Xet transfers

  • Xet is supported, behind BASERT_HF_XET=1. Its advantage is
    deduplication across models: chunks shared with something already fetched
    never cross the wire. It cannot resume, though — an interrupted transfer
    starts over — so it is not what basert pull reaches for by default.

    basert pull basecompute/Qwen3.5-35B-A3B                  # resumable (default)
    BASERT_HF_XET=1 basert pull basecompute/Qwen3.5-35B-A3B  # dedup, no resume

Parallel, resumable downloads

  • Files are fetched over many connections at once, in fixed chunks,
    instead of through one stream.

  • An interrupted pull resumes where it stopped, per chunk: completion is
    recorded next to the partial file as each chunk lands and fsynced, so a
    crash, a ^C, or a link that drops overnight continues from what is on disk,
    across process restarts. Just re-run the same basert pull.

  • A stalled transfer fails instead of hanging. Connections that stop
    delivering bytes without closing are now cut off and retried, and a transfer
    shorter than the size the Hub advertised is reported as an error rather than
    installed as a truncated model that fails later at load time.

  • Tunable when the defaults do not suit a link:

    Variable Default Meaning
    BASERT_HF_CONNECTIONS 24 Concurrent range requests
    BASERT_HF_CHUNK_MB 16 Bytes per range request
    BASERT_HF_READ_TIMEOUT_SECS 60 Stall timeout for a single read
    BASERT_HF_MAX_RETRIES 5 Retries per chunk (0 disables)
    BASERT_HF_XET unset Use Xet's CAS path for large files (dedup, no resume)

    Peak memory is connections × chunk — 384 MB at the defaults, since each
    worker holds one chunk in flight. Lower either knob on a constrained
    machine.

Fixed

  • Muse Glimmer's tool calls are parsed, not printed. The model frames a
    call as a message addressed to the tool:

    <|start|>assistant to=fn<|message|><atem:function_calls>
    <atem:invoke name="fn">…</atem:invoke></atem:function_calls><|eot|>
    

    which matched none of the dialects the parser knew (Gemma-native, ChatML
    JSON, Qwen XML), so it fell through to content and the raw markup was
    streamed to the user — the reported symptom was a reply that read
    <atem:invoke name="bash">. The dialect is now recognized and emitted as a
    proper tool call.

  • --kquant-passthrough embeddings on non-Glimmer models. Every model but
    Muse Glimmer handed k-quant embedding lookups the slab parameter struct,
    whose second field is the vocabulary size where the k-quant kernels expect
    the row width. Rows were read at the wrong stride, and the output was
    incoherent — Qwen3-0.6B was the reported case.

  • developer and function roles no longer break the native template.
    A model's chat template accepts the roles it was written for — Qwen 3.8's
    takes system, user, assistant, tool and raises on anything else — and
    a raise drops the whole request into the generic ChatML fallback, quietly
    costing the model its think prefill and tool-call dialect. developer (which
    superseded system in the newer OpenAI API, and is what agent clients send)
    and function (the deprecated spelling of tool) are now resolved before
    rendering. When a template does reject a role, the warning names the roles
    the request carried instead of suggesting a re-convert that would not help.

  • Chat templates using is undefined no longer fall back to ChatML.
    Qwen 3.8's template gates its reasoning-effort block on
    enable_thinking is undefined, a standard Jinja2 test the bundled template
    engine did not implement. Rendering threw on every request and the server
    quietly fell back to generic ChatML — which still looked right for
    ChatML-shaped models while silently dropping the reasoning-effort system
    block, the think prefill, and the model's entire tool-call dialect. Tool
    calling on affected models works again.

Changed

  • Building tools/base-convert now requires Rust 1.85+ (was 1.80).

BaseRT engine 0.2.2

Choose a tag to compare

@prabod prabod released this 12 Aug 07:14
3a0358a

BaseRT 0.2.2

Muse Glimmer, and native GGUF k-quants. BaseRT now runs Muse Glimmer 30B —
text and vision — on Apple Silicon, and reads GGUF k-quantized weights
(Q4_K / Q5_K / Q6_K) directly, without a dequantize-and-repack step. Short
prompts prefill 1.4–1.55× faster on M5.

Added

Muse Glimmer

  • Muse Glimmer 30B, text and vision, including the perception tower — so
    the model answers about images, not just text.

  • Pre-converted bundles at
    basecompute/Muse-Glimmer-30B:

    basert pull basecompute/Muse-Glimmer-30B              # 21 GB, recommended
    basert pull basecompute/Muse-Glimmer-30B:q4k-17gb     # 18 GB, smaller

    Both are the model's own published k-quants, repackaged — the super-block
    bytes are copied through unchanged, so the weights are bit-identical to the
    upstream GGUFs rather than a re-quantization of them. Both carry the vision
    tower; there is no separate projector file to manage.

  • Images use the model's own <|patch|> placeholder:

    basert complete basecompute/Muse-Glimmer-30B --chat \
      --image photo.png --prompt "<|patch|>Describe this image."

Native GGUF k-quants

  • Q4_K / Q5_K / Q6_K run natively — prefill, decode and embedding
    lookup all read the quantized weights directly. Q2_K and Q3_K are
    supported as well.
  • --kquant-passthrough copies GGUF super-blocks into a .base bundle
    verbatim, so the weights stay bit-identical to the source. Re-quantizing an
    already-quantized model compounds its error; copying the blocks through
    introduces none, so this is not a quality trade and needs no override flag.
  • --mmproj folds a companion mmproj-*.gguf perception tower into the
    same bundle. Without it a bundle converted from GGUF is text-only, because
    the tower ships as a separate file. A bundle built this way is
    interchangeable with one converted from the original safetensors.

Changed

  • Faster short-prompt prefill on M5 — 1.4–1.55× measured on M5 Max. Short
    prompts previously fell back to a slower path than the hardware supports;
    they now use the same accelerated one longer prompts already did. Long
    prefill is unchanged, because it was already there.
  • Faster k-quant prefill — 3.4×, from running k-quantized weights on the
    accelerated path rather than the general one.
  • Wider accelerated coverage for the 5-, 2- and 3-bit quantizations,
    alongside the 4-, 6- and 8-bit ones.
  • Decode speed is unchanged. At these model shapes it is limited by memory
    bandwidth rather than compute, so the changes above do not move it, and the
    decode-side ideas tried for this release measured slower and were dropped.

Fixed

  • M3 Ultra was identified as a base M3, so it was tuned for a far smaller
    chip than it is. Reachable only when the system does not report a GPU core
    count, but it would have affected every future per-chip decision.
  • Language bindings could corrupt memory. The model config is returned by
    value, and the Python, Rust and Node definitions of that struct had fallen
    behind the C header — a mirror even four bytes short lets the engine write
    past the caller's allocation. All four bindings are back in step, and a test
    now holds them there.
  • Models converted from GGUF ignored the end-of-turn token and kept
    generating past the end of a reply.
  • baseRT_format_chat is available without a chat-capable build —
    it no longer depends on the chat layer being present.

Notes

  • k-quant support is Apple Silicon (Metal) only in this release.
  • Folding a perception tower from a companion mmproj GGUF is exact for still
    images, which is what these models take today. Such a file does not carry
    everything a video-frame path would need.

BaseRT engine 0.2.1

Choose a tag to compare

@prabod prabod released this 09 Aug 14:10
bbded2c

BaseRT 0.2.1

Speech-to-text. BaseRT now runs the full OpenAI Whisper family — all twelve
variants from tiny to large-v3-turbo, including the English-only builds — as
.base bundles on both Apple Silicon (Metal) and NVIDIA GB10 (CUDA), served
through an OpenAI-compatible transcription API with live streaming.

Added

Whisper speech-to-text

  • All twelve Whisper variants convert to .base and run end-to-end:
    tiny, base, small, medium, large, large-v2, large-v3,
    large-v3-turbo, and the .en English-only builds. Special-token IDs are
    read from each model's own tokenizer rather than inferred from vocabulary
    size, so large-v3 and large-v3-turbo (100 languages, 128 mel bins) map
    correctly.
  • Quantized bundles — whisper-q8 and whisper-q4 profiles quantize the
    linear projections while the convolutional front end, positional and token
    embeddings, norms, and biases stay f16. Q8 meets the same word-error-rate
    parity gates as f16 against the reference implementation.
  • CUDA support — Whisper runs on NVIDIA GB10 with transcripts byte-identical
    to the Metal path.
  • POST /v1/audio/transcriptions and POST /v1/audio/translations,
    OpenAI-compatible, including verbose_json with per-segment
    avg_logprob, no_speech_prob, compression_ratio and temperature, plus
    the detected language.
  • Streaming transcription — stream: true emits segments from the decode
    loop as each window completes, in the documented transcript.text.delta /
    transcript.text.done wire format, so official OpenAI SDKs parse it directly.
    Time-to-first-event is one window's decode rather than the whole file.
  • Decode parity with the reference implementation — task=translate,
    language auto-detection (language: "auto"), initial_prompt with
    condition-on-previous-text, and the temperature-fallback ladder.
    Word-level timestamps remain future work.

C API

  • baseRT_load_model_ex with BaseRTLoadOptions — per-load KV width, paged-KV,
    batch size, prefix cache, prefill chunk and paged-weight mode, replacing the
    process-wide setters. The options struct is ABI-guarded by struct_size.
  • Transcription accessors: baseRT_set_task, baseRT_set_initial_prompt,
    baseRT_set_condition_on_previous_text, per-segment results, detected
    language, and source-audio duration.
  • Batched logits: baseRT_batch_logits_stride, baseRT_argmax_logits_row,
    baseRT_sample_logits_row, baseRT_logits_row_logprobs,
    baseRT_mask_logits_row.
  • baseRT_capabilities reports what the running backend supports.

All of the above are bound in the Python, Node, Swift and Rust bindings.

Changed

  • Faster shared-prefix serving — concurrent requests that share a prompt
    prefix now detect and reuse it through a cascade dispatch path.
  • Faster MoE decode on CUDA — fused SiLU gate/up for 4-bit experts.
  • Metal GEMM accuracy — f32 accumulation on the paths where f16
    accumulation was measurably lossy.

Fixed

  • context_length_exceeded reported a capped token count. The number came
    from a fixed tokenizer buffer, so any over-long prompt reported the same
    figure no matter how far over it was — a client could not trim to a number
    that was itself the ceiling. It now reports the real count.
  • Image requests under-reported usage and never returned
    finish_reason: "length": the first token produced during image prefill was
    streamed to the client but not counted, so a truncated answer looked complete.
  • Gemma 4 audio produced unusable embeddings. The channel-last convolution
    indexed its weights as [out, kh, kw, in] where the bundle stores
    [out, in, kh, kw]. For a single input channel the two are identical, so the
    first subsampling convolution was correct and masked the fault while the
    second convolved with transposed weights.
  • Repetition penalty silently did nothing on some Apple GPUs — a sampling
    kernel read a scratch buffer that was never bound, which is undefined
    behaviour and therefore hardware-dependent.
  • A KV-cache shift kernel had a data race — an in-place shift whose source
    and destination overlap, spread across threads, so results depended on
    scheduling order.
  • Rate limiting could grow without bound. Idle buckets expired on time
    alone, so a source rotating addresses added an entry per request and nothing
    became evictable within the window. Bucket count is now capped.
  • Absolute --media-dir paths are honoured as written and must resolve
    inside the configured root, instead of being re-anchored under it — which
    made a legitimate absolute path fail as an image-decoding error.
  • Rejections now appear in the access log, so an operator can tell a
    refused request from a hung one.
  • Second images and audio parts in a chat request are rejected with a clear
    400 instead of a 500 or, for audio, a confident answer about audio the model
    never received.
  • Image-decoding failures no longer echo the server's resolved filesystem path
    back to the caller.
  • basert inspect reported 0 tensors for working vision and audio towers: it
    counted name prefixes that converted bundles do not use.
  • Chat-template cache eviction and a Qwen3 tool round-trip fault.

Notes

  • BASERT_VERSION now advertises the patch component, and the release pipeline
    verifies the tag, the compiled-in version, and every package manifest agree
    before publishing.

BaseRT engine 0.2.0

Choose a tag to compare

@prabod prabod released this 31 Jul 10:17
6ded5ba

BaseRT 0.2.0

A second hardware backend. BaseRT now runs on NVIDIA GB10 (DGX Spark, Linux/arm64) alongside Apple Silicon — the same .base bundles, the same basert CLI and OpenAI-compatible server, now on CUDA. This release also brings the Qwen3.5 / Qwen3.6 hybrid model family and a continuous-batching overhaul of basert serve.

Added

  • CUDA backend — NVIDIA GB10 / DGX Spark — full inference on Linux/arm64 + CUDA: dense, MoE, and hybrid (Gated-DeltaNet) architectures, across prefill, single-stream decode, and continuous-batching serving. install.sh auto-detects the platform (macOS/arm64 → Metal, Linux/arm64 → CUDA) and basert pull resolves the right bundle for the host.
  • Qwen3.5 & Qwen3.6 hybrid models — support for the Gated-DeltaNet hybrid architecture (dense and MoE), including vision (Qwen3.5-VL), on both Metal and CUDA. GB10 continuous batching for the dense hybrids.
  • Continuous-batching serving — basert serve now drives tool calls (streaming + non-streaming), grammar-constrained decode, per-token logprobs, and n>1 choices through the continuous-batching engine, with a decode-priority scheduler and radix prefix-cache reuse (dense and hybrid, via a keyed GDN snapshot store).
  • CUDA-native model bundles — cuda-q8 and cuda-q4mix variants for 11 models (Qwen3-0.6B/30B-A3B, Qwen3.5-2B/35B-A3B, Qwen3.6-27B/35B, Llama-3.2-1B/3B, Gemma-3-1B, Gemma-4-E2B/26B), on Hugging Face and the catalog.
  • Faster Metal GEMMs — large-tile prefill GEMM kernels for native bf16 / f16 weights and Q8, plus batched Gated-DeltaNet decode/prefill kernels for the M1–M5 families.
  • Tokenizer conformance — an HF-exact conformance harness and a Unicode-category pretokenizer; Phi-3-mini un-quarantined after revalidation.

Fixed

  • Chunked-decode correctness on CUDA — free generation past the decode chunk size no longer repeats an earlier block (an eager-dispatch write-after-read hazard); prefill and teacher-forced decode were always correct.
  • Large quantized tensors — weight tensors with more than ~512M elements no longer decode to garbage (64-bit byte-offset fix).
  • Gemma normalization precision — the canonical post-attention / feed-forward norm weights are kept in floating point in the shipped bundles (were quantized in some profiles).
  • Hybrid prefix reuse — correct KV/GDN dtype handling and prefix-cache reuse for hybrid-GDN models under serving.
  • MoE quantization parity — Qwen3.5 MoE quant rules matched to canonical tensor names so experts quantize as intended.

Changed

  • Backend-aware model resolution — CUDA hosts prefer CUDA-native bundles, Apple Silicon prefers Metal, with a universal fallback; catalog entries are backend-tagged and can be derived from a .base header.
  • Performance — Metal decode/prefill tuning across the M1 Max / M4 Pro / M5 Pro families (large-tile GEMM, higher-occupancy Gemma-4 MoE gate/up). Benchmarks vs llama.cpp and vLLM on GB10 and Apple Silicon are in the repo.

Install

# macOS (Apple Silicon, Metal) or Linux (arm64, CUDA) — auto-detected
curl -fsSL https://raw.githubusercontent.com/basecompute/baseRT/main/install.sh | sh
basert pull <model>          # resolves the right bundle for your host
basert serve <model>.base    # OpenAI-compatible server

BaseRT engine 0.1.7

Choose a tag to compare

@prabod prabod released this 21 Jul 03:10
7a63702

Maintenance release focused on the three most-reported serve/CLI issues, plus converter and chat quality-of-life fixes.

Fixed

  • Tool calls for Qwen 3.5 / 3.6 (#21) — basert serve now emits structured tool_calls for the Qwen XML tool-call dialect, in both streaming and non-streaming modes. Truncated generations never execute partial arguments, and logprobs stay aligned across tool-call segments.
  • Deterministic HTTP 500 on some short CJK prompts (#22) — fixed a byte-alphabet hole (soft hyphen 0xAD) in GPT-2-style byte tokenization; tokenization is now HF-exact for the affected models. Server errors are structured and logged instead of opaque 500s.
  • --no-think and chat-template handling (#23) — --no-think now works across all thinking-capable architectures; the CLI renders model-embedded chat templates; fixed a Gemma 4 channel-marker leak and KV-reuse correctness on divergent history.
  • MLX 3/5/6-bit conversion — MLX quantized checkpoints at 3/5/6-bit widths now unpack correctly (little-endian bitstream); previously these converted to wrong weights.
  • basert pull HF-cache duplication — pulled models are no longer duplicated in the global Hugging Face cache (staged download with resume preserved; set BASERT_KEEP_HF_SOURCES=1 to keep sources).

Added

  • Chat line editing — cursor movement, kill ops, and word delete (Ctrl+W / Alt+Backspace / Ctrl+Backspace) in basert chat, with UTF-8/CJK support.
  • q6 MoE kernels — Q6 expert GEMM (plus q8/q6 scale-dtype siblings) and a bundled default-q6 conversion profile (--target base-q6).
  • baseRT_decode_token_raw() — length-preserving raw-byte token decode in the C API (byte-level BPE tokens can legally contain NUL; the C-string variants truncate there). No struct-layout changes.

Changed

  • Chat / complete / bench banners show the model id (org/repo · variant) instead of model.base.
  • basert help output regrouped runtime-first.

Performance is unchanged vs 0.1.6 (validated dense + MoE on M4 Pro and M5 Pro, within noise).

BaseRT engine 0.1.6

Choose a tag to compare

@prabod prabod released this 18 Jul 07:10
67c2dac

BaseRT 0.1.6

  • Qwen 3.5 / 3.6 support: dense, MoE, and vision-language models (hybrid Gated-DeltaNet architecture), in both the runtime and basert convert.
  • M5 tensor-core kernels included: release binaries now ship the Metal 4 kernel library (baseRT_tensor.metallib) — earlier releases were missing it, so M5 GPUs fell back to slower kernels. Loads automatically; older Macs are unaffected.
  • Big-model support: models larger than the GPU wiring limit now load and decode at full speed (e.g. a 28 GB model on a 48 GB Mac) instead of failing.
  • Qwen3.5-VL accuracy: image preprocessing and multimodal RoPE now match the reference implementation token-for-token.
  • CLI: basert-complete accepts -n for max tokens; unknown flags are now errors.
  • Requires macOS 15+ (Metal 4 kernels engage on macOS 26 / M5).

Contents: libbaseRT.dylib (embedded kernels) + basert CLI + basert-* runtime tools + headers.