EN HU
← AI Lab
Scrap to Spark / Field study 01

The BC-250 Is a Capable Local AI PC, Too

My AMD BC-250 local AI experiments: Ornith inference at 89–90 tokens per second sustained, Vulkan tuning, and a 70K-token Qwen context.

Illustrated AMD BC-250 board and heatsink with the measured 90 tokens per second sustained result
Illustrated cover based on the BC-250 hardware. Performance figures come from the experiments described below.

Linus Tech Tips recently shared a video about the BC-250, showing what this unusual board can do as a budget gaming PC.

There’s another side of that story I want to add: it’s also a surprisingly capable local AI machine.

I’ve been using the AMD BC-250 to run language models, investigate Vulkan kernels, and explore what a useful local assistant needs beyond a fast response. With Ornith 1.5 35B-A3B, I took generation from approximately 64 tokens per second to 89–90 sustained, with short-context peaks of 97–101 tokens per second.

That’s roughly 39–41% higher sustained throughput, and 52–58% higher peak throughput, relative to my starting configuration.

This work is part of the independent AI lab I’m building toward: practical experiments with local models, accessible hardware, and results I can measure.

The results first

Ornith throughput: initial configuration approximately 64 tokens/s; tuned sustained 89–90; short-context peak 97–101; one approximately 10K-token generation averaged 84.9.

Results from my own experimental runs. Context lengths and configurations varied; the peak is not a sustained average.

The 50%+ peak improvement is a useful headline. The result I care about most is holding around 90 tokens per second across several thousand generated tokens.

In one longer generation of approximately 10,000 tokens, the average was still 84.9 tokens per second. As context grew, attention took a larger share of the time needed to produce each token.

These gains came from the combined configuration: the 40-CU setup, Vulkan workgroup tuning, MTP speculative decoding, and additional system/GPU tuning. I haven’t isolated a percentage contribution for each change.

What I was running

My setup used the BC-250 in a 40-compute-unit configuration, with llama.cpp’s Vulkan backend.

I used two workloads:

  • Ornith 1.5 35B-A3B: a sparse mixture-of-experts model. My mixed-precision GGUF used heavily quantized Q2_K experts, with higher precision for selected components and multi-token prediction.
  • Qwen3.8 27B: a different workload for exploring quantization, memory constraints, and long conversations.

Ornith’s nominal 35 billion parameters don’t all activate for every token. That distinction matters when interpreting its speed: parameter count alone doesn’t tell you how much computation a model performs per token.

Getting the models running was the starting point. Then I wanted to understand where the decode time actually went.

Finding the time between tokens

Profiling showed several contributors: Q2_K expert operations, IQ matrix-vector operations, attention, the output projection, and many small Vulkan dispatches.

I compared a BC-250-specific llama.cpp fork against upstream and experimented with increasing the number of K-quant output rows processed per workgroup:

rm_kq: 2 → 4

I also tested the corresponding IQ path:

rm_iq = 2 * rm_kq

Some individual operations improved substantially. The Q5_K vocabulary/output operation reached approximately 2.03 TFLOPS, and some IQ2_S operations exceeded 2.0 TFLOPS.

Those are kernel measurements. To decide whether a change was useful, I went back to complete generation runs.

Saving microseconds in a small operation has limited value when several milliseconds are spent elsewhere. The metric that guided this work was total time per generated token.

Letting the model propose its next token

Ornith includes a multi-token prediction head, or MTP. It can propose a token that the model then verifies, without loading a separate draft model.

The best configuration in my experiments used one draft token:

--spec-type draft-mtp
--spec-draft-n-max 1
--spec-draft-p-min 0.0

A typical long run showed approximately 80% draft acceptance and a mean accepted length of about 1.8. Draft acceptance describes how often the proposed tokens were accepted during verification; it isn’t a measure of answer quality.

MTP also changed the kernel workload. Verification evaluates multiple tokens together, so Vulkan performance and speculative decoding had to be considered together.

This combination helped move the system toward 89–90 tokens per second sustained, with the higher short-context peaks shown above.

Qwen: making the model fit is only the beginning

The Qwen3.8 27B files I tested included a roughly 51 GB BF16 version and a much more practical 10 GB UD-Q2_K_XL quantization.

Qwen3.8 27B model files: approximately 51 GB for BF16 and 10 GB for UD-Q2_K_XL. File size excludes runtime memory and KV cache.

Model-file sizes, not total runtime memory. This comparison does not establish equivalent output quality.

My Vulkan configuration used forced MMVQ, Flash Attention, a Q8_0 KV cache, and GPU layer offload:

GGML_VK_FORCE_MMVQ=1
-fa on
-ctk q8_0
-ctv q8_0
-ngl all

That gave me a platform for the next experiment: how far could I push an active conversation?

A 70K conversation still needs memory management

I pushed an active conversation to approximately 70,000 tokens inside a 71,680-token context.

That demonstrated context capacity. It didn’t establish perfect recall across the whole conversation, and it didn’t make the context window unlimited. Eventually, it fills.

KV-cache shifting also doesn’t preserve the meaning of everything removed. A useful long-running assistant needs a deliberate way to retain decisions, important facts, and experimental results.

The direction I want to explore is periodic context compaction. At roughly 58K active tokens, a candidate approach would summarize around 25K tokens of older history into 3–5K of structured memory, while keeping the newest approximately 30K tokens verbatim.

Proposed working-memory design: summarize approximately 25K older tokens into 3–5K structured tokens, retain approximately 30K recent tokens, preserve instructions and the current request, and archive the full history for retrieval.

Proposed architecture, not a measured memory-retention result. The approximate budgets leave room for instructions, the current request, and other overhead.

The summary would retain facts, decisions, and results. The full conversation would remain on disk for later retrieval.

The 70K context was tested. This compaction design is the next research direction.

One optimization I’m still investigating

Ornith’s MoE graph revealed a promising opportunity: its gate and up expert projections consume the same activation and selected expert IDs.

A combined kernel could share activation preparation and reduce intermediate memory traffic and dispatch overhead. I confirmed the graph topology, including the [512, 8, 2] workload used in batch-2 MTP verification, and built an experimental Q2_K MMVQ Vulkan shader that compiled to SPIR-V.

The runtime integration into the generic MUL_MAT_ID path became complicated, and it wasn’t stable enough to merge.

I’m not counting fusion as a performance gain. The opportunity is worth investigating, but it still needs reliable end-to-end validation.

What this means for the lab

The BC-250 has become a useful research platform for me. It gives me somewhere to study inference kernels, speculative decoding, quantization, and the memory architecture of local agents on hardware I own.

Ornith showed how much the complete software configuration can affect generation speed. Qwen made the next challenge concrete: once a model is useful locally, managing its context becomes part of building the system around it.

That’s the kind of independent AI lab I want to create. My interests in local AI and hardware meet in these experiments: understand the machine, measure what changes, and turn the results into something useful.

If you’re working on local inference, unusual hardware, or long-running agents, I’d like to compare notes.

Follow the lab: @knowixbuilds.

Measurement note: these are approximate results from my own experiments across changing configurations and context lengths. The write-up does not include exact build revisions, model-file identifiers, clocks, or raw run logs. Treat it as a practical case study of the combined setup, rather than a controlled comparison of individual optimizations.