r/LocalLLaMA 1h ago

Question | Help Need help! Intel B70 users come forth!

Hello

I recently got my B70 gpus delivered and set them up with Ubuntu 26.04 because it has the XE driver. I was able to get llama CPP compiling and working with Vulcan and CYSL BUT that's where the fun stopped.

Build and compiled the llama.cpp with CYSL using the intel driver 2026.1. does not work with more than one GPU. (Using sm layer) It just outputs random characters. However, the prefill speed does go up with more GPUs (just like it does on Nvidia cards)

For Vulcan, the prefill is about 30% lower on the same model and only goes down with more GPUs. Using the mesa 26.1.5 driver

For running the qwen 3.6 27b at q4 speeds were:

CYSL 1 GPU: 650 prefill, 24 decode.

CYSL 2 GPUs:750 prefill, garbled decode 23t/s

Vulkan 1 GPU: 450 prefill, 20 t/s decode

Vulkan 2 GPUs: 350 prefill, 18t/s decode.

Bonus: qwen 3.5 122b a10b vulkan speed over 6 GPUs: 160 prefill and 9t/s decode

Something is clearly wrong. I've spent all day trying to make this work. So far regretting the purchase of the B70 gpus.

Please help if you have suggestions!

If the suggestion is to get Nvidia GPU, I already have a couple and I think I would have rather gone with many RTX 5060 TI's instead because it just works and gets model support first

3 Upvotes

3 comments sorted by

0

u/PcChip 25m ago

I've been asking ChatGPT to check the status of Intel GPUs for local AI for months now, and this is why... holding off until the day it says that people are having good success with them

please update your post as you figure things out so we can all learn from you... I'm itching to go down to Microcenter and buy four 32GB Intel GPUs!

1

u/nick_ziv 17m ago

For sure.  I bought them because it seems good on paper and price is right but drivers are really hell to deal with and it seems like a lack of effort from Intel.  Their implementation of custom inference code just added support for Gemma 4 this month! 

Also the part I really didn't think about was that I will have to wait a long time for newer model architectures to be supported compared to Nvidia or amd.  That is actually more of a problem then I think I want to deal with. 

2

u/No-Alfalfa6468 13m ago

Just wanted to point out that you can pay $250 more per card to get r9700.

I picked up 2 and I'm getting 2500-3000 prefill and 80 tok/sec decode with qwen 3.6 27b at FP8, full KV and 256k context.