
AI
Cerebras serves Qwen 3.8 27B at 1500 tokens/s, and the agent stops waiting
Cerebras's inference documentation lists Qwen 3.8 27B served at roughly 1500 tokens per second on its public endpoints. For a chatbot that is comfort; for an agent loop, it changes what the thing is.