196 tok/s per user
GLM 5.3 on AMD MI350X node, 64 concurrent users, 500k context.
Per-user decode speed
SwarmOne tunes your serving stack to your real agent traffic. Same model. Same hardware.
196 tok/s per user
GLM 5.3 on AMD MI350X node, 64 concurrent users, 500k context.
Per-user decode speed
680 ms TTFT (P95)
GLM 5.3, on NVIDIA B300 node, at 172 tok/s per user.
Time to first token
6 TO 104
17× faster
Qwen3.8-27B on Tenstorrent Loudbox, 104 tok/s per user vs 6 tok/s untuned.
Before to after
Most published speed numbers use 1k-token prompts. Agents send 100k to 500k. That is where default stacks slow down, and where we measure.
Agent sessions grow to 500k tokens, pause for tool calls, then come back. Cache placement, eviction, and routing decide how fast they run. Default configs and generic benchmarks miss all of it.
150-200
Commodity hardware typically lands at 30-40 tokens per second per user on real agent traffic. After optimization: 150-200.
Record a few real agent sessions. Search the stack against that traffic. Ship only what measured better. Then do it again the next night.
01
Capture the real shape
Measure with SwarmSim. Search with SwarmOptimizer. Start with one, or run both.
Test on your real traffic, not a benchmark.
Turns a few recorded agent sessions into thousands of new ones with the same shape, and runs them against any endpoint. See tok/s per user, TTFT, completion time, and cost before you ship a change.
Record: A few real agent traces and tool calls
Perturb: Same work, new conversations
Simulate: Thousands of fresh test runs
Measure: tok/s per user, TTFT, completion time, cost
Finds the fastest setup for that traffic.
Searches configs, KV cache policy, routing, and kernels on SwarmSim traffic, then ships the measured winner. Every test uses fresh traffic, so the optimizer can't game the benchmark. Bring your own optimizer if you have one.
Searches configs, KV cache, routing, and kernels
Fresh traffic on every test
Ships only the measured winner
Works with your optimizer or ours
NVIDIA, AMD, and more. Runs on your cluster, connected or air-gapped. No rewrites. No lock-in.
SILICON
Send us a few traces. We'll show you tok/s per user, TTFT, and cost on your own traffic, before and after.
Talk to Us