Running tasks (slots)
Connecting…
no active instance IDLE
No active model instance.
Processing speed (last 5 minutes)
Prompt / prefill (tokens/s) Decode / generation (tokens/s)
Global metrics

Prompt/decode rates are derived from the server counters (Δtokens / Δbusy-time) and attributed across the busy window — the server only flushes some counters when a task completes.

Requests in progress
0
active slots
Requests queued
0
in waiting queue
Prompt
0
tok/s · busy-time rate
Decode
0
tok/s · busy-time rate
Prompt tokens (total)
0
excl. cache
Generated tokens (total)
0
 
Served from cache
0
prompt cache hits
Sequence peak
0
tokens (prompt + gen.)
llama_decode() calls
0
 
Server uptime
—
since session start
Speculation (spec-decode)
Drafts (steps)
0
Draft tokens
0
Accepted tokens
0
Acceptance rate
—