Usage Statistics

Historical request usage loaded from the Docker-mounted telemetry log. Choose a time window and the model profiles to include.

Models included (1 / 10 selected)
Dell_L40_VLLM1 model
GB10_15 models
GB10_1_openai1 model
gb10_1_vllm1 model
NVIDIA 50902 models
Showing 6 requests from 2026-09-23 16:40 through 2026-10-06 17:09. Averages exclude zero-valued observations/buckets. Dashboard Reset does not delete this retained history. Per-model Clear resets a live model's usage or removes a historical-only entry from the selector; the retained file and global all-time totals are never rewritten. History file: /persist/data/token-events.jsonl
Requests
6
4 errors ยท 4 missing usage
Input Tokens
16,546
prompt/input usage
Output Tokens
896
generated/completion usage
Avg Duration
16.46s
TTFT 18.81s average
Avg Output Rate
14.7
tokens/sec during active generation

Requests

Completed requests, errors, and requests without upstream usage data.

requestserrorsmissing usagerequest avg 6 reqerror avg 4 reqmissing avg 4 req
6 req0 req
09-23 16:4010-06 16:40

Input Tokens

Prompt/input tokens consumed in each time bucket.

input tokensavg 16.55K tok
16.55K tok0 tok
09-23 16:4010-06 16:40

Output Tokens

Generated/completion tokens produced in each time bucket.

output tokensavg 896 tok
896 tok0 tok
09-23 16:4010-06 16:40

Latency

Average end-to-end request duration and time to first token.

durationTTFTduration avg 16.46 sTTFT avg 18.81 s
18.81 s0 s
09-23 16:4010-06 16:40

Generation Throughput

Output tokens divided by active generation seconds for each bucket.

output tok/savg 14.7 tok/s
14.7 tok/s0 tok/s
09-23 16:4010-06 16:40

Breakdown by Model

SystemModelRequestsErrorsInputOutputTotalAvg durationAvg TTFTAvg tok/s
GB10_1_openaiGB10_1_Qwen_38_openai
GB10_1_Qwen_38_openai
6416,54689617,44216.46s18.81s14.7