Model Profiles
Profiles are ordered as endpoint name → runtime type → profile name. Each profile is bound to one exact endpoint/runtime pair.
Binding:
GB10_1_openai (openai_compatible) GB10_1_Qwen_38_openai · record legacy_a71Use Refresh Models From Endpoints to query /v1/models. Discovery is scoped to the exact endpoint name/runtime pair.
Saved Model Records
This list is read directly from /persist/config/model_profiles.json. endpoints.json is never a model-profile source. Routing and /v1/models read this same file authority.
| Status | Endpoint / profile | Display name | Upstream | Record | Actions |
|---|---|---|---|---|---|
| PUBLISHED / routable | 5090 (lm_studio) nvidia_5090___qwen3.5-9b-uncensored-hauhaucs-aggressive | qwen3.5-9b-uncensored-hauhaucs-aggressive via NVIDIA 5090 | qwen3.5-9b-uncensored-hauhaucs-aggressive | legacy_21e | |
| PUBLISHED / routable | Dell_L40 (openai_compatible) Qwem_3_8_Flash_Next | Qwem_3_8_Flash_Next_L40 | qwen3.8-flash-next-q2_0 | 505bad7da0 | |
| PUBLISHED / routable | gb10_1 (vllm) gb10_1_vllm_deepseek_ai_deepseek_v4_1_flash | deepseek-ai/DeepSeek-V4.1-Flash | deepseek-ai/DeepSeek-V4.1-Flash | legacy_821 | |
| PUBLISHED / routable | gb10_1 (vllm) gb10_1_vllm_deepseek_v41_flash_exl3 | deepseek-v41-flash-exl3 | deepseek-v41-flash-exl3 | legacy_545 | |
| PUBLISHED / routable | gb10_1 (vllm) gb10_1_vllm_deepseek_v4_flash_dspark_abliterated | deepseek-v4-flash-dspark-abliterated | deepseek-v4-flash-dspark-abliterated | legacy_c62 | |
| PUBLISHED / routable | gb10_1 (vllm) gb10_1_vllm_qwen3_8_flash_next | qwen3.8-flash-next | qwen3.8-flash-next | legacy_21d | |
| PUBLISHED / routable | GB10_1_openai (openai_compatible) GB10_1_Qwen_38_openai | GB10_1_Qwen_38_openai | qwen3.8-flash-next-nvfp4 | legacy_a71 |
Current Model Catalog
Ordered by endpoint name, runtime type, then profile name.
| Endpoint / profile | Display name | Upstream model | Online | Discovered | Preset | Benchmark | Metadata |
|---|---|---|---|---|---|---|---|
5090 (lm_studio) nvidia_5090___qwen3.5-9b-uncensored-hauhaucs-aggressive | qwen3.5-9b-uncensored-hauhaucs-aggressive via NVIDIA 5090 | qwen3.5-9b-uncensored-hauhaucs-aggressive | False | False | creative_brainstorming | — tok/s TTFT — ms | total / active B; ; ctx |
Dell_L40 (openai_compatible) Qwem_3_8_Flash_Next | Qwem_3_8_Flash_Next_L40 | qwen3.8-flash-next-q2_0 | True | True | general_assistant | 166.3 tok/s TTFT 48 ms | total / active B; ; ctx |
gb10_1 (vllm) gb10_1_vllm_deepseek_ai_deepseek_v4_1_flash | deepseek-ai/DeepSeek-V4.1-Flash | deepseek-ai/DeepSeek-V4.1-Flash | True | False | glm_5_3_flash_speculative | — tok/s TTFT — ms | total / active B; ; ctx |
gb10_1 (vllm) gb10_1_vllm_deepseek_v41_flash_exl3 | deepseek-v41-flash-exl3 | deepseek-v41-flash-exl3 | True | False | glm_5_3_flash_speculative | — tok/s TTFT — ms | total / active B; ; ctx |
gb10_1 (vllm) gb10_1_vllm_deepseek_v4_flash_dspark_abliterated | deepseek-v4-flash-dspark-abliterated | deepseek-v4-flash-dspark-abliterated | True | False | deepseek_speculative | — tok/s TTFT — ms | total / active B; ; ctx |
gb10_1 (vllm) gb10_1_vllm_qwen3_8_flash_next | qwen3.8-flash-next | qwen3.8-flash-next | True | True | glm_5_3_flash_speculative | 45.3 tok/s TTFT 175 ms | 6 total / 200 active B; NVFP4; 240000 ctx |
GB10_1_openai (openai_compatible) GB10_1_Qwen_38_openai | GB10_1_Qwen_38_openai | qwen3.8-flash-next-nvfp4 | True | False | creative_brainstorming | 20.9 tok/s TTFT 237 ms | total / active B; ; ctx |