Model Profiles

Profiles are ordered as endpoint name → runtime type → profile name. Each profile is bound to one exact endpoint/runtime pair.

Binding: GB10_1_openai (openai_compatible) GB10_1_Qwen_38_openai · record legacy_a71
Use Refresh Models From Endpoints to query /v1/models. Discovery is scoped to the exact endpoint name/runtime pair.

Saved Model Records

This list is read directly from /persist/config/model_profiles.json. endpoints.json is never a model-profile source. Routing and /v1/models read this same file authority.

StatusEndpoint / profileDisplay nameUpstreamRecordActions
PUBLISHED / routable5090 (lm_studio) nvidia_5090___qwen3.5-9b-uncensored-hauhaucs-aggressiveqwen3.5-9b-uncensored-hauhaucs-aggressive via NVIDIA 5090 qwen3.5-9b-uncensored-hauhaucs-aggressivelegacy_21e
PUBLISHED / routableDell_L40 (openai_compatible) Qwem_3_8_Flash_NextQwem_3_8_Flash_Next_L40qwen3.8-flash-next-q2_0505bad7da0
PUBLISHED / routablegb10_1 (vllm) gb10_1_vllm_deepseek_ai_deepseek_v4_1_flashdeepseek-ai/DeepSeek-V4.1-Flashdeepseek-ai/DeepSeek-V4.1-Flashlegacy_821
PUBLISHED / routablegb10_1 (vllm) gb10_1_vllm_deepseek_v41_flash_exl3deepseek-v41-flash-exl3deepseek-v41-flash-exl3legacy_545
PUBLISHED / routablegb10_1 (vllm) gb10_1_vllm_deepseek_v4_flash_dspark_abliterateddeepseek-v4-flash-dspark-abliterateddeepseek-v4-flash-dspark-abliteratedlegacy_c62
PUBLISHED / routablegb10_1 (vllm) gb10_1_vllm_qwen3_8_flash_nextqwen3.8-flash-nextqwen3.8-flash-nextlegacy_21d
PUBLISHED / routableGB10_1_openai (openai_compatible) GB10_1_Qwen_38_openaiGB10_1_Qwen_38_openaiqwen3.8-flash-next-nvfp4legacy_a71

Current Model Catalog

Ordered by endpoint name, runtime type, then profile name.

Endpoint / profileDisplay nameUpstream modelOnlineDiscoveredPresetBenchmarkMetadata
5090 (lm_studio) nvidia_5090___qwen3.5-9b-uncensored-hauhaucs-aggressiveqwen3.5-9b-uncensored-hauhaucs-aggressive via NVIDIA 5090 qwen3.5-9b-uncensored-hauhaucs-aggressiveFalseFalsecreative_brainstorming— tok/s
TTFT — ms
total / active B; ; ctx
Dell_L40 (openai_compatible) Qwem_3_8_Flash_NextQwem_3_8_Flash_Next_L40qwen3.8-flash-next-q2_0TrueTruegeneral_assistant166.3 tok/s
TTFT 48 ms
total / active B; ; ctx
gb10_1 (vllm) gb10_1_vllm_deepseek_ai_deepseek_v4_1_flashdeepseek-ai/DeepSeek-V4.1-Flashdeepseek-ai/DeepSeek-V4.1-FlashTrueFalseglm_5_3_flash_speculative— tok/s
TTFT — ms
total / active B; ; ctx
gb10_1 (vllm) gb10_1_vllm_deepseek_v41_flash_exl3deepseek-v41-flash-exl3deepseek-v41-flash-exl3TrueFalseglm_5_3_flash_speculative— tok/s
TTFT — ms
total / active B; ; ctx
gb10_1 (vllm) gb10_1_vllm_deepseek_v4_flash_dspark_abliterateddeepseek-v4-flash-dspark-abliterateddeepseek-v4-flash-dspark-abliteratedTrueFalsedeepseek_speculative— tok/s
TTFT — ms
total / active B; ; ctx
gb10_1 (vllm) gb10_1_vllm_qwen3_8_flash_nextqwen3.8-flash-nextqwen3.8-flash-nextTrueTrueglm_5_3_flash_speculative45.3 tok/s
TTFT 175 ms
6 total / 200 active B; NVFP4; 240000 ctx
GB10_1_openai (openai_compatible) GB10_1_Qwen_38_openaiGB10_1_Qwen_38_openaiqwen3.8-flash-next-nvfp4TrueFalsecreative_brainstorming20.9 tok/s
TTFT 237 ms
total / active B; ; ctx