Skip to content

fix(runtime): support vLLM 0.26 and SGLang 0.5.16 - #2991

Merged
Qubitium merged 1 commit into
mainfrom
fix/vllm-sglang-latest-compat
Jul 31, 2026
Merged

fix(runtime): support vLLM 0.26 and SGLang 0.5.16#2991
Qubitium merged 1 commit into
mainfrom
fix/vllm-sglang-latest-compat

Conversation

@ZX-ModelCloud

@ZX-ModelCloud ZX-ModelCloud commented Jul 31, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • update vLLM optional-import handling, sampling conversion, token-prompt generation, and config/device discovery for current engine layouts
  • support SGLang Engine while retaining the Runtime fallback, including normalized runtime aliases, device IDs, sampling parameters, and token inputs
  • accept runtime-supported checkpoint formats, preserve explicit GPU selection, and filter deprecated Evalution constructor arguments

Comment thread gptqmodel/utils/sglang.py Dismissed
Comment thread gptqmodel/utils/vllm.py Dismissed
@Qubitium
Qubitium merged commit 45e5866 into main Jul 31, 2026
6 checks passed
@Qubitium
Qubitium deleted the fix/vllm-sglang-latest-compat branch July 31, 2026 14:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants