Auto Model Detection, Simplified Args, and Overflow Safeguards - #203
Merged
Conversation
ephemeral2eternity
pushed a commit
to ephemeral2eternity/candle-vllm
that referenced
this pull request
Nov 20, 2025
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR includes three major improvements:
1) Automatic model type detection
The model type is now automatically inferred based on the input path or file, significantly simplifying usage.
Examples:
Safetensors model:
GGUF model:
Multi-rank inference:
Automatically enables multi-process mode (via daemon processes):
ISQ inference:
2) Simplified CLI parameters for improved usability
To make the command-line interface more concise and user-friendly, key arguments have been shortened:
--weight-path→--w--weight-file→--f--model-id→--m--device-ids→--d--port→--pAlso:
--multi-processflag has been removed. Multi-process mode is now automatically enabled for multi-rank inference.--multithreadflag.3) Improved safeguards against overflow
This update introduces a mechanism to prevent generation overflow.
It checks the available KV cache and model context length before accepting new requests, ensuring that inputs exceeding the model’s capacity are rejected early and safely.