Skip to content

Commit f31b8a9

Browse files
committed
docs/tests: update default Qwen model to 3.5 0.8B
1 parent 3a70766 commit f31b8a9

5 files changed

Lines changed: 13 additions & 13 deletions

File tree

‎.github/workflows/test.yaml‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -8,8 +8,8 @@ on:
88
- main
99

1010
env:
11-
REPO_ID: Qwen/Qwen2-0.5B-Instruct-GGUF
12-
MODEL_FILE: qwen2-0_5b-instruct-q8_0.gguf
11+
REPO_ID: lmstudio-community/Qwen3.5-0.8B-GGUF
12+
MODEL_FILE: Qwen3.5-0.8B-Q8_0.gguf
1313

1414
jobs:
1515
download-model:

‎README.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -322,8 +322,8 @@ You'll need to install the `huggingface-hub` package to use this feature (`pip i
322322

323323
```python
324324
llm = Llama.from_pretrained(
325-
repo_id="Qwen/Qwen2-0.5B-Instruct-GGUF",
326-
filename="*q8_0.gguf",
325+
repo_id="lmstudio-community/Qwen3.5-0.8B-GGUF",
326+
filename="*Q8_0.gguf",
327327
verbose=False
328328
)
329329
```
@@ -685,7 +685,7 @@ For possible options, see [llama_cpp/llama_chat_format.py](llama_cpp/llama_chat_
685685
If you have `huggingface-hub` installed, you can also use the `--hf_model_repo_id` flag to load a model from the Hugging Face Hub.
686686

687687
```bash
688-
python3 -m llama_cpp.server --hf_model_repo_id Qwen/Qwen2-0.5B-Instruct-GGUF --model '*q8_0.gguf'
688+
python3 -m llama_cpp.server --hf_model_repo_id lmstudio-community/Qwen3.5-0.8B-GGUF --model '*Q8_0.gguf'
689689
```
690690

691691
### Web Server Features

‎examples/gradio_chat/local.py‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -4,10 +4,10 @@
44
import gradio as gr
55

66
llama = llama_cpp.Llama.from_pretrained(
7-
repo_id="Qwen/Qwen1.5-0.5B-Chat-GGUF",
8-
filename="*q8_0.gguf",
7+
repo_id="lmstudio-community/Qwen3.5-0.8B-GGUF",
8+
filename="*Q8_0.gguf",
99
tokenizer=llama_cpp.llama_tokenizer.LlamaHFTokenizer.from_pretrained(
10-
"Qwen/Qwen1.5-0.5B"
10+
"Qwen/Qwen3.5-0.8B"
1111
),
1212
verbose=False,
1313
)

‎examples/hf_pull/main.py‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -3,10 +3,10 @@
33

44

55
llama = llama_cpp.Llama.from_pretrained(
6-
repo_id="Qwen/Qwen1.5-0.5B-Chat-GGUF",
7-
filename="*q8_0.gguf",
6+
repo_id="lmstudio-community/Qwen3.5-0.8B-GGUF",
7+
filename="*Q8_0.gguf",
88
tokenizer=llama_cpp.llama_tokenizer.LlamaHFTokenizer.from_pretrained(
9-
"Qwen/Qwen1.5-0.5B"
9+
"Qwen/Qwen3.5-0.8B"
1010
),
1111
verbose=False,
1212
)

‎tests/test_llama.py‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -58,8 +58,8 @@ def test_llama_cpp_tokenization():
5858

5959
@pytest.fixture
6060
def llama_cpp_model_path():
61-
repo_id = "Qwen/Qwen2-0.5B-Instruct-GGUF"
62-
filename = "qwen2-0_5b-instruct-q8_0.gguf"
61+
repo_id = "lmstudio-community/Qwen3.5-0.8B-GGUF"
62+
filename = "Qwen3.5-0.8B-Q8_0.gguf"
6363
model_path = hf_hub_download(repo_id, filename)
6464
return model_path
6565

0 commit comments

Comments
 (0)