You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 1de1569
Browse filesBrowse the repository at this point in the historyBrowse files
# int32_t n_gpu_layers; // number of layers to store in VRAM
804
+
# int32_t n_gpu_layers; // number of layers to store in VRAM, a negative value means all layers
793
805
# enum llama_split_mode split_mode; // how to split the model across multiple GPUs
806
+
# enum llama_load_mode load_mode; // how to load the model
794
807
795
808
# // the GPU that is used for the entire model when split_mode is LLAMA_SPLIT_MODE_NONE
796
809
# int32_t main_gpu;
@@ -812,9 +825,6 @@ class llama_model_imatrix_data(ctypes.Structure):
812
825
813
826
# // Keep the booleans together to avoid misalignment during copy-by-value.
814
827
# bool vocab_only; // only load the vocabulary, no weights
815
-
# bool use_mmap; // use mmap if possible
816
-
# bool use_direct_io; // use direct io, takes precedence over use_mmap when supported
817
-
# bool use_mlock; // force system to keep model in RAM
818
828
# bool check_tensors; // validate model tensor data
819
829
# bool use_extra_bufts; // use extra buffer types (used for weight repacking)
820
830
# bool no_host; // bypass host buffer allowing extra buffers to be used
@@ -826,17 +836,15 @@ class llama_model_params(ctypes.Structure):
826
836
Attributes:
827
837
devices (ctypes.Array[ggml_backend_dev_t]): NULL-terminated list of devices to use for offloading (if NULL, all available devices are used)
828
838
tensor_buft_overrides (ctypes.Array[llama_model_tensor_buft_override]): NULL-terminated list of buffer types to use for tensors that match a pattern
829
-
n_gpu_layers (int): number of layers to store in VRAM
839
+
n_gpu_layers (int): number of layers to store in VRAM, a negative value means all layers
830
840
split_mode (int): how to split the model across multiple GPUs
841
+
load_mode (int): how to load the model
831
842
main_gpu (int): the GPU that is used for the entire model when split_mode is LLAMA_SPLIT_MODE_NONE
832
843
tensor_split (ctypes.Array[ctypes.ctypes.c_float]): proportion of the model (layers or rows) to offload to each GPU, size: llama_max_devices()
833
844
progress_callback (llama_progress_callback): called with a progress value between 0.0 and 1.0. Pass NULL to disable. If the provided progress_callback returns true, model loading continues. If it returns false, model loading is immediately aborted.
834
845
progress_callback_user_data (ctypes.ctypes.c_void_p): context pointer passed to the progress callback
835
846
kv_overrides (ctypes.Array[llama_model_kv_override]): override key-value pairs of the model meta data
836
847
vocab_only (bool): only load the vocabulary, no weights
837
-
use_mmap (bool): use mmap if possible
838
-
use_direct_io (bool): use direct io, takes precedence over use_mmap when supported
839
-
use_mlock (bool): force system to keep model in RAM
840
848
check_tensors (bool): validate model tensor data
841
849
use_extra_bufts (bool): use extra buffer types (used for weight repacking)
842
850
no_host (bool): bypass host buffer allowing extra buffers to be used
@@ -849,15 +857,13 @@ class llama_model_params(ctypes.Structure):
0 commit comments