Common params
82print usage and exit
show version and build info
show list of models in cache
print source-able bash completion script for llama.cpp
number of CPU threads to use during generation
number of threads to use during batch and prompt processing
CPU affinity mask: arbitrarily long hex. Complements cpu-range
range of CPUs for affinity. Complements --cpu-mask
use strict CPU placement
set process/thread priority : low(-1), normal(0), medium(1), high(2), realtime(3)
use polling level to wait for work (0 - no polling, default: 50)
CPU affinity mask: arbitrarily long hex. Complements cpu-range-batch
ranges of CPUs for affinity. Complements --cpu-mask-batch
use strict CPU placement
set process/thread priority : 0-normal, 1-medium, 2-high, 3-realtime
use polling to wait for work
size of the prompt context
number of tokens to predict
logical maximum batch size
physical maximum batch size
number of tokens to keep from the initial prompt
use full-size SWA cache
set Flash Attention use ('on', 'off', or 'auto', default: 'auto')
whether to enable internal libllama performance timings
whether to process escapes sequences (\n, \r, \t, \', \", \\)
RoPE frequency scaling method, defaults to linear unless specified by the model
RoPE context scaling factor, expands context by a factor of N
RoPE base frequency, used by NTK-aware scaling
RoPE frequency scaling factor, expands context by a factor of 1/N
YaRN: original context size of model
YaRN: extrapolation mix factor
YaRN: scale sqrt(t) or attention magnitude
YaRN: high correction dim or alpha
YaRN: low correction dim or beta
whether to enable KV cache offloading
whether to enable weight repacking
bypass host buffer allowing extra buffers to be used
KV cache data type for K
- allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1
KV cache data type for V
- allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1
KV cache defragmentation threshold (DEPRECATED)
comma-separated list of RPC servers (host:port)
model loading mode
- - auto: mmap, unless a device does not support it
- - none: no special loading mode
- - mmap: memory-map model (if mmap disabled, slower load but may reduce pageouts if not using mlock)
- - mlock: force system to keep model in RAM rather than swapping or compressing
- - mmap+mlock: mmap + force system to keep model in RAM rather than swapping or compressing
- - dio: use DirectIO if available
on-demand reading of certain tensors, for example per-layer embeddings
- - on: read the rows of such tensors from disk on demand instead of keeping them resident (requires mmap)
- - auto: on, but only for tensors larger than 4 GiB
- - off: always keep them resident
attempt optimizations that help on some NUMA systems
- - distribute: spread execution evenly over all nodes
- - isolate: only spawn threads on CPUs on the node that execution started on
- - numactl: use the CPU map provided by numactl
- if run without this previously, it is recommended to drop the system page cache before using this
- see https://github.com/ggml-org/llama.cpp/issues/1437
comma-separated list of devices to use for offloading (none = don't offload)
- use --list-devices to see a list of available devices
print list of available devices and exit
override tensor buffer type
keep all Mixture of Experts (MoE) weights in the CPU
keep the Mixture of Experts (MoE) weights of the first N layers in the CPU
keep the dense FFN weights of the first N layers in the CPU
- (dense models; for MoE expert weights use --n-cpu-moe)
max. number of layers to store in VRAM, either an exact number, 'auto', or 'all'
how to split the model across multiple GPUs, one of:
- - none: use one GPU only
- - layer (default): split layers and KV across GPUs (pipelined)
- - row: split weight across GPUs by rows (parallelized)
- - tensor: split weights and KV across GPUs (parallelized, EXPERIMENTAL)
fraction of the model to offload to each GPU, comma-separated list of proportions, e.g. 3,1
the GPU to use for the model (with split-mode = none), or for intermediate results and KV (with split-mode = row)
whether to adjust unset arguments to fit in device memory ('on' or 'off', default: 'on')
target margin per device for --fit, comma-separated list of values, single value is broadcast across all devices, default: 1024
minimum ctx size that can be set by --fit option, default: 4096
check model tensor data for invalid values
advanced option to override model metadata by key. to specify multiple overrides, either use comma-separated values.
- types: int, float, bool, str. example: --override-kv tokenizer.ggml.add_bos_token=bool:false,tokenizer.ggml.add_eos_token=bool:false
whether to offload host tensor operations to device
path to LoRA adapter (use comma-separated values to load multiple adapters)
path to LoRA adapter with user defined scaling (format: FNAME:SCALE,...)
- note: use comma-separated values
add a control vector
- note: use comma-separated values to add multiple control vectors
add a control vector with user defined scaling SCALE
- note: use comma-separated values (format: FNAME:SCALE,...)
layer range to apply the control vector(s) to, start and end inclusive
model path to load
model download url
Docker Hub model repository. repo is optional, default to ai/. quant is optional, default to :latest.
- example: gemma3
Hugging Face model repository; quant is optional, case-insensitive, default to Q4_K_M, or falls back to the first file in the repo if Q4_K_M doesn't exist.
- mmproj is also downloaded automatically if available. to disable, add --no-mmproj
- example: ggml-org/GLM-4.7-Flash-GGUF:Q4_K_M
Hugging Face model file. If specified, it will override the quant in --hf-repo
Hugging Face access token
Log disable
Log to file
Log as JSONL (one JSON object per line) to stdout, this also disables colored logging
Set colored logging ('on', 'off', or 'auto', default: 'auto')
- 'auto' enables colors when output is to a terminal
Set verbosity level to infinity (i.e. log all messages, useful for debugging)
Offline mode: forces use of cache, prevents network access
Set the verbosity threshold. Messages with a higher verbosity will be ignored. Values:
- - 0: generic output
- - 1: error
- - 2: warning
- - 3: info
- - 4: trace (more info)
- - 5: debug
Enable prefix in log messages
Enable timestamps in log messages
KV cache data type for K for the draft model
- allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1
KV cache data type for V for the draft model
- allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1
Sampling params
34samplers that will be used for generation in the order, separated by ';'
RNG seed
simplified sequence for samplers that will be used
ignore end of stream token and continue generating (implies --logit-bias EOS-inf)
temperature
top-k sampling
top-p sampling
min-p sampling
top-n-sigma sampling
xtc probability
xtc threshold
locally typical sampling, parameter p
last n tokens to consider for penalize
penalize repeat sequence of tokens
repeat alpha presence penalty
repeat alpha frequency penalty
set DRY sampling multiplier
set DRY sampling base value
set allowed length for DRY sampling
set DRY penalty for the last n tokens
add sequence breaker for DRY sampling, clearing out default breakers ('\n', ':', '"', '*') in the process; use "none" to not use any sequence breakers
adaptive-p: select tokens near this probability (valid range 0.0 to 1.0; negative = disabled)
adaptive-p: decay rate for target adaptation over time. lower values are more reactive, higher values are more stable.
- (valid range 0.0 to 0.99) (default: 0.90)
dynamic temperature range
dynamic temperature exponent
use Mirostat sampling.
- Top K, Nucleus and Locally Typical samplers are ignored if used.
Mirostat learning rate, parameter eta
Mirostat target entropy, parameter tau
modifies the likelihood of token appearing in the completion,
- i.e. `--logit-bias 15043+1` to increase likelihood of token ' Hello',
- or `--logit-bias 15043-1` to decrease likelihood of token ' Hello'
BNF-like grammar to constrain generations (see samples in grammars/ dir)
file to read grammar from
JSON schema to constrain generations (https://json-schema.org/), e.g. `{"type": "object"}` for any JSON object
File containing a JSON schema to constrain generations (https://json-schema.org/), e.g. `{"type": "object"}` for any JSON object
enable backend sampling (experimental)
Server-specific params
139path to static lookup cache to use for lookup decoding (not updated by generation)
path to dynamic lookup cache to use for lookup decoding (updated by generation)
context limit per parallel slot.
- when set without -c/--ctx-size, the shared KV pool is sized to n_parallel*N
max number of context checkpoints to create per slot
minimum spacing between context checkpoints in tokens
set the maximum cache size in MiB
use single unified KV buffer shared across all sequences
save idle slots to the prompt cache on new task, and clear them when using unified KV
whether to use context shift on infinite text generation
halt generation at PROMPT, return control in interactive mode
special tokens output enabled
whether to perform warmup with an empty run
use Suffix/Prefix/Middle pattern for infill (instead of Prefix/Suffix/Middle) as some models prefer this.
pooling type for embeddings, use model default if unspecified
number of server slots
whether to enable continuous batching (a.k.a dynamic batching)
path to a multimodal projector file. see tools/mtmd/README.md
- note: if -hf is used, this argument can be omitted
URL to a multimodal projector file. see tools/mtmd/README.md
whether to use multimodal projector file (if available), useful when using -hf
whether to enable GPU offloading for multimodal projector
device to use for multimodal projector (none = don't offload, default: follows --device)
- use --list-devices to see a list of available devices
minimum number of tokens each image can take, only used by vision models with dynamic resolution
maximum number of tokens each image can take, only used by vision models with dynamic resolution
maximum number of image tokens per batch when encoding images
target video frame rate
interval in milliseconds between text timestamps
path to the directory containing ffmpeg and ffprobe
set model name aliases, comma-separated (to be used by API)
set model tags, comma-separated (informational, not used for routing)
normalisation for embeddings (-1=none, 0=max absolute int16, 1=taxicab, 2=euclidean, >2=p-norm)
IP addresses to listen on, comma-separated, or UNIX socket paths ending in .sock; with multiple TCP addresses, :: binds IPv6 only; overlapping addresses result in undefined behavior
port to listen
allow multiple sockets to bind to the same port
path to serve static files from
comma-separated list of allowed origins for CORS
- if set to special value 'localhost', reflect the Origin header only if it is localhost
comma-separated list of allowed methods for CORS
comma-separated list of allowed headers for CORS
whether to allow credentials for CORS
- note: if this is enabled and --cors-origins is set to * (default), the Origin header will be echoed back, and credentials will always be allowed
prefix path the server serves from, without the trailing slash
JSON that provides default UI settings (overrides UI defaults)
JSON file that provides default UI settings (overrides UI defaults)
experimental: whether to enable MCP CORS proxy - do not enable in untrusted environments
experimental: whether to enable built-in tools for AI agents - do not enable in untrusted environments
- specify "all" to enable all tools
- available tools: read_file, file_glob_search, grep_search, exec_shell_command, write_file, edit_file, get_info
- note: for security reasons, this will limit --cors-origins to localhost by default
experimental: run tools in a separate runtime environment
- available options:
- 'docker:<image>', 'podman:<image>': spin up a new container and reuse it for all invocations, clean up on server exit
- 'docker-container:<id>', 'podman-container:<id>': use an existing container by ID, won't stop on server exit
- 'ssh:<target>': run tools on a remote POSIX host over SSH, key-based auth and a trusted host key are required
experimental: path to JSON file with MCP server definitions (Cursor-compatible format) - do not enable in untrusted environments
- note: for security reasons, this will limit --cors-origins to localhost by default
experimental: inline JSON with MCP server definitions (Cursor-compatible format) - do not enable in untrusted environments
- note: for security reasons, this will limit --cors-origins to localhost by default
whether to enable CORS proxy and all built-in tools - do not enable in untrusted environments
- note: for security reasons, this will limit --cors-origins to localhost by default
whether to enable the Web UI
restrict to only support embedding use case; use only with dedicated embedding models
enable reranking endpoint on server
API key to use for authentication, multiple keys can be provided as a comma-separated list
path to file containing API keys, one per line; lines starting with a hash are treated as comments
path to file a PEM-encoded SSL private key
path to file a PEM-encoded SSL certificate
sets additional params for the json template parser, must be a valid json object string, e.g. '{"key1":"value1","key2":"value2"}'
server read/write timeout in seconds
server SSE ping interval in seconds (-1 = disabled, default: 30)
number of threads used to process HTTP requests
whether to enable prompt caching
min chunk size to attempt reusing from the cache via KV shifting, requires prompt caching to be enabled
enable prometheus compatible metrics endpoint
enable changing global properties via POST /props
expose slots monitoring endpoint
path to save slot kv cache
directory for loading local media files; files can be accessed via file:// URLs using relative paths
directory containing models for the router server
path to INI file containing model presets for the router server
for router server, maximum number of models to load simultaneously
for router server, whether to automatically load models
whether to use jinja template engine for chat
controls whether thought tags are allowed and/or extracted from the response, and in which format they're returned; one of:
- - none: leaves thoughts unparsed in `message.content`
- - deepseek: puts thoughts in `message.reasoning_content`
- - deepseek-legacy: keeps `<think>` tags in `message.content` while also populating `message.reasoning_content`
Use reasoning/thinking in the chat ('on', 'off', or 'auto', default: 'auto' (detect from template))
reasoning effort level given to the chat template: 'default' to keep the template default,
- or a level such as 'minimal', 'low', 'medium', 'high', 'xhigh' or 'max' (default: default)
token budget for thinking: -1 for unrestricted, 0 for immediate end, N>0 for token budget
message injected before the end-of-thinking tag when reasoning budget is exhausted
preserve reasoning trace in the full history, not just the last assistant message
- compatible with certain templates having 'supports_preserve_reasoning' capability
- example: https://docs.z.ai/guides/capabilities/thinking-mode#preserved-thinking
set custom jinja chat template
- if suffix/prefix are specified, template will be disabled
- only commonly used templates are accepted (unless --jinja is set before this flag):
- list of built-in templates:
- bailing, bailing-think, bailing2, chatglm3, chatglm4, chatml, command-r, deepseek, deepseek-ocr, deepseek2, deepseek3, exaone-moe, exaone3, exaone4, falcon3, gemma, gigachat, glmedge, gpt-oss, granite, granite-4.0, granite-4.1, grok-2, hunyuan-dense, hunyuan-moe, hunyuan-vl, kimi-k2, llama2, llama2-sys, llama2-sys-bos, llama2-sys-strip, llama3, llama4, megrez, minicpm, mistral-v1, mistral-v3, mistral-v3-tekken, mistral-v7, mistral-v7-tekken, monarch, openchat, orion, pangu-embedded, phi3, phi4, rwkv-world, seed_oss, smolvlm, solar-open, vicuna, vicuna-orca, yandex, zephyr
set custom jinja chat template file
- if suffix/prefix are specified, template will be disabled
- only commonly used templates are accepted (unless --jinja is set before this flag):
- list of built-in templates:
- bailing, bailing-think, bailing2, chatglm3, chatglm4, chatml, command-r, deepseek, deepseek-ocr, deepseek2, deepseek3, exaone-moe, exaone3, exaone4, falcon3, gemma, gigachat, glmedge, gpt-oss, granite, granite-4.0, granite-4.1, grok-2, hunyuan-dense, hunyuan-moe, hunyuan-vl, kimi-k2, llama2, llama2-sys, llama2-sys-bos, llama2-sys-strip, llama3, llama4, megrez, minicpm, mistral-v1, mistral-v3, mistral-v3-tekken, mistral-v7, mistral-v7-tekken, monarch, openchat, orion, pangu-embedded, phi3, phi4, rwkv-world, seed_oss, smolvlm, solar-open, vicuna, vicuna-orca, yandex, zephyr
force a pure content parser, even if a Jinja template is specified; model will output everything in the content section, including any reasoning and/or tool calls
whether to prefill the assistant's response if the last message is an assistant message
- when this flag is set, if the last message is an assistant message then it will be treated as a full message and not prefilled
how much the prompt of a request must match the prompt of a slot in order to use that slot
load LoRA adapters without applying them (apply later via POST /lora-adapters)
number of seconds of idleness after which the server will sleep
Log prompts to directory (auto-created if not present; only used for debugging, default: disabled)
Same as --hf-repo, but for the draft model
number of threads to use during generation
number of threads to use during batch and prompt processing
Draft model CPU affinity mask. Complements cpu-range-draft
Ranges of CPUs for affinity. Complements --cpu-mask-draft
Use strict CPU placement for draft model
set draft process/thread priority : 0-normal, 1-medium, 2-high, 3-realtime
Use polling to wait for draft model work
Draft model CPU affinity mask. Complements cpu-range-draft
Use strict CPU placement for draft model
set draft process/thread priority : 0-normal, 1-medium, 2-high, 3-realtime
Use polling to wait for draft model work
override tensor buffer type for draft model
keep all Mixture of Experts (MoE) weights in the CPU for the draft model
keep the Mixture of Experts (MoE) weights of the first N layers in the CPU for the draft model
number of tokens to draft for speculative decoding
minimum number of draft tokens to use for speculative decoding
target mean synthetic acceptance length, including the target token (benchmarking only)
comma-separated unconditional per-position synthetic acceptance probabilities (benchmarking only)
speculative decoding split probability
minimum speculative decoding probability (greedy)
offload draft sampling to the backend
comma-separated list of devices to use for offloading the draft model (none = don't offload, default: follows --device)
- use --list-devices to see a list of available devices
max. number of draft model layers to store in VRAM, either an exact number, 'auto', or 'all'
draft model for speculative decoding
comma-separated list of types of speculative decoding to use
minimum number of ngram tokens to use for ngram-based speculative decoding
maximum number of ngram tokens to use for ngram-based speculative decoding
ngram-mod lookup length
ngram size N for ngram-simple speculative decoding, length of lookup n-gram
ngram size M for ngram-simple speculative decoding, length of draft m-gram
minimum hits for ngram-simple speculative decoding
ngram size N for ngram-map-k speculative decoding, length of lookup n-gram
ngram size M for ngram-map-k speculative decoding, length of draft m-gram
minimum hits for ngram-map-k speculative decoding
ngram size N for ngram-map-k4v speculative decoding, length of lookup n-gram
ngram size M for ngram-map-k4v speculative decoding, length of draft m-gram
minimum hits for ngram-map-k4v speculative decoding
the argument has been removed. use --spec-draft-n-max or --spec-ngram-mod-n-max
the argument has been removed. use --spec-draft-n-min or --spec-ngram-mod-n-min
the argument has been removed. use the respective --spec-ngram-*-size-n or --spec-ngram-mod-n-match
the argument has been removed. use the respective --spec-ngram-*-size-m
the argument has been removed. use the respective --spec-ngram-*-min-hits
use default EmbeddingGemma model (note: can download weights from the internet)
use default Qwen 2.5 Coder 1.5B (note: can download weights from the internet)
use default Qwen 2.5 Coder 3B (note: can download weights from the internet)
use default Qwen 2.5 Coder 7B (note: can download weights from the internet)
use Qwen 2.5 Coder 7B + 0.5B draft for speculative decoding (note: can download weights from the internet)
use Qwen 2.5 Coder 14B + 0.5B draft for speculative decoding (note: can download weights from the internet)
use default Qwen 3 Coder 30B A3B Instruct (note: can download weights from the internet)
use gpt-oss-20b (note: can download weights from the internet)
use gpt-oss-120b (note: can download weights from the internet)
use Gemma 3 4B QAT (note: can download weights from the internet)
use Gemma 3 12B QAT (note: can download weights from the internet)
enable default speculative decoding config