llama-quantize command. --run is required to
execute it, and the output directory is created automatically. Install a
llama.cpp build exposing llama-quantize before running conversions.Create smaller GGUF models with a safe preview-first workflow.
llama-quantize command. --run is required to
execute it, and the output directory is created automatically. Install a
llama.cpp build exposing llama-quantize before running conversions.