Arm AArch64: update docs/build.md README to include compile time flags for buiilding the Q4_0_4_4 quant type

2024-07-09 19:19:24 +00:00 · 2024-07-09 19:19:24 +00:00 · c653eb1f1b
commit c653eb1f1b
parent 0e84ef1aa7
1 changed files with 2 additions and 0 deletions
--- a/docs/build.md
+++ b/docs/build.md
@ -28,6 +28,7 @@ In order to build llama.cpp you have four different options.
        ```

  - Notes:
+    - For `Q4_0_4_4` quantization type build, add the `GGML_NO_LLAMAFILE=1` flag. For example, use `make GGML_NO_LLAMAFILE=1`.
    - For faster compilation, add the `-j` argument to run multiple jobs in parallel. For example, `make -j 8` will run 8 jobs in parallel.
    - For faster repeated compilation, install [ccache](https://ccache.dev/).
    - For debug builds, run `make LLAMA_DEBUG=1`
@ -41,6 +42,7 @@ In order to build llama.cpp you have four different options.

  **Notes**:

+    - For `Q4_0_4_4` quantization type build, add the `-DGGML_LLAMAFILE=OFF` cmake option. For example, use `cmake -B build -DGGML_LLAMAFILE=OFF`.
    - For faster compilation, add the `-j` argument to run multiple jobs in parallel. For example, `cmake --build build --config Release -j 8` will run 8 jobs in parallel.
    - For faster repeated compilation, install [ccache](https://ccache.dev/).
    - For debug builds, there are two cases: