update readme sycl for new update

2024-03-19 09:54:39 +08:00 · 2024-03-19 09:54:39 +08:00 · 84da0f553e
commit 84da0f553e
parent 2d15886bb0
2 changed files with 80 additions and 33 deletions
--- a/README-sycl.md
+++ b/README-sycl.md
@ -29,6 +29,7 @@ For Intel CPU, recommend to use llama.cpp for X86 (Intel MKL building).
 ## News
 - 2024.3
  - New base line is ready: tag b2437.
  - Support multiple cards: **--split-mode**: [none|layer]; not support [row], it's on developing.
  - Support to assign main GPU by **--main-gpu**, replace $GGML_SYCL_DEVICE.
  - Support detecting all GPUs with level-zero and same top **Max compute units**.
@ -81,7 +82,7 @@ For dGPU, please make sure the device memory is enough. For llama-2-7b.Q4_0, rec
 |-|-|-|
 |Ampere Series| Support| A100|
-### oneMKL
+### oneMKL for CUDA
 The current oneMKL release does not contain the oneMKL cuBlas backend.
 As a result for Nvidia GPU's oneMKL must be built from source.
@ -254,16 +255,16 @@ Run without parameter:
 Check the ID in startup log, like:
 ```
-found 4 SYCL devices:
+found 6 SYCL devices:
-  Device 0: Intel(R) Arc(TM) A770 Graphics,	compute capability 1.3,
+|  |                  |                                             |Compute   |Max compute|Max work|Max sub|               |
-    max compute_units 512,	max work group size 1024,	max sub group size 32,	global mem size 16225243136
+|ID|       Device Type|                                         Name|capability|units      |group   |group  |Global mem size|
-  Device 1: Intel(R) FPGA Emulation Device,	compute capability 1.2,
+|--|------------------|---------------------------------------------|----------|-----------|--------|-------|---------------|
-    max compute_units 24,	max work group size 67108864,	max sub group size 64,	global mem size 67065057280
+| 0|[level_zero:gpu:0]|               Intel(R) Arc(TM) A770 Graphics|       1.3|        512|    1024|     32|    16225243136|
-  Device 2: 13th Gen Intel(R) Core(TM) i7-13700K,	compute capability 3.0,
+| 1|[level_zero:gpu:1]|                    Intel(R) UHD Graphics 770|       1.3|         32|     512|     32|    53651849216|
-    max compute_units 24,	max work group size 8192,	max sub group size 64,	global mem size 67065057280
+| 2|    [opencl:gpu:0]|               Intel(R) Arc(TM) A770 Graphics|       3.0|        512|    1024|     32|    16225243136|
-  Device 3: Intel(R) Arc(TM) A770 Graphics,	compute capability 3.0,
+| 3|    [opencl:gpu:1]|                    Intel(R) UHD Graphics 770|       3.0|         32|     512|     32|    53651849216|
-    max compute_units 512,	max work group size 1024,	max sub group size 32,	global mem size 16225243136
+| 4|    [opencl:cpu:0]|         13th Gen Intel(R) Core(TM) i7-13700K|       3.0|         24|    8192|     64|    67064815616|
-
+| 5|    [opencl:acc:0]|               Intel(R) FPGA Emulation Device|       1.2|         24|67108864|     64|    67064815616|
 ```
 |Attribute|Note|
@ -271,12 +272,35 @@ found 4 SYCL devices:
 |compute capability 1.3|Level-zero running time, recommended |
 |compute capability 3.0|OpenCL running time, slower than level-zero in most cases|
-4. Set device ID and execute llama.cpp
+4. Device selection and execute llama.cpp
-Set device ID = 0 by **GGML_SYCL_DEVICE=0**
+There are two modes to select device:
 - Single device: Use one device assigned by user.
 - Multiple devices: Detect and choose all devices which have same top Max compute units automatically.
 |Device selection|Parameter|
 |-|-|
 |Single device|--split-mode none --main-gpu DEVICE_ID |
 |Multiple devices|--split-mode layer (default)|
 Examples:
 - Use device 0:
 ```sh
-GGML_SYCL_DEVICE=0 ./build/bin/main -m models/llama-2-7b.Q4_0.gguf -p "Building a website can be done in 10 simple steps:" -n 400 -e -ngl 33
+ZES_ENABLE_SYSMAN=1 ./build/bin/main -m models/llama-2-7b.Q4_0.gguf -p "Building a website can be done in 10 simple steps:" -n 400 -e -ngl 33 -sm none -mg 0
 ```
 or run by script:
 ```sh
 ./examples/sycl/run_llama2.sh 0
 ```
 - Use multiple devices:
 ```sh
 ZES_ENABLE_SYSMAN=1 ./build/bin/main -m models/llama-2-7b.Q4_0.gguf -p "Building a website can be done in 10 simple steps:" -n 400 -e -ngl 33 -sm layer
 ```
 or run by script:
@ -293,7 +317,9 @@ Note:
 Like:
 ```
-Using device **0** (Intel(R) Arc(TM) A770 Graphics) as main device
+detect 1 SYCL GPUs: [0] with top Max compute units:512
 or
 use 1 SYCL GPUs: [0] with Max compute units:512
 ```
 ## Windows
@ -430,15 +456,16 @@ build\bin\main.exe
 Check the ID in startup log, like:
 ```
-found 4 SYCL devices:
+found 6 SYCL devices:
-  Device 0: Intel(R) Arc(TM) A770 Graphics,	compute capability 1.3,
+|  |                  |                                             |Compute   |Max compute|Max work|Max sub|               |
-    max compute_units 512,	max work group size 1024,	max sub group size 32,	global mem size 16225243136
+|ID|       Device Type|                                         Name|capability|units      |group   |group  |Global mem size|
-  Device 1: Intel(R) FPGA Emulation Device,	compute capability 1.2,
+|--|------------------|---------------------------------------------|----------|-----------|--------|-------|---------------|
-    max compute_units 24,	max work group size 67108864,	max sub group size 64,	global mem size 67065057280
+| 0|[level_zero:gpu:0]|               Intel(R) Arc(TM) A770 Graphics|       1.3|        512|    1024|     32|    16225243136|
-  Device 2: 13th Gen Intel(R) Core(TM) i7-13700K,	compute capability 3.0,
+| 1|[level_zero:gpu:1]|                    Intel(R) UHD Graphics 770|       1.3|         32|     512|     32|    53651849216|
-    max compute_units 24,	max work group size 8192,	max sub group size 64,	global mem size 67065057280
+| 2|    [opencl:gpu:0]|               Intel(R) Arc(TM) A770 Graphics|       3.0|        512|    1024|     32|    16225243136|
-  Device 3: Intel(R) Arc(TM) A770 Graphics,	compute capability 3.0,
+| 3|    [opencl:gpu:1]|                    Intel(R) UHD Graphics 770|       3.0|         32|     512|     32|    53651849216|
-    max compute_units 512,	max work group size 1024,	max sub group size 32,	global mem size 16225243136
+| 4|    [opencl:cpu:0]|         13th Gen Intel(R) Core(TM) i7-13700K|       3.0|         24|    8192|     64|    67064815616|
 | 5|    [opencl:acc:0]|               Intel(R) FPGA Emulation Device|       1.2|         24|67108864|     64|    67064815616|
 ```
@ -447,13 +474,31 @@ found 4 SYCL devices:
 |compute capability 1.3|Level-zero running time, recommended |
 |compute capability 3.0|OpenCL running time, slower than level-zero in most cases|
 4. Set device ID and execute llama.cpp
-Set device ID = 0 by **set GGML_SYCL_DEVICE=0**
+4. Device selection and execute llama.cpp
 There are two modes to select device:
 - Single device: Use one device assigned by user.
 - Multiple devices: Detect and choose all devices which have same top Max compute units automatically.
 |Device selection|Parameter|
 |-|-|
 |Single device|--split-mode none --main-gpu DEVICE_ID |
 |Multiple devices|--split-mode layer (default)|
 Examples:
 - Use device 0:
 ```
-set GGML_SYCL_DEVICE=0
+build\bin\main.exe -m models\llama-2-7b.Q4_0.gguf -p "Building a website can be done in 10 simple steps:\nStep 1:" -n 400 -e -ngl 33 -s 0 -sm none -mg 0
-build\bin\main.exe -m models\llama-2-7b.Q4_0.gguf -p "Building a website can be done in 10 simple steps:\nStep 1:" -n 400 -e -ngl 33 -s 0
+```
 - Use multiple devices:
 ```
 build\bin\main.exe -m models\llama-2-7b.Q4_0.gguf -p "Building a website can be done in 10 simple steps:\nStep 1:" -n 400 -e -ngl 33 -s 0 -sm layer
 ```
 or run by script:
@ -470,7 +515,9 @@ Note:
 Like:
 ```
-Using device **0** (Intel(R) Arc(TM) A770 Graphics) as main device
+detect 1 SYCL GPUs: [0] with top Max compute units:512
 or
 use 1 SYCL GPUs: [0] with Max compute units:512
 ```
 ## Environment Variable
@ -489,7 +536,6 @@ Using device **0** (Intel(R) Arc(TM) A770 Graphics) as main device
 |Name|Value|Function|
 |-|-|-|
 |GGML_SYCL_DEVICE|0 (default) or 1|Set the device id used. Check the device ids by default running output|
 |GGML_SYCL_DEBUG|0 (default) or 1|Enable log function by macro: GGML_SYCL_DEBUG|
 |ZES_ENABLE_SYSMAN| 0 (default) or 1|Support to get free memory of GPU by sycl::aspect::ext_intel_free_memory.<br>Recommended to use when --split-mode = layer|
@ -507,6 +553,9 @@ Using device **0** (Intel(R) Arc(TM) A770 Graphics) as main device
 ## Q&A
 Note: please add prefix **[SYCL]** in issue title, so that we will check it as soon.
 - Error:  `error while loading shared libraries: libsycl.so.7: cannot open shared object file: No such file or directory`.
  Miss to enable oneAPI running environment.
@ -538,4 +587,4 @@ Using device **0** (Intel(R) Arc(TM) A770 Graphics) as main device
 ## Todo
- Support multiple cards.
+- NA
--- a/examples/sycl/win-run-llama2.bat
+++ b/examples/sycl/win-run-llama2.bat
@ -6,8 +6,6 @@ set INPUT2="Building a website can be done in 10 simple steps:\nStep 1:"
@call "C:\Program Files (x86)\Intel\oneAPI\setvars.bat" intel64 --force
 set GGML_SYCL_DEVICE=0
 rem set GGML_SYCL_DEBUG=1
 .\build\bin\main.exe -m models\llama-2-7b.Q4_0.gguf -p %INPUT2% -n 400 -e -ngl 33 -s 0