CPU/CUDA: Gemma 2 FlashAttention support (#8542)

* CPU/CUDA: Gemma 2 FlashAttention support * apply logit_softcap to scale in kernel * disable logit softcapping tests on Metal * remove metal check
2024-08-24 21:34:59 +02:00 · 2024-08-24 21:34:59 +02:00 · e11bd856d5
commit e11bd856d5
parent 8f824ffe8e
12 changed files with 319 additions and 79 deletions
--- a/ggml/include/ggml.h
+++ b/ggml/include/ggml.h
@ -1760,7 +1760,8 @@ extern "C" {
            struct ggml_tensor  * v,
            struct ggml_tensor  * mask,
            float                 scale,
-            float                 max_bias);
+            float                 max_bias,
+            float                 logit_softcap);

    GGML_API void ggml_flash_attn_ext_set_prec(
            struct ggml_tensor * a,