sampling : refactor + optimize penalties sampler (#10803)

* sampling : refactor + optimize penalties sampler ggml-ci * common : apply ignore_eos as logit bias ggml-ci * batched : remove penalties sampler * params : allow penalty_last_n == -1 to be equal to context size ggml-ci * common : by default, move the penalties at the end of the sampling chain ggml-ci * common : ignore all EOG tokens Co-authored-by: Diego Devesa <slarengh@gmail.com> * common : move back the penalties at the front of the sampling chain ggml-ci * readme : restore hint about --ignore-eos flag [no ci] * llama : minor ggml-ci * webui : update --------- Co-authored-by: Diego Devesa <slarengh@gmail.com>
2024-12-16 12:31:14 +02:00 · 2024-12-16 12:31:14 +02:00 · 644fd71b44
commit 644fd71b44
parent 4ddd199f6f
17 changed files with 111 additions and 152 deletions
--- a/examples/batched/batched.cpp
+++ b/examples/batched/batched.cpp
@ -65,6 +65,7 @@ int main(int argc, char ** argv) {
    llama_context * ctx = llama_new_context_with_model(model, ctx_params);

    auto sparams = llama_sampler_chain_default_params();
+    sparams.no_perf = false;

    llama_sampler * smpl = llama_sampler_chain_init(sparams);