sampling : refactor init to use llama_sampling_params (#3696)

* sampling : refactor init to use llama_sampling_params * llama : combine repetition, frequency and presence penalties in 1 call * examples : remove embd-input and gptneox-wip * sampling : rename penalty params + reduce size of "prev" vector * sampling : add llama_sampling_print helper * sampling : hide prev behind API and apply #3661 ggml-ci
2023-10-20 21:07:23 +03:00 · 2023-10-20 21:07:23 +03:00 · d1031cf49c
commit d1031cf49c
parent 8cf19d60dc
30 changed files with 365 additions and 4502 deletions
--- a/README.md
+++ b/README.md
@ -962,7 +962,6 @@ docker run --gpus all -v /path/to/models:/models local/llama.cpp:light-cuda -m /

 - [main](./examples/main/README.md)
 - [server](./examples/server/README.md)
- [embd-input](./examples/embd-input/README.md)
 - [jeopardy](./examples/jeopardy/README.md)
 - [BLIS](./docs/BLIS.md)
 - [Performance troubleshooting](./docs/token_generation_performance_tips.md)