llama.cpp

Author	SHA1	Message	Date
Concedo	b13a768813	added softprompt endpoint	2023-03-25 10:12:47 +08:00
Concedo	8e383f1895	gitignore	2023-03-24 23:02:25 +08:00
LostRuins	1c78ffb964	Update README.md	2023-03-24 22:45:54 +08:00
Concedo	e791827973	added a GUI for selection of models if none was passed in through command line.	2023-03-24 22:03:57 +08:00
Concedo	c6c60332a4	Optimizations	2023-03-24 21:33:53 +08:00
Concedo	3879d84400	Merge branch 'master' into concedo # Conflicts: # .devops/tools.sh # CMakeLists.txt # README.md # flake.nix	2023-03-24 19:28:27 +08:00
Concedo	706e19e9b4	added ability to fast forward in time through partially duplicated prompts	2023-03-24 18:50:16 +08:00
Georgi Gerganov	b6b268d441	Add link to Roadmap discussion	2023-03-24 09:13:35 +02:00
Georgi Gerganov	3cd8dde0d1	Revert "Fix memory allocation issues and seg faults" This reverts commit `4870e455b3`. Will provide the correct fix later	2023-03-24 06:22:28 +02:00
Georgi Gerganov	4870e455b3	Fix memory allocation issues and seg faults	2023-03-24 00:11:53 +02:00
Georgi Gerganov	483bab2e3d	Avoid the transposed X branch in the Z = X * Y matrix multiplication (#439 ) Should make results reproducible for different number of threads and batch sizes	2023-03-23 23:22:01 +02:00
Jed Fox	404e1da38e	Fix quantize script not finding models in parent directory (#428 )	2023-03-23 22:42:52 +02:00
Georgi Gerganov	4cc053b6d5	Remove oboslete command from Docker script	2023-03-23 22:39:44 +02:00
Georgi Gerganov	0ba5a3a9a5	Obsolete	2023-03-23 22:32:21 +02:00
rabidcopy	2e17dfd80a	Replace EOS with newline to prevent context/memory being flushed by EOS in interactive mode (#333 ) * Improve interactive mode's coherence after EOS Aims to improve coherence and ability to resume the interactive session when the user is given input back after an end of text token is reached. Not sure what token 13 is or why it seems to help. See conversation for examples. * Make newline token a constant * dynamically determine newline token * relocate previous newline token const * cleanup whitespace * print a new line on end of text in interactive this may need to be looked into further when not using a reverse prompt * only print manual newline with reverse prompt fix formatting of reverse prompts so they don't end up at the end of the current line while not introducing unnecessary new lines otherwise * alternate approach to replace end of text tokens * Inject the reverse prompt again after eos in interactive mode * tokenize reverse prompt when needed makes this PR compatible with https://github.com/ggerganov/llama.cpp/pull/330 * tokenize and inject only first reverse prompt thanks to tjohnman * tokenize first reverse prompt once * add newline token * add newline token * tokenize/inject reverse prompt for refactor this doesn't seem right though * tokenize nothing for antiprompt if no reverse * Update main.cpp * Update main.cpp * tokenize and inject reverse prompt as needed this doesn't seem to work if the reverse prompt is tokenized outside earlier on * not needed * remove newline token * remove newline token * tokenize newline token * add space to comment * Update main.cpp Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: Slaren <2141330+slaren@users.noreply.github.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-03-23 22:22:47 +02:00
Timmy Knight	20a1a4e09c	Fix GPTQ converter (#423 ) * Fix GPTQ converter * Fix comment --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-03-23 22:18:13 +02:00
nusu-github	ad072fc5ad	Generate library with CMake (#430 ) * Generate library with CMake BUILD_SHARED_LIBS to allow llama library to be generated. * Turn ON PIC when BUILD_SHARED_LIBS is ON	2023-03-23 21:16:48 +01:00
anzz1	ea10d3ded2	Command line args bounds checking (#424 ) * command line args bounds checking * unknown and invalid param exit codes 0 -> 1	2023-03-23 19:54:28 +02:00
Ben Siraphob	a18c19259a	Fix Nix build	2023-03-23 17:51:26 +01:00
Concedo	1166fda943	Merge branch 'master' into concedo # Conflicts: # .github/workflows/build.yml # CMakeLists.txt # Makefile # README.md	2023-03-23 23:51:07 +08:00
Stephan Walter	a50e39c6fe	Revert "Delete SHA256SUMS for now" (#429 ) * Revert "Delete SHA256SUMS for now (#416)" This reverts commit `8eea5ae0e5`. * Remove ggml files until they can be verified * Remove alpaca json * Add also model/tokenizer.model to SHA256SUMS + update README --------- Co-authored-by: Pavol Rusnak <pavol@rusnak.io>	2023-03-23 15:15:48 +01:00
Kerfuffle	a140219e81	Fix Makefile echo escape codes (by removing them). (#418 )	2023-03-23 12:41:32 +01:00
Gary Mulder	8a3e5ef801	Move model section from issue template to README.md (#421 ) * Update custom.md * Removed Model section as it is better placed in README.md * Updates to README.md model section * Inserted text that was removed from issue template about obtaining models from FB and links to papers describing the various models * Removed IPF down links for the Alpaca 7B models as these look to be in the old data format and probably shouldn't be directly linked to, anyway * Updated the perplexity section to point at Perplexity scores #406 discussion	2023-03-23 11:30:40 +00:00
anzz1	8eea5ae0e5	Delete SHA256SUMS for now (#416 ) Delete this for now to avoid confusion since it contains some wrong checksums from the old tokenizer format Re-add after #374 is resolved	2023-03-23 11:26:19 +01:00
Georgi Gerganov	93208cfb92	Adjust repetition penalty ..	2023-03-23 10:46:58 +02:00
LostRuins	47ea33ab59	Update README.md	2023-03-23 16:02:19 +08:00
Georgi Gerganov	03ace14cfd	Add link to recent podcast about whisper.cpp and llama.cpp	2023-03-23 09:48:51 +02:00
anzz1	e4412b45e3	CI: CMake: Separate build and test steps (#376 ) * CI: Separate Build and Test steps (CMake) * CI: Make sure build passes before running tests (CMake) * CI: Standardise step id names	2023-03-23 04:20:34 +02:00
tjohnman	f7dc43bc0d	Fix instruct mode broken by PR #354 (#409 ) Co-authored-by: Johnman <tjohnman@github>	2023-03-23 01:30:23 +01:00
Gary Mulder	ee8a788786	Update issue template so people will use it (#404 )	2023-03-22 19:06:18 +00:00
Stephan Walter	69c92298a9	Deduplicate q4 quantization functions (#383 ) * Deduplicate q4 quantization functions * Use const; add basic test * Re-enable quantization test * Disable AVX2 flags in CI --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-03-22 19:29:06 +02:00
Valentyn Bezshapkin	97940520e8	fix: add POSIX functionality for Linux compilation (#51 ) * fix: add POSIX functionality for Linux compilation * fix: older standard for compatibility	2023-03-22 19:20:25 +02:00
tjohnman	305ba6f0e6	Don't force immediate interactive without `-i` (#354 ) * Don't force immediate interactive without -i Sometimes we might want to use a reverse prompt but we want to let the model generate tokens right after the initial prompt. So we don't force user input mode if the -i flag wasn't specified and instead let it run until we encounter the reverse prompt. This gives use some more flexibility, since it doesn't force the user to enter a newline if they want to let the model generate text right after the initial prompt and only be asked for input if the reverse prompt is encountered. The `--interactive-first` flag is reintroduced to force the old behavior. `-r` behaves like `-i` plus introduces a reverse prompt (it can be specified more than once). * Update help output. --------- Co-authored-by: Johnman <tjohnman@github>	2023-03-22 19:16:35 +02:00
Erik Scholz	4122dffff9	cmake: make llama an actual library (#392 )	2023-03-22 18:37:10 +02:00
Erik Scholz	56e659a0b2	fix perplexity after c-api refactor (#390 ) * preallocate a buffer of fitting size for tokenization (utils.cpp) * don't create a new std::string (especially here, where it's usually large)	2023-03-22 18:09:38 +02:00
Gary Linscott	40ea807a97	Add details on perplexity to README.md (#395 )	2023-03-22 08:53:54 -07:00
LostRuins	c5c1c8d5ce	Update README.md	2023-03-22 22:54:27 +08:00
Concedo	4ff58f73e5	Merge branch 'master' into concedo	2023-03-22 22:32:11 +08:00
Concedo	86c7457e24	Merge branch 'master' into concedo # Conflicts: # .github/workflows/build.yml # CMakeLists.txt # Makefile # README.md # main.cpp	2023-03-22 22:31:45 +08:00
Yusuf Kağan Hanoğlu	d5850c53ca	Add missing header for memcpy (#386 ) fixed: memcpy is not defined	2023-03-22 10:55:45 +02:00
Concedo	5c475503ce	resize image	2023-03-22 16:21:40 +08:00
Concedo	4e95e7f87f	Updated readme	2023-03-22 16:20:37 +08:00
Concedo	5f142df76e	dynamic max context size defaulting to 1024, also implemented the basic API as a fallback	2023-03-22 15:56:47 +08:00
Georgi Gerganov	ae44e23ee3	When seed <= 0 - use the clock to generate one	2023-03-22 07:47:15 +02:00
Georgi Gerganov	928480ef5b	Init llama_context_params properly from CLI (#370 )	2023-03-22 07:45:14 +02:00
Georgi Gerganov	56817b1f88	Remove temporary notice and update hot topics	2023-03-22 07:34:02 +02:00
Georgi Gerganov	f5a77a629b	Introduce C-style API (#370 ) * Major refactoring - introduce C-style API * Clean up * Add <cassert> * Add <iterator> * Add <algorithm> .... * Fix timing reporting and accumulation * Measure eval time only for single-token calls * Change llama_tokenize return meaning	2023-03-22 07:32:36 +02:00
Gary Mulder	da0e9fe90c	Add SHA256SUMS file and instructions to README how to obtain and verify the downloads Hashes created using: sha256sum models/B/.pth models/[7136]B/ggml-model-f16.bin models/[7136]B/ggml-model-q4_0.bin > SHA256SUMS	2023-03-21 23:19:11 +01:00
anzz1	e6c9e0986c	Fix bin dir for win ci	2023-03-22 00:01:08 +02:00
Erik Scholz	01a297b099	specify build type for ctest on windows (#371 )	2023-03-21 23:34:25 +02:00

1 2 3 4 5

216 commits