llama.cpp

Author	SHA1	Message	Date
Bach Le	c9c74b4e3f	llama : add classifier-free guidance (#2135 ) * Initial implementation * Remove debug print * Restore signature of llama_init_from_gpt_params * Free guidance context * Make freeing of guidance_ctx conditional * Make Classifier-Free Guidance a sampling function * Correct typo. CFG already means context-free grammar. * Record sampling time in llama_sample_classifier_free_guidance * Shift all values by the max value before applying logsoftmax * Fix styling based on review	2023-07-11 19:18:43 +03:00
Jinwoo Jeong	3ec7e596b2	docker : add '--server' option (#2174 )	2023-07-11 19:12:35 +03:00
Chad Brewbaker	917831c63a	readme : fix zig build instructions (#2171 )	2023-07-11 19:03:06 +03:00
Howard Su	2347463201	Support using mmap when applying LoRA (#2095 ) * Support using mmap when applying LoRA * Fix Linux * Update comment to reflect the support lora with mmap	2023-07-11 22:37:01 +08:00
LostRuins	bbef28218f	Possible solution to allow K-quants on models with n_vocab!=32000 (#2148 ) * This allows LLAMA models that were previously incompatible with K quants to function mostly as normal. This happens when a model has a vocab != 32000, e.g 32001 which means it's not divisible by 256 or 64. Since the problematic dimensions only apply for `tok_embeddings.weight` and `output.weight` (dimentions 4096 x n_vocab), we can simply quantize these layers to Q8_0 whereas the majority of the hidden layers are still K-quanted since they have compatible dimensions. * Fix indentation Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * As an alternative, to avoid failing on Metal due to lack of Q8_0 support, instead quantize tok_embeddings.weight to Q4_0 and retain output.weight as F16. This results in a net gain of about 55mb for a 7B model compared to previous approach, but should minimize adverse impact to model quality. --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-07-11 22:01:08 +08:00
Concedo	a286776435	updated lite	2023-07-11 21:48:01 +08:00
Concedo	1d1111e10f	expose timing info in web api	2023-07-11 18:56:06 +08:00
Concedo	7222877069	Merge remote-tracking branch 'ren/concedo' into concedo_experimental	2023-07-11 18:45:36 +08:00
Concedo	5ca204d527	Merge remote-tracking branch 'yellowrose/pr/open/LostRuins/koboldcpp/multigpu-cuda-gui' into concedo_experimental # Conflicts: # koboldcpp.py	2023-07-11 18:22:54 +08:00
Concedo	4be167915a	added linear rope option, added warning for bad samplers	2023-07-11 18:08:19 +08:00
Concedo	b0b131499f	Merge branch 'master' into concedo_experimental # Conflicts: # .github/workflows/build.yml # CMakeLists.txt # Makefile # README.md # tests/test-tokenizer-0.cpp	2023-07-11 16:12:15 +08:00
Evan Miller	5656d10599	mpi : add support for distributed inference via MPI (#2099 ) * MPI support, first cut * fix warnings, update README * fixes * wrap includes * PR comments * Update CMakeLists.txt * Add GH workflow, fix test * Add info to README * mpi : trying to move more MPI stuff into ggml-mpi (WIP) (#2099) * mpi : add names for layer inputs + prep ggml_mpi_graph_compute() * mpi : move all MPI logic into ggml-mpi Not tested yet * mpi : various fixes - communication now works but results are wrong * mpi : fix output tensor after MPI compute (still not working) * mpi : fix inference * mpi : minor * Add OpenMPI to GH action * [mpi] continue-on-error: true * mpi : fix after master merge * [mpi] Link MPI C++ libraries to fix OpenMPI * tests : fix new llama_backend API * [mpi] use MPI_INT32_T * mpi : factor out recv / send in functions and reuse * mpi : extend API to allow usage with outer backends (e.g. Metal) --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-07-10 18:49:56 +03:00
Concedo	11ebfea8c0	Merge branch 'kquant_vocab_fix' into concedo_experimental	2023-07-10 23:28:48 +08:00
Concedo	fd9a2fdfe2	As an alternative, to avoid failing on Metal due to lack of Q8_0 support, instead quantize tok_embeddings.weight to Q4_0 and retain output.weight as F16. This results in a net gain of about 55mb for a 7B model compared to previous approach, but should minimize adverse impact to model quality.	2023-07-10 23:22:45 +08:00
LostRuins	048dca9809	Fix indentation Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-07-10 22:57:15 +08:00
Concedo	9324cb804a	reimplemented save and load	2023-07-10 22:49:27 +08:00
Concedo	50097e6c7f	Merge branch 'master' into concedo_experimental # Conflicts: # CMakeLists.txt # README.md # llama.cpp	2023-07-10 20:08:27 +08:00
Concedo	523fc3be52	fixed rwkv, standardized new ctx usage	2023-07-10 20:05:53 +08:00
Concedo	2827920044	fix compile errors, rwkv not working	2023-07-10 18:23:25 +08:00
YellowRoseCx	f1014f3cc7	remove unused .re	2023-07-10 00:26:40 -05:00
YellowRoseCx	242f01e983	Add Multi-GPU CuBLAS support in the new GUI	2023-07-09 17:10:14 -05:00
oobabooga	1d16309969	llama : remove "first token must be BOS" restriction (#2153 )	2023-07-09 11:59:53 +03:00
Nigel Bosch	db4047ad5c	main : escape prompt prefix/suffix (#2151 )	2023-07-09 11:56:18 +03:00
JackJollimore	18780e0a5e	readme : update Termux instructions (#2147 ) The file pathing is significant when running models inside of Termux on Android devices. llama.cpp performance is improved with loading a .bin from the $HOME directory.	2023-07-09 11:20:43 +03:00
clyang	3bbc1a11f0	ggml : fix buidling with Intel MKL but ask for "cblas.h" issue (#2104 ) (#2115 ) * Fix buidling with Intel MKL but ask for "cblas.h" issue * Use angle brackets to indicate the system library	2023-07-09 11:12:20 +03:00
rankaiyx	2492a53fd0	readme : add more docs indexes (#2127 ) * Update README.md to add more docs indexes * Update README.md to add more docs indexes	2023-07-09 10:38:42 +03:00
Johannes Gäßler	64639555ff	Fixed OpenLLaMA 3b CUDA mul_mat_vec_q (#2144 )	2023-07-08 20:01:44 +02:00
Concedo	15576bc865	Merge branch 'kquant_vocab_fix' into concedo_experimental # Conflicts: # .github/workflows/build.yml # Makefile # README.md # llama.cpp # tests/CMakeLists.txt # tests/test-grad0.c # tests/test-opt.c	2023-07-08 20:43:20 +08:00
Concedo	1854168841	This allows LLAMA models that were previously incompatible with K quants to function mostly as normal. This happens when a model has a vocab != 32000, e.g 32001 which means it's not divisible by 256 or 64. Since the problematic dimensions only apply for `tok_embeddings.weight` and `output.weight` (dimentions 4096 x n_vocab), we can simply quantize these layers to Q8_0 whereas the majority of the hidden layers are still K-quanted since they have compatible dimensions.	2023-07-08 20:38:03 +08:00
callMeMakerRen	4e46673f80	Merge branch 'LostRuins:concedo' into concedo	2023-07-08 09:33:26 +08:00
Johannes Gäßler	061f5f8d21	CUDA: add __restrict__ to mul mat vec kernels (#2140 )	2023-07-08 00:25:15 +02:00
dylan	84525e7962	docker : add support for CUDA in docker (#1461 ) Co-authored-by: canardleteer <eris.has.a.dad+github@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-07-07 21:25:25 +03:00
Georgi Gerganov	a7e20edf22	ci : switch threads to 1 (#2138 )	2023-07-07 21:23:57 +03:00
Qingyou Meng	1d656d6360	ggml : change ggml_graph_compute() API to not require context (#1999 ) * ggml_graph_compute: deprecate using ggml_context, try resolve issue #287 * rewrite: no longer consider backward compitability; plan and make_plan * minor: rename ctx as plan; const * remove ggml_graph_compute from tests/test-grad0.c, but current change breaks backward * add static ggml_graph_compute_sugar() * minor: update comments * reusable buffers * ggml : more consistent naming + metal fixes * ggml : fix docs * tests : disable grad / opt + minor naming changes * ggml : add ggml_graph_compute_with_ctx() - backwards compatible API - deduplicates a lot of copy-paste * ci : enable test-grad0 * examples : factor out plan allocation into a helper function * llama : factor out plan stuff into a helper function * ci : fix env * llama : fix duplicate symbols + refactor example benchmark * ggml : remove obsolete assert + refactor n_tasks section * ggml : fix indentation in switch * llama : avoid unnecessary bool * ggml : remove comments from source file and match order in header --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-07-07 19:24:01 +03:00
Concedo	8edcb337c6	added ability to select "all devices"	2023-07-07 23:37:55 +08:00
Georgi Gerganov	7242140283	ggml : remove sched_yield() call in ggml_graph_compute_thread() (#2134 )	2023-07-07 18:37:10 +03:00
Concedo	ddaa4f2a26	fix cuda garbage results and gpu selection issues	2023-07-07 22:14:14 +08:00
Aarni Koskela	3e08ae99ce	convert.py: add mapping for safetensors bf16 (#1598 ) Fixes #1473	2023-07-07 09:12:49 -04:00
Concedo	95eca51bef	add gpu choice for GUI for cuda	2023-07-07 18:39:47 +08:00
Concedo	a689a66068	make it work with pyinstaller	2023-07-07 17:52:34 +08:00
Concedo	9ee9a77f12	warn outdated GUI (+1 squashed commits) Squashed commits: [15aec3d] spelling error	2023-07-07 16:39:17 +08:00
Concedo	32102c2064	Merge branch 'master' into concedo_experimental # Conflicts: # README.md	2023-07-07 14:15:39 +08:00
shutup	894c72819c	Merge branch 'concedo' of https://github.com/callMeMakerRen/koboldcpp into concedo	2023-07-07 11:57:25 +08:00
shutup	1727e652f1	expose some useful info that can be used in statistics of performence	2023-07-07 11:52:58 +08:00
Howard Su	481f793acc	Fix opencl by wrap #if-else-endif with \n (#2086 )	2023-07-07 05:34:18 +02:00
Georgi Gerganov	dfd9fce6d6	ggml : fix restrict usage	2023-07-06 19:41:31 +03:00
Judd	36680f6e40	convert : update for baichuan (#2081 ) 1. guess n_layers; 2. relax warnings on context size; 3. add a note that its derivations are also supported. Co-authored-by: Judd <foldl@boxvest.com>	2023-07-06 19:23:49 +03:00
tslmy	a17a2683d8	alpaca.sh : update model file name (#2074 ) The original file name, `ggml-alpaca-7b-q4.bin`, implied the first-generation GGML. After the breaking changes (mentioned in https://github.com/ggerganov/llama.cpp/issues/382), `llama.cpp` requires GGML V3 now. Those model files are named `ggmlv3.bin`. We should change the example to an actually working model file, so that this thing is more likely to run out-of-the-box for more people, and less people would waste time downloading the old Alpaca model.	2023-07-06 19:17:50 +03:00
Concedo	8424a35c62	added the ability to ban any substring tokens	2023-07-06 23:24:21 +08:00
Concedo	27a0907cfa	backport MM256_SET_M128I to ggml_v2, updated lite, added support for selecting the GPU for cublas	2023-07-06 22:33:46 +08:00

1 2 3 4 5 ...

1544 commits