llama.cpp

Author	SHA1	Message	Date
Jiri Podivin	6035abe170	Setting new target for test binaries Programs in the tests directory are now build with target tests and placed in the same location. * clean target was expanded to remove new binaries * test target binaries are listed in a variable * Locations of binaries were added to the .gitignore Signed-off-by: Jiri Podivin <jpodivin@gmail.com>	2023-07-16 17:46:41 +02:00
Xiao-Yong Jin	6e7cca4047	llama : add custom RoPE (#2054 ) * Implement customizable RoPE The original RoPE has pre-defined parameters theta_i = 10000^(−2(i−1)/d), for i in [1, 2, ..., d/2] Our customizable RoPE, ggml_rope_custom_inplace, uses theta_i = scale * base^(−2(i−1)/d), for i in [1, 2, ..., d/2] with the default matches the original scale = 1.0 base = 10000 The new command line arguments --rope-freq-base --rope-freq-scale set the two new RoPE parameter. Recent researches show changing these two parameters extends the context limit with minimal loss. 1. Extending Context to 8K kaiokendev https://kaiokendev.github.io/til#extending-context-to-8k 2. Extending Context Window of Large Language Models via Positional Interpolation Shouyuan Chen, Sherman Wong, Liangjian Chen, Yuandong Tian https://arxiv.org/abs/2306.15595 3. NTK-Aware Scaled RoPE allows LLaMA models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation. https://www.reddit.com/user/bloc97 https://www.reddit.com/r/LocalLLaMA/comments/14lz7j5/ntkaware_scaled_rope_allows_llama_models_to_have/ For the bold, try adding the following command line parameters to your favorite model: -c 16384 --rope-freq-base 80000 --rope-freq-scale 0.5 * ggml-metal: fix custom rope * common: fix argument names in help * llama: increase MEM_REQ_EVAL for MODEL_3B It avoids crashing for quantized weights on CPU. Better ways to calculate the required buffer size would be better. * llama: make MEM_REQ_EVAL depend on n_ctx * server: use proper Content-Type in curl examples Without the header Content-Type: application/json, curl will POST with Content-Type: application/x-www-form-urlencoded Though our simple server doesn't care, the httplib.h used has a limit with CPPHTTPLIB_FORM_URL_ENCODED_PAYLOAD_MAX_LENGTH 8192 With Content-Type: application/json, we can send large json data. * style : minor fixes, mostly indentations * ggml : fix asserts --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-07-15 13:34:16 +03:00
Dave Della Costa	a6803cab94	flake : add runHook preInstall/postInstall to installPhase so hooks function (#2224 )	2023-07-14 22:13:38 +03:00
wzy	7dabc66f3c	make : use pkg-config for OpenBLAS (#2222 )	2023-07-14 22:05:08 +03:00
Bach Le	7cdd30bf1f	cuda : allocate all temporary ggml_tensor_extra_gpu from a fixed-size buffer (#2220 )	2023-07-14 22:00:58 +03:00
Evan Miller	e8035f141e	ggml : fix static_assert with older compilers #2024 (#2218 )	2023-07-14 21:55:56 +03:00
Bach Le	7513b7b0a1	llama : add functions that work directly on model (#2197 ) * Remove vocab reference from context * Add functions that works directly with model	2023-07-14 21:55:24 +03:00
Ali Chraghi	de8342423d	build.zig : install config header (#2216 )	2023-07-14 21:50:58 +03:00
Shangning Xu	c48c525f87	examples : fixed path typos in embd-input (#2214 )	2023-07-14 21:40:05 +03:00
Jiahao Li	206e01de11	cuda : support broadcast add & mul (#2192 ) Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-07-14 21:38:24 +03:00
Johannes Gäßler	4304bd3cde	CUDA: mul_mat_vec_q kernels for k-quants (#2203 )	2023-07-14 19:44:08 +02:00
James Reynolds	229aab351c	make : fix combination of LLAMA_METAL and LLAMA_MPI (#2208 ) Fixes https://github.com/ggerganov/llama.cpp/issues/2166 by moving commands after the CFLAGS are changed.	2023-07-14 20:34:40 +03:00
Georgi Gerganov	697966680b	ggml : sync (ggml_conv_2d, fix mul_mat bug, CUDA GLM rope)	2023-07-14 16:36:41 +03:00
Kawrakow	27ad57a69b	Metal: faster Q4_0 and Q4_1 matrix x vector kernels (#2212 ) * 3-5% faster Q4_0 on Metal * 7-25% faster Q4_1 on Metal * Oops, forgot to delete the original Q4_1 kernel --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2023-07-14 11:46:21 +02:00
Howard Su	32c5411631	Revert "Support using mmap when applying LoRA (#2095 )" (#2206 ) Has perf regression when mlock is used. This reverts commit `2347463201`.	2023-07-13 21:58:25 +08:00
Howard Su	ff5d58faec	Fix compile error on Windows CUDA (#2207 )	2023-07-13 21:58:09 +08:00
Bodo Graumann	b782422a3e	devops : add missing quotes to bash script (#2193 ) This prevents accidentally expanding arguments that contain spaces.	2023-07-13 16:49:14 +03:00
Shouzheng Liu	1cbf561466	metal : new q4_0 matrix-vector kernel (#2188 ) Prefetch data to improve GPU utilization. ~48% faster for 33B model.	2023-07-12 23:10:55 +03:00
Georgi Gerganov	975221e954	ggml : broadcast mul_mat + conv batch support (#2199 ) * ggml : broadcast mul_mat + conv batch support * ggml : apply mul_mat broadcast fix by @jploski	2023-07-12 20:51:29 +03:00
Georgi Gerganov	4523d10d0c	ggml : add ggml_pool_1d and ggml_pool_2d	2023-07-12 20:32:15 +03:00
Georgi Gerganov	680e6f9177	cuda : add gelu support	2023-07-12 20:32:15 +03:00
Howard Su	4e7464ef88	FP16 is supported in CM=6.0 (#2177 ) * FP16 is supported in CM=6.0 * Building PTX code for both of 60 and 61 Co-authored-by: Johannes Gäßler <johannesg@5d6.de>	2023-07-12 20:18:40 +08:00
Johannes Gäßler	2b5eb72e10	Fixed __dp4a compute capability: 6.0 -> 6.1 (#2189 )	2023-07-12 10:38:52 +02:00
Georgi Gerganov	f7d278faf3	ggml : revert CUDA broadcast changes from #2183 (#2191 )	2023-07-12 10:54:19 +03:00
Georgi Gerganov	20d7740a9b	ggml : sync (abort callback, mul / add broadcast, fix alibi) (#2183 )	2023-07-11 22:53:34 +03:00
Spencer Sutton	5bf2a27718	ggml : remove src0 and src1 from ggml_tensor and rename opt to src (#2178 ) * Add ggml changes * Update train-text-from-scratch for change * mpi : adapt to new ggml_tensor->src --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-07-11 19:31:10 +03:00
Bach Le	c9c74b4e3f	llama : add classifier-free guidance (#2135 ) * Initial implementation * Remove debug print * Restore signature of llama_init_from_gpt_params * Free guidance context * Make freeing of guidance_ctx conditional * Make Classifier-Free Guidance a sampling function * Correct typo. CFG already means context-free grammar. * Record sampling time in llama_sample_classifier_free_guidance * Shift all values by the max value before applying logsoftmax * Fix styling based on review	2023-07-11 19:18:43 +03:00
Jinwoo Jeong	3ec7e596b2	docker : add '--server' option (#2174 )	2023-07-11 19:12:35 +03:00
Chad Brewbaker	917831c63a	readme : fix zig build instructions (#2171 )	2023-07-11 19:03:06 +03:00
Howard Su	2347463201	Support using mmap when applying LoRA (#2095 ) * Support using mmap when applying LoRA * Fix Linux * Update comment to reflect the support lora with mmap	2023-07-11 22:37:01 +08:00
LostRuins	bbef28218f	Possible solution to allow K-quants on models with n_vocab!=32000 (#2148 ) * This allows LLAMA models that were previously incompatible with K quants to function mostly as normal. This happens when a model has a vocab != 32000, e.g 32001 which means it's not divisible by 256 or 64. Since the problematic dimensions only apply for `tok_embeddings.weight` and `output.weight` (dimentions 4096 x n_vocab), we can simply quantize these layers to Q8_0 whereas the majority of the hidden layers are still K-quanted since they have compatible dimensions. * Fix indentation Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * As an alternative, to avoid failing on Metal due to lack of Q8_0 support, instead quantize tok_embeddings.weight to Q4_0 and retain output.weight as F16. This results in a net gain of about 55mb for a 7B model compared to previous approach, but should minimize adverse impact to model quality. --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-07-11 22:01:08 +08:00
Evan Miller	5656d10599	mpi : add support for distributed inference via MPI (#2099 ) * MPI support, first cut * fix warnings, update README * fixes * wrap includes * PR comments * Update CMakeLists.txt * Add GH workflow, fix test * Add info to README * mpi : trying to move more MPI stuff into ggml-mpi (WIP) (#2099) * mpi : add names for layer inputs + prep ggml_mpi_graph_compute() * mpi : move all MPI logic into ggml-mpi Not tested yet * mpi : various fixes - communication now works but results are wrong * mpi : fix output tensor after MPI compute (still not working) * mpi : fix inference * mpi : minor * Add OpenMPI to GH action * [mpi] continue-on-error: true * mpi : fix after master merge * [mpi] Link MPI C++ libraries to fix OpenMPI * tests : fix new llama_backend API * [mpi] use MPI_INT32_T * mpi : factor out recv / send in functions and reuse * mpi : extend API to allow usage with outer backends (e.g. Metal) --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-07-10 18:49:56 +03:00
oobabooga	1d16309969	llama : remove "first token must be BOS" restriction (#2153 )	2023-07-09 11:59:53 +03:00
Nigel Bosch	db4047ad5c	main : escape prompt prefix/suffix (#2151 )	2023-07-09 11:56:18 +03:00
JackJollimore	18780e0a5e	readme : update Termux instructions (#2147 ) The file pathing is significant when running models inside of Termux on Android devices. llama.cpp performance is improved with loading a .bin from the $HOME directory.	2023-07-09 11:20:43 +03:00
clyang	3bbc1a11f0	ggml : fix buidling with Intel MKL but ask for "cblas.h" issue (#2104 ) (#2115 ) * Fix buidling with Intel MKL but ask for "cblas.h" issue * Use angle brackets to indicate the system library	2023-07-09 11:12:20 +03:00
rankaiyx	2492a53fd0	readme : add more docs indexes (#2127 ) * Update README.md to add more docs indexes * Update README.md to add more docs indexes	2023-07-09 10:38:42 +03:00
Johannes Gäßler	64639555ff	Fixed OpenLLaMA 3b CUDA mul_mat_vec_q (#2144 )	2023-07-08 20:01:44 +02:00
Johannes Gäßler	061f5f8d21	CUDA: add __restrict__ to mul mat vec kernels (#2140 )	2023-07-08 00:25:15 +02:00
dylan	84525e7962	docker : add support for CUDA in docker (#1461 ) Co-authored-by: canardleteer <eris.has.a.dad+github@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-07-07 21:25:25 +03:00
Georgi Gerganov	a7e20edf22	ci : switch threads to 1 (#2138 )	2023-07-07 21:23:57 +03:00
Qingyou Meng	1d656d6360	ggml : change ggml_graph_compute() API to not require context (#1999 ) * ggml_graph_compute: deprecate using ggml_context, try resolve issue #287 * rewrite: no longer consider backward compitability; plan and make_plan * minor: rename ctx as plan; const * remove ggml_graph_compute from tests/test-grad0.c, but current change breaks backward * add static ggml_graph_compute_sugar() * minor: update comments * reusable buffers * ggml : more consistent naming + metal fixes * ggml : fix docs * tests : disable grad / opt + minor naming changes * ggml : add ggml_graph_compute_with_ctx() - backwards compatible API - deduplicates a lot of copy-paste * ci : enable test-grad0 * examples : factor out plan allocation into a helper function * llama : factor out plan stuff into a helper function * ci : fix env * llama : fix duplicate symbols + refactor example benchmark * ggml : remove obsolete assert + refactor n_tasks section * ggml : fix indentation in switch * llama : avoid unnecessary bool * ggml : remove comments from source file and match order in header --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-07-07 19:24:01 +03:00
Georgi Gerganov	7242140283	ggml : remove sched_yield() call in ggml_graph_compute_thread() (#2134 )	2023-07-07 18:37:10 +03:00
Aarni Koskela	3e08ae99ce	convert.py: add mapping for safetensors bf16 (#1598 ) Fixes #1473	2023-07-07 09:12:49 -04:00
Howard Su	481f793acc	Fix opencl by wrap #if-else-endif with \n (#2086 )	2023-07-07 05:34:18 +02:00
Georgi Gerganov	dfd9fce6d6	ggml : fix restrict usage	2023-07-06 19:41:31 +03:00
Judd	36680f6e40	convert : update for baichuan (#2081 ) 1. guess n_layers; 2. relax warnings on context size; 3. add a note that its derivations are also supported. Co-authored-by: Judd <foldl@boxvest.com>	2023-07-06 19:23:49 +03:00
tslmy	a17a2683d8	alpaca.sh : update model file name (#2074 ) The original file name, `ggml-alpaca-7b-q4.bin`, implied the first-generation GGML. After the breaking changes (mentioned in https://github.com/ggerganov/llama.cpp/issues/382), `llama.cpp` requires GGML V3 now. Those model files are named `ggmlv3.bin`. We should change the example to an actually working model file, so that this thing is more likely to run out-of-the-box for more people, and less people would waste time downloading the old Alpaca model.	2023-07-06 19:17:50 +03:00
Tobias Lütke	31cfbb1013	Expose generation timings from server & update completions.js (#2116 ) * use javascript generators as much cleaner API Also add ways to access completion as promise and EventSource * export llama_timings as struct and expose them in server * update readme, update baked includes * llama : uniform variable names + struct init --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-07-05 16:51:13 -04:00
Jesse Jojo Johnson	983b555e9d	Update Server Instructions (#2113 ) * Update server instructions for web front end * Update server README * Remove duplicate OAI instructions * Fix duplicate text --------- Co-authored-by: Jesse Johnson <thatguy@jessejojojohnson.com>	2023-07-05 21:03:19 +03:00

1 2 3 4 5 ...

844 commits