llama.cpp

Author	SHA1	Message	Date
Yazan Agha-Schrader	39a163f76e	add missing char	2024-05-29 13:32:33 +02:00
Yazan Agha-Schrader	513406ab60	add more comon stop tokens	2024-05-29 13:29:00 +02:00
Yazan Agha-Schrader	80b6143f78	more prompt format fixes	2024-05-29 13:19:22 +02:00
Yazan Agha-Schrader	ca565f4ed6	fix llama3 prompt template	2024-05-29 12:08:39 +02:00
Yazan Agha-Schrader	9fa0aa53f5	fix chatml & add llama3 format	2024-05-29 11:26:34 +02:00
Yazan Agha-Schrader	5fa255edfb	add user message suffix	2024-05-29 10:28:07 +02:00
Yazan Agha-Schrader	eac8d739a5	update forgotten css theme	2024-05-29 08:54:04 +02:00
Yazan Agha-Schrader	aa493e022d	add css class	2024-05-29 08:45:20 +02:00
Yazan Agha-Schrader	9bb074e1f6	add phi3 to dropdown	2024-05-29 06:28:27 +02:00
Yazan Agha-Schrader	be675948d4	add phi-3 prompt template	2024-05-29 05:28:52 +02:00
Yazan Agha-Schrader	efbbc95321	remove ms per token, since not relevant for most webui users and use cases	2024-05-28 02:51:47 +02:00
Yazan Agha-Schrader	6bf7ae08dd	update const ModelGenerationInfo	2024-05-28 02:46:18 +02:00
Yazan Agha-Schrader	8768b4f5ea	use template literales for promptFormats.js	2024-05-28 02:25:49 +02:00
Yazan Agha-Schrader	e2d917c5f0	fix grammar field width	2024-05-28 00:29:08 +02:00
Yazan Agha-Schrader	7055d84020	fix FloatField and BoolField tooltips	2024-05-28 00:23:18 +02:00
Yazan Agha-Schrader	569cbd8bbe	Merge branch 'ggerganov:master' into server-ui-pr	2024-05-27 23:49:47 +02:00
Yazan Agha-Schrader	0f077968c0	add tooltips to the parameters with comprehensible explanations	2024-05-27 23:46:18 +02:00
Yazan Agha-Schrader	c7803876ce	move API to the top, rearrange param sliders. update css	2024-05-27 22:33:18 +02:00
Johannes Gäßler	10b1e45876	make: add --device-debug to NVCC debug flags (#7542 )	2024-05-27 19:34:40 +02:00
agray3	197c00681b	Allow multiple copy function pointers for CUDA graph kernel param updates (#7565 ) CUDA graphs require parameter updates to kernels associated with GGML_OP_CPY nodes. Previously the implementation only checked for a single CUDA kernel in such nodes, but this caused a bug in cases where 2 such kernels exist. This fixes the issue by using a vector to allow multiple function pointers to be stored and checked against. Fixes #7942	2024-05-27 19:33:42 +02:00
AidanBeltonS	95f84d5ce8	Fix q_xxs using mul_mat_q (#7459 )	2024-05-27 22:04:51 +05:30
Yazan Agha-Schrader	450471454c	clean the code	2024-05-27 17:00:03 +02:00
Yazan Agha-Schrader	a2edaf48c3	Add API key CSS classes and update styling in style.css	2024-05-27 16:33:45 +02:00
Yazan Agha-Schrader	b16e10bb69	some necessary fixes	2024-05-27 16:28:09 +02:00
Yazan Agha-Schrader	5a14ef1dca	add api-key css classes	2024-05-27 15:15:00 +02:00
AidanBeltonS	5487593bc7	Add freq factors (#7495 )	2024-05-27 18:04:09 +05:30
Georgi Gerganov	1d8fca72ae	metal : add GGML_OP_REPEAT kernels (#7557 ) ggml-ci	2024-05-27 12:10:19 +03:00
Georgi Gerganov	62bfef5194	metal : disable FA kernel for HS=256 (#7556 ) ggml-ci	2024-05-27 10:38:39 +03:00
Yazan Agha-Schrader	5d455f2789	chore: Update HTML meta tags in index.html file	2024-05-27 09:20:10 +02:00
Yazan Agha-Schrader	8d49b9906a	de prompts	2024-05-27 09:18:19 +02:00
Yazan Agha-Schrader	bd2c97c51a	add the belonging stuff: css,favicon etc	2024-05-27 09:14:54 +02:00
Yazan Agha-Schrader	902862a505	migrate my eary work	2024-05-27 08:33:36 +02:00
Georgi Gerganov	eaf6e03174	llama : add comments about experimental flags (#7544 )	2024-05-27 09:24:13 +03:00
Yazan Agha-Schrader	0a30b6e082	ic	2024-05-27 07:05:08 +02:00
Brian	d6ef0e77dd	github: add self sorted issue ticket forms (#7543 ) * github: add self sorted issue ticket forms [no ci] * github: consolidate BSD in bug issue ticket * github: remove contact from bug ticket template [no ci] * github: remove bios from os dropdown in bug report [no ci]	2024-05-27 10:54:30 +10:00
Georgi Gerganov	dff451cfa1	flake.lock: Update (#7540 ) Flake lock file updates: • Updated input 'nixpkgs': 'github:NixOS/nixpkgs/4a6b83b05df1a8bd7d99095ec4b4d271f2956b64?narHash=sha256-%2BNpbZRCRisUHKQJZF3CT%2Bxn14ZZQO%2BKjxIIanH3Pvn4%3D' (2024-05-17) → 'github:NixOS/nixpkgs/bfb7a882678e518398ce9a31a881538679f6f092?narHash=sha256-4zSIhSRRIoEBwjbPm3YiGtbd8HDWzFxJjw5DYSDy1n8%3D' (2024-05-24) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>	2024-05-26 08:54:56 -07:00
Brian	d298382ad9	main: replace --no-special with --special (#7534 ) This also flips the default behavior of the output to not include control token by default.	2024-05-27 00:10:17 +10:00
Galunid	32a28217f4	Fix aya-23 conversion scripts (#7539 )	2024-05-26 16:02:34 +02:00
Bartowski	c429b33beb	llama : add Smaug 70B support (#7402 )	2024-05-26 15:28:35 +03:00
Aarni Koskela	9146d36fe7	Readme: add akx/ggify to tools (#1484 )	2024-05-26 22:09:42 +10:00
HanishKVC	b9adcbbf92	SimpleChat Completion Mode flexibility and cleanup, Settings gMe, Optional sliding window (#7480 ) * SimpleChat: A placeholder system prompt, Use usage msg in code Just have a alert msg wrt needing javascript enabled in html. And have usage message from js file. Update the usage message a bit. So also enable switch session wrt setup_ui call. Add a possible system prompt as a placeholder for the system-input. * SimpleChat:CompletionMode: Allow control of Role: prefix * SimpleChat:Completion: Avoid Role: prefix; Newline only in between In completion mode * avoid inserting Role: prefix before each role's message * avoid inserting newline at the begin and end of the prompt message. However if there are multiple role messages, then insert newline when going from one role's message to the next role's message. * SimpleChat:CompletionMode: Update readme/usage, trim textarea newline Readme update wrt completion mode behavior. Usage help updated wrt completion mode behavior. When changing from input to textarea elment wrt user input, the last newline at the end of the user input wrt textarea, was forgotten to be filtered, this is fixed now. However if user wants to have a explicit newline they can using shift+enter to insert a newline, that wont be removed. The extra newline removal logic uses substring and keyup to keep things simple and avoid some previously noted bugs wrt other events in the key path as well as IME composition etal. * SimpleChat:SC: Ensure proper clearing/reseting previous logic would have cleared/reset the xchat, without doing the same wrt iLastSys, thus leading to it pointing to a now non existent role-content entry. So if a user set a system prompt and used completion mode, it would have done the half stupid clear, after the model response was got. Inturn when user tries to send a new completion query, it would inturn lead to handle_user_submit trying to add/update system prompt if any, which will fail, bcas iLastSys will be still pointing to a non existant entry. This is fixed now, by having a proper clear helper wrt SC class. * SimpleChat: Update usage note and readme a bit * SimpleChat:Completion: clear any prev chat history at begining Previously any chat history including model response to a completion query would have got cleared, after showing the same to the user, at the end of handle_user_submit, rather than at the begining. This gave the flexibility that user could switch from chat mode to completion mode and have the chat history till then sent to the ai model, as part of the completion query. However this flow also had the issue that, if user switches between different chat sessions, after getting a completion response, they can no longer see the completion query and its response that they had just got. The new flow changes the clearing of chat history wrt completion mode to the begining of handle_user_submit, so that user doesnt lose the last completion mode query and response, till a new completion mode query is sent to the model, even if they were to switch between the chat sessions. At the same time the loss of flexibility wrt converting previous chat history into being part of the completion query implicitly doesnt matter, because now the end user can enter multiline queries. * SimpleChat:Try read json early, if available For later the server flow doesnt seem to be sending back data early, atleast for the request (inc options) that is currently sent. if able to read json data early on in future, as and when ai model is generating data, then this helper needs to indirectly update the chat div with the recieved data, without waiting for the overall data to be available. * SimpleChat: Rename the half asleep mis-spelled global var * SimpleChat: Common chat request options from a global object * SimpleChat: Update title, usage and readme a bit Keep the title simple so that print file name doesnt have chars that need to be removed. Update readme wrt some of the new helpers and options. Change Usage list to a list of lists, add few items and style it to reduce the margin wrt lists. * SimpleChat:ChatRequestOptions: max_tokens As some times based on the query from the user, the ai model may get into a run away kind of generation with repeatations etal, so adding max_tokens to try and limit this run away behaviour, if possible. * SimpleChat: Reduce max_tokens to be small but still sufficient * SimpleChat: Consolidate global vars into gMe, Display to user This allows the end user to see the settings used by the logic, as well as allows users to change/update the settings if they want to by using devel-tools/console * SimpleChat:SlidingWindow: iRecentUserMsgCnt to limit context load This is disabled by default. However if enabled, then in addition to latest system message, only the last N user messages, after the latest system message and its reponses from the ai model will be sent to the ai-model, when querying for a new response. This specified N also includes the latest user query. * SimpleChat: placeholder based usage hint for user-in textarea * SimpleChat: Try make user experience better, if possible Reduce chat history context sent to the server/ai-model to be just the system-prompt, prev-user-request-and-ai-response and cur-user-request, instead of the previous full chat history. This way if there is any response with garbage/repeatation, it doesnt mess with things beyond the next question, in some ways. Increase max_tokens to 1024, so that a relatively large previous reponse doesnt eat up the space available wrt next query-response. However dont forget that the server when started should also be started with a model context size of 1k or more, to be on safe side. Add frequency and presence penalty fields set to 1.2 to the set of fields sent to server along with the user query. So that the model is partly set to try avoid repeating text in its response. * SimpleChat:Add n_predict (equiv max_tokens) for llamacpp server The /completions endpoint of examples/server doesnt take max_tokens, instead it takes the internal n_predict, for now add the same on the client side, maybe later add max_tokens to /completions endpoint handling. * SimpleChat: Note about trying to keep things simple yet flexible	2024-05-26 10:56:34 +10:00
Georgi Gerganov	9588f196b1	train : change default FA argument (#7528 )	2024-05-25 15:22:35 +03:00
Brian	3cbd23ed88	labeler: added Apple Metal detector (+Kompute) (#7529 ) * labeler: added Apple Metal detector [no ci] * labeler: add Kompute to detector [no ci]	2024-05-25 19:30:42 +10:00
Justine Tunney	00c6390793	main : don't print special tokens with --grammar (#6923 ) * main : don't print special tokens with --grammar The CLI interface was recently changed to print special control tokens like the </s> stop message one. This token shouldn't be printed if the grammar flag was passed, unless the grammar specifies it, because that breaks shell-scriptability. * main: use seperate stream for control characters * main: use dprintf and add --ctrl-token-no-out and --ctrl-token-fd-out * main: dprintf isn't part of the IEEE POSIX standard. Just use write(). * main: remove --ctrl-token-fd-out in favor for fcntl() based detection * common.cpp: accidentally removed --interactive-first * main: only merge stdout and control token if not in conversation or grammar mode * main: rejig control token descriptor handling * main: must check pipe status on very top of program * main: renamed --no-special from --ctrl-token-no-out and other refactoring * main: refactor ctrl_token_no_out --> no_special * llama: rename llama_token_is_control_token() to llama_token_is_control() * main: remove special token file descriptor feature (#5) --------- Co-authored-by: Brian <mofosyne@gmail.com>	2024-05-25 19:04:03 +10:00
Masaya, Kato	faa0e6979a	ggml: aarch64: SVE kernels for q8_0_q8_0, q4_0_q8_0 vector dot (#7433 ) * Add SVE support for q4_0_q8_0 q8_0_q8_0 * remove ifdef	2024-05-25 11:42:31 +03:00
Elton Kola	9791f40258	android : module (#7502 ) * move ndk code to a new library * add gradle file	2024-05-25 11:11:33 +03:00
Xuan Son Nguyen	902184dd3a	fix missing slash in `fs_get_cache_directory()` (#7503 ) * fix missing slash in fs_get_cache_directory() * use LOCALAPPDATA for fs_get_cache_directory() * better code style	2024-05-25 13:30:59 +10:00
Mikko Juola	57684331fc	Make tokenize CLI tool have nicer command line arguments. (#6188 ) * Make tokenizer.cpp CLI tool nicer. Before this commit, tokenize was a simple CLI tool like this: tokenize MODEL_FILENAME PROMPT [--ids] This simple tool loads the model, takes the prompt, and shows the tokens llama.cpp is interpreting. This changeset makes the tokenize more sophisticated, and more useful for debugging and troubleshooting: tokenize [-m, --model MODEL_FILENAME] [--ids] [--stdin] [--prompt] [-f, --file] [--no-bos] [--log-disable] It also behaves nicer on Windows now, interpreting and rendering Unicode from command line arguments and pipes no matter what code page the user has set on their terminal. * style fix: strlen(str) == 0 --> str == 0 Simplify tokenize.cpp; by getting rid of handling positional style arguments. It must now be invoked with long --model, --prompt etc. arguments only. Shortens the code. * tokenize.cpp: iostream header no longer required --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> Co-authored-by: brian khuu <mofosyne@gmail.com>	2024-05-25 11:14:42 +10:00
compilade	b83bab15a5	gguf-py : fix and simplify quantized shape round-trip (#7483 ) * gguf-py : fix and simplify quantized shape round-trip * gguf-py : remove unused import	2024-05-25 11:11:48 +10:00
Georgi Gerganov	d041d2ceaa	flake.lock: Update (#7232 ) Flake lock file updates: • Updated input 'flake-parts': 'github:hercules-ci/flake-parts/e5d10a24b66c3ea8f150e47dfdb0416ab7c3390e?narHash=sha256-yzcRNDoyVP7%2BSCNX0wmuDju1NUCt8Dz9%2BlyUXEI0dbI%3D' (2024-05-02) → 'github:hercules-ci/flake-parts/8dc45382d5206bd292f9c2768b8058a8fd8311d9?narHash=sha256-/GJvTdTpuDjNn84j82cU6bXztE0MSkdnTWClUCRub78%3D' (2024-05-16) • Updated input 'nixpkgs': 'github:NixOS/nixpkgs/63c3a29ca82437c87573e4c6919b09a24ea61b0f?narHash=sha256-4cPymbty65RvF1DWQfc%2BBc8B233A1BWxJnNULJKQ1EY%3D' (2024-05-02) → 'github:NixOS/nixpkgs/4a6b83b05df1a8bd7d99095ec4b4d271f2956b64?narHash=sha256-%2BNpbZRCRisUHKQJZF3CT%2Bxn14ZZQO%2BKjxIIanH3Pvn4%3D' (2024-05-17) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>	2024-05-24 08:59:06 -07:00

1 2 3 4 5 ...

3039 commits