cosmopolitan

mirror of https://github.com/jart/cosmopolitan.git synced 2025-02-14 10:18:02 +00:00

Author	SHA1	Message	Date
Justine Tunney	daf4454a06	Validate privileged code relationships - Work towards improving non-optimized build support - Introduce MODE=zero which is -O0 without ASAN/UBSAN - Use system GCC when ~/.cosmo.mk has USE_SYSTEM_TOOLCHAIN=1 - Have package.com check .privileged code doesn't call non-privileged	2023-06-08 04:38:06 -07:00
Justine Tunney	8fdb31681a	Introduce support for GGJT v3 file format llama.com can now load weights that use the new file format which was introduced a few weeks ago. Note that, unlike llama.cpp, we will keep support for old file formats in our tool so you don't need to convert your weights when the upstream project makes breaking changes. Please note that using ggjt v3 does make avx2 inference go 5% faster for me.	2023-06-03 15:46:21 -07:00
Justine Tunney	e7eb0b3070	Make more ML improvements - Fix UX issues with llama.com - Do housekeeping on libm code - Add more vectorization to GGML - Get GGJT quantizer programs working well - Have the quantizer keep the output layer as f16c - Prefetching improves performance 15% if you use fewer threads	2023-05-16 08:07:23 -07:00
Justine Tunney	210187cf77	Perform some code cleanup	2023-05-15 16:32:10 -07:00
Justine Tunney	282dd8e7b7	Get radpajama to build make -j8 o//third_party/radpajama/radpajama.com make -j8 o//third_party/radpajama/radpajama-chat.com This change gets the radpajama.mk config working. This package depends on THIRD_PARTY_GGML but it's configured to call ggjt_v1(), so that the library will provide the old quantizers. The ggml_quantize_chunk() API will now dispatch to older quantizers based on the configured version.	2023-05-13 20:44:36 -07:00
Justine Tunney	4a8a81eb9f	Fix llama.com interactive mode regressions	2023-05-13 00:09:38 -07:00
Justine Tunney	45186c74ac	Introduce -q (quiet flag) and improve ctrl-c ux	2023-05-12 09:46:07 -07:00
Justine Tunney	80c174d494	Clean up llama.com anti/stop/reverse-prompt code Example use case for JSON completion: $ m=opt $ make -j16 m=$m o/$m/third_party/ggml/llama.com $ o/$m/third_party/ggml/llama.com -m llama.bin -p '{"key": "life", "val": ' -r '}' 42} This provides better control. More sophisticated facilities for controlling text generation will be provided soon enough.	2023-05-12 08:20:58 -07:00
Justine Tunney	ca990ef091	Make `llama.com -h` print to stdout	2023-05-10 04:55:59 -07:00
Justine Tunney	5f57fc1f59	Upgrade llama.cpp to e6a46b0ed1884c77267dc70693183e3b7164e0e0	2023-05-10 04:20:48 -07:00
Justine Tunney	3dac9f8999	Use Companion AI in llama.com by default	2023-04-30 23:08:15 -07:00
Justine Tunney	b31ba86ace	Introduce prompt caching so prompts load instantly This change also introduces an ephemeral status line in non-verbose mode to display a load percentage status when slow operations are happening.	2023-04-28 16:15:26 -07:00
Justine Tunney	1c2da3a55a	Make shell usability improvements to llama.cpp - Introduce -v and --verbose flags - Don't print stats / diagnostics unless -v is passed - Reduce --top_p default from 0.95 to 0.70 - Change --reverse-prompt to no longer imply --interactive - Permit --reverse-prompt specifying custom EOS if non-interactive	2023-04-28 02:54:11 -07:00
Justine Tunney	e8b43903b2	Import llama.cpp https://github.com/ggerganov/llama.cpp 0b2da20538d01926b77ea237dd1c930c4d20b686 See third_party/ggml/README.cosmo for changes	2023-04-27 14:37:14 -07:00

14 commits