llama.cpp

History

Carolinabanana 5dc9dd7152 llama : add Command R Plus support (#6491 ) * Add Command R Plus GGUF * Add Command R Plus GGUF * Loading works up to LayerNorm2D * Export new tensors in 1D so they are not quantized. * Fix embedding layer based on Noeda's example * Whitespace * Add line * Fix unexpected tokens on MPS. Re-add F16 fix. ((Noeda) * dranger003: Fix block index overflow in CUDA dequantizing. * Reverted blocked multiplication code as it still has issues and could affect other Llama arches * export norms as f32 * fix overflow issues during quant and other cleanup * Type convention Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * dranger003: Fix more int overflow during quant. --------- Co-authored-by: S <seast@Ss-Mac-Studio.local> Co-authored-by: S <s@example.com> Co-authored-by: slaren <slarengh@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>		2024-04-09 11:16:13 +03:00
..
__init__.py	gguf-py: Refactor and allow reading/modifying existing GGUF files (#3981 )	2023-11-11 08:04:50 +03:00
constants.py	llama : add Command R Plus support (#6491 )	2024-04-09 11:16:13 +03:00
gguf.py	gguf-py: Refactor and allow reading/modifying existing GGUF files (#3981 )	2023-11-11 08:04:50 +03:00
gguf_reader.py	gguf : add support for I64 and F64 arrays (#6062 )	2024-03-15 10:46:51 +02:00
gguf_writer.py	gguf.py : add licence and version to gguf writer (#6504 )	2024-04-05 21:41:38 +03:00
py.typed	convert : various script cleanups/fixes + merges and special token handling (#2842 )	2023-08-30 11:25:50 +03:00
tensor_mapping.py	llama : add Command R Plus support (#6491 )	2024-04-09 11:16:13 +03:00
vocab.py	fix(gguf-py): special tokens are no longer skipped when add_<token>_token is set to false (#5487 )	2024-02-15 14:14:37 +01:00