Add quantize-stats command for testing quantization (#728)

Command that calculates some statistics over the errors introduced by quantization, like mean square error, max error and some percentile errors for layer weights. Should be useful for testing quantization improvements. Exposes some internal state from ggml and llama for testing
2023-04-08 00:09:18 +02:00 · 2023-04-08 00:09:18 +02:00 · 62cfc54f77
commit 62cfc54f77
parent 698f7b5d63
9 changed files with 415 additions and 17 deletions
--- a/llama.cpp
+++ b/llama.cpp
@ -1852,3 +1852,8 @@ const char * llama_print_system_info(void) {

    return s.c_str();
 }
+
+// For internal test use
+std::unordered_map<std::string, struct ggml_tensor *>& llama_internal_get_tensor_map(struct llama_context * ctx) {
+    return ctx->model.tensors;
+}