llama : offload to RPC in addition to other backends (#7640)

* llama : offload to RPC in addition to other backends * - fix copy_tensor being called on the src buffer instead of the dst buffer - always initialize views in the view_src buffer - add RPC backend to Makefile build - add endpoint to all RPC object names * add rpc-server to Makefile * Update llama.cpp Co-authored-by: slaren <slarengh@gmail.com> --------- Co-authored-by: slaren <slarengh@gmail.com>
2024-06-03 20:03:26 +03:00 · 2024-06-03 20:03:26 +03:00 · bde7cd3cd9
commit bde7cd3cd9
parent a5735e4426
6 changed files with 86 additions and 53 deletions
--- a/ggml-backend.h
+++ b/ggml-backend.h
@ -225,7 +225,7 @@ extern "C" {

    // Tensor initialization
    GGML_API void ggml_backend_tensor_alloc(ggml_backend_buffer_t buffer, struct ggml_tensor * tensor, void * addr);
-    GGML_API void ggml_backend_view_init(ggml_backend_buffer_t buffer, struct ggml_tensor * tensor);
+    GGML_API void ggml_backend_view_init(struct ggml_tensor * tensor);


 #ifdef  __cplusplus