examples: support LLaVA v1.5 (multimodal model) (#3436)

* WIP: start implementing LLaVA

* rm scratch buf for now, will revert after cleanup

* LLaVA image encoder is working. will combine with llama

* Add llava inference code, but it's buggy. debugging

* LLaVA is working e2e, needs to optimize memory allocation + cleanup

* Use ggml_allocr + rm unnecessary code

* fix: crlf -> lf

* fix: new line at EoF

* fix: trailing whitespace

* Add readme

* Update readme

* Some cleanup

* Are you happy editorconfig?

* rm unused batch image preprocessing

* rm unused import

* fix: rm designated initializers

* introduce pad-to-square mode for non-square images

* are you happy editorconfig?

* gitignore /llava

* Handle cases where image file does not exist

* add llava target to Makefile

* add support for 13b model variant

* Maybe seed is unlucky?

* Check if apples are compared to apples

* are you happy editorconfig?

* Use temperature = 0.1 by default

* command line: use gpt_params_parse()

* minor

* handle default n_predict

* fix typo

* llava : code formatting, rename files, fix compile warnings

* do not use Wno-cast-qual for MSVC

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

This commit is contained in:

M. Yusuf Sarıgöz

2023-10-12 18:23:18 +03:00

• committed by

GitHub

parent 9e24cc6e2e

commit 370359e5ba

No known key found for this signature in database

GPG key ID: 4AEE18F83AFDEB23

15 changed files with 10214 additions and 2 deletions

									
										2

ggml.c
									
										View file
										
				@ -14428,7 +14428,7 @@ static void ggml_compute_forward_conv_2d_f16_f32(

				    int64_t t0 = ggml_perf_time_us();

				    UNUSED(t0);

				    GGML_TENSOR_BINARY_OP_LOCALS

				    GGML_TENSOR_BINARY_OP_LOCALS;

				    const int ith = params->ith;

				    const int nth = params->nth;

Rows
Columns

examples: support LLaVA v1.5 (multimodal model) (#3436)

2 ggml.c Unescape Escape View file

2

ggml.c

View file