Merge upstream #1

pi6am · 2024-07-09T07:10:06Z

I have read the contributing guidelines

* whisper : use ggml_backend_sched (wip) * use sched in whisper_allocr * whisper : single backend in whisper_context * whisper : remove whisper_state->backends_used * whisper : remove whisper_context->backend * whisper : reset scheduler after init * whisper : fix external encoder (e.g. CoreML) * whisper : cleanup * whisper : handle null GPU buffer types + fix sycl --------- Co-authored-by: slaren <slarengh@gmail.com>

Signed-off-by: thxCode <thxcode0824@gmail.com>

On hosts which are not prepared/dedicated to execute code using CUDA it is still possible to compile llama.cpp with CUDA support by just installing the development packages. Missing are the runtime libraries like /usr/lib64/libcuda.so* and currently the link step will fail. The development environment is prepared for such situations. There are stub libraries for all the CUDA libraries available in the $(CUDA_PATH)/lib64/stubs directory. Adding this directory to the end of the search path will not change anything for environments which currently work fine but will enable compiling llama.cpp also in case the runtime code is not available.

* Only use FIM middle if it exists * Only use FIM middle if it exists

* Random test: add_bos_token, add_eos_token * Random test: add BPE models for testing * Custom regex split fails with codepoint 0 * Fix falcon punctuation regex * Refactor llm_tokenizer_bpe: move code to constructor * Move 'add_special_bos/eos' logic to llm_tokenizer_bpe * Move tokenizer flags to vocab structure. * Default values for special_add_bos/eos * Build vocab.special_tokens_cache using vocab token types * Generalize 'jina-v2' per token attributes * Fix unicode whitespaces (deepseek-coder, deepseek-llm) * Skip missing byte tokens (falcon) * Better unicode data generation * Replace char32_t with uint32_t

* seperate lower precision GEMM from the main files * fix workgroup size hardcode

@slaren

* un-ignore `build-info.cmake` and `build-info.sh` I am assuming that ignoring them was unintentional. If they are ignored, some tools, like cargo, will consider the files inexistent, even if they're comitted, for the purpose of publishing. This leads to the build failing in such cases. * un-ignore `build-info.cpp.in` For the same reason as the previous two files. * Reorganize `.gitignore` * Add exceptions for files mentioned by @slaren I did leave .clang-tidy since it was explicitly ignored before. * Add comments for organization * Sort some lines for pretty * Test with `make` and `cmake` builds to ensure no build artifacts might be comitted * Remove `.clang-tidy` from `.gitignore` Per comment by @ggerganov * Remove `IDEWorkspaceChecks.plist` from root-level `.gitignore`

Currently the Metal backend does not support BF16. `ggml_metal_supports_op` was returning true in these cases, leading to a crash with models converted with `--leave-output-tensor`. This commit checks if the first few sources types are BF16 and returns false if that's the case.

* CUDA: stream-k decomposition for MMQ * fix undefined memory reads for small matrices

* add sycl preset * fix debug link error. fix windows crash * update README

* common: fix warning * Update common/common.cpp Co-authored-by: slaren <slarengh@gmail.com> --------- Co-authored-by: slaren <slarengh@gmail.com>

…ml-org#8040)

* create append_pooling operation; allow to specify attention_type; add last token pooling; update examples * find result_norm/result_embd tensors properly; update output allocation logic * only use embd output for pooling_type NONE * get rid of old causal_attn accessor * take out attention_type; add in llama_set_embeddings * bypass logits when doing non-NONE pooling

ggml-ci

* initial iq4_xs * fix ci * iq4_nl * iq1_m * iq1_s * iq2_xxs * iq3_xxs * iq2_s * iq2_xs * iq3_s before sllv * iq3_s * iq3_s small fix * iq3_s sllv can be safely replaced with sse multiply

…ml-org#8022) * vulkan: detect multiple devices by deviceUUID instead of deviceID * vulkan: remove unneeded variables * vulkan: fix id query

# Conflicts: # .github/labeler.yml # .github/workflows/server.yml # .gitignore # CMakeLists.txt # Makefile # README-sycl.md # README.md # llama.cpp # requirements/requirements-convert-hf-to-gguf-update.txt # requirements/requirements-convert-hf-to-gguf.txt # requirements/requirements-convert-legacy-llama.txt # scripts/sync-ggml.last # tests/test-tokenizer-random.py

@ochafik

* Adding simple bare-bones test for end-to-end integration test for json validation against auto-generated JSON-schema grammars. * Adding additional examples as documented in ggml-org#7789 . Also adding the ability to automatically output improperly failing grammars to debug output files so they can more easily be examined in the gbnf-validator program. * Uncommenting formerly commented tests so that they fail for others who are attempting to reproduce the bugs. * Merging improved schema test methods added by @ochafik in ggml-org#7797 * Adding #define to temporarily remove failing tests so that this PR can pass CI, but still be useful for other PRs that want to leverage the framework. * Fixing nits from ochafik. Removing escape slashes, adding additional failing cases, fixing some other strings. * Fixing grammar indentation to be consistent throughout file.

@JohannesGaessler

…alues (ggml-org#8058) Uses the values computed by @JohannesGaessler in PR ggml-org#7413

* Give the CI builds a recognizable AVX1 name * Chat Adapters

# Conflicts: # .devops/full-cuda.Dockerfile # .devops/full-rocm.Dockerfile # .devops/llama-cli-cuda.Dockerfile # .devops/llama-cli-rocm.Dockerfile # .devops/llama-cli-vulkan.Dockerfile # .devops/llama-cpp-cuda.srpm.spec # .devops/llama-server-cuda.Dockerfile # .devops/llama-server-rocm.Dockerfile # .devops/llama-server-vulkan.Dockerfile # .github/workflows/build.yml # .github/workflows/docker.yml # CMakeLists.txt # Makefile # README.md # examples/llama.android/llama/src/main/cpp/CMakeLists.txt # flake.lock # ggml/CMakeLists.txt # ggml/src/CMakeLists.txt # grammars/README.md # scripts/sync-ggml-am.sh # scripts/sync-ggml.last # tests/test-chat-template.cpp # tests/test-grammar-integration.cpp # tests/test-json-schema-to-grammar.cpp

…tor to Gemma2 (ggml-org#8197) * Add attention and final logit softcapping. * fix * Add custom add_ functions * Disable flash attention for Gemma2 * Update src/llama.cpp Co-authored-by: slaren <slarengh@gmail.com> * Add default value for attention and final logit softcap value * Add custom kq scaling from Gemma2Attention * Remove custom pre attention scaling and use computed value instead. --------- Co-authored-by: slaren <slarengh@gmail.com>

…ggml-metal.o'. Stop` error on macOS Metal (#957)

…x/suffix is set (ggml-org#8203) * preserve new line llama_chat_format_single * disable chat template if in-prefix/suffix is set * remove redundant change

* align with rope.cu and move sycl-op to a single file

* Update README.md document BERT support * Update README.md

* nix : remove OpenCL remnants * minor : remove parentheses

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* Added gppm to Tool list in README * Update README.md --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* gemma2: add sliding window mask * fix data_swa uninitialized * better naming * add co-author Co-authored-by: Arlo Phoenix <arlo-phoenix@users.noreply.github.com> * replace list with single tensor * update * llama : minor styling * convert : add sanity check for query_pre_attn_scalar * fix small typo in README --------- Co-authored-by: Arlo Phoenix <arlo-phoenix@users.noreply.github.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* CUDA: refactor and optimize IQ MMVQ * uint -> uint32_t * __dp4a -> ggml_cuda_dp4a * remove MIN_CC_DP4A checks * change default * try CI fix

* fix gemma2 tokenizer convert * remove scores * improve code, fix new line issue

* use warp_size macro for all sycl kernels * fix mask of permute_sub_group_by_xor * fix rms_norm with correct warp number * fix rms_norm_f32/group_norm_f32 * move norm to norm.cpp file * fix quantize bug * fix mmvq's batch size

* fix win build conflict of math library * fix the condition: !(win32 & SYCL) * revert warp_size=16

* convert-hf : print output file name when completed This commit adds the output file name to the log message when the conversion is completed. The motivation for this change is that when `--outfile` option is not specified it migth not be obvious where the output file is written. With this change the output of running the script will be something like the following: ```console INFO:hf-to-gguf:Model successfully exported to models/gemma-2-9b-it.gguf. ``` Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com> * squash! convert-hf : print output file name when completed Updates the output of to support printing the directory if the output is split into multiple files. Also the output file name is now retrieved from the model_instance object. Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com> * squash! convert-hf : print output file name when completed Use parent attribute of Path object and string interpolation. Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com> * squash! convert-hf : print output file name when completed Use os.sep instead of hardcoding the path separator. Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com> --------- Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com>

# Conflicts: # .devops/nix/package.nix # CMakePresets.json # README.md # flake.lock # ggml/src/CMakeLists.txt # tests/test-backend-ops.cpp # tests/test-chat-template.cpp

(cherry picked from commit 572aba8)

* fstring #1 * fstring #2

* dictionary #1 * dictionary #2

* [example] batched-bench "segmentation fault" When `llama-batched-bench` is invoked _without_ setting `-npl`, "number of parallel prompts", it segfaults. The segfault is caused by invoking `max_element()` on a zero-length vector, `n_pl` This commit addresses that by first checking to see if the number of parallel prompts is zero, and if so sets the maximum sequence size to 1; otherwise, sets it to the original, the result of `max_element()`. Fixes, when running `lldb build/bin/llama-batched-bench -- -m models/Meta-Llama-3-8B.gguf` ``` * thread #1, queue = 'com.apple.main-thread', stop reason = EXC_BAD_ACCESS (code=1, address=0x0) frame #0: 0x000000010000366c llama-batched-bench`main(argc=3, argv=0x000000016fdff268) at batched-bench.cpp:72:28 69 llama_context_params ctx_params = llama_context_params_from_gpt_params(params); 70 71 // ensure enough sequences are available -> 72 ctx_params.n_seq_max = *std::max_element(n_pl.begin(), n_pl.end()); ``` * Update examples/batched-bench/batched-bench.cpp Co-authored-by: compilade <git@compilade.net> --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> Co-authored-by: compilade <git@compilade.net>

ggerganov and others added 30 commits June 18, 2024 09:50

ggml : sync

5326bcc

readme : update UI list (ggml-org#7943)

1193778

chore: clean useless beam search param (ggml-org#7985)

b96f9af

Signed-off-by: thxCode <thxcode0824@gmail.com>

Fix no gcc pragma on Windows (ggml-org#7751)

84f6de1

Only use FIM middle token if it exists (ggml-org#7648)

91c188d

* Only use FIM middle if it exists * Only use FIM middle if it exists

[SYCL] refactor (ggml-org#6408)

623494a

* seperate lower precision GEMM from the main files * fix workgroup size hardcode

codecov : remove (ggml-org#8004)

a04a953

ggml : synchronize threads using barriers (ggml-org#7993)

9c77ec1

server : fix smart slot selection (ggml-org#8020)

ba58993

CUDA: stream-k decomposition for MMQ (ggml-org#8018)

d50f889

* CUDA: stream-k decomposition for MMQ * fix undefined memory reads for small matrices

[SYCL] Fix windows build and inference (ggml-org#8003)

de391e4

* add sycl preset * fix debug link error. fix windows crash * update README

common: fix warning (ggml-org#8036)

abd894a

* common: fix warning * Update common/common.cpp Co-authored-by: slaren <slarengh@gmail.com> --------- Co-authored-by: slaren <slarengh@gmail.com>

convert-hf : Fix the encoding in the convert-hf-to-gguf-update.py (gg…

17b291a

…ml-org#8040)

requirements : Bump torch and numpy for python3.12 (ggml-org#8041)

b1ef562

swiftui : enable stream updating (ggml-org#7754)

0e64591

llama : optimize long word tokenization with WPM (ggml-org#8034)

a927b0f

ggml-ci

ggml : AVX IQ quants (ggml-org#7845)

7d5e877

* initial iq4_xs * fix ci * iq4_nl * iq1_m * iq1_s * iq2_xxs * iq3_xxs * iq2_s * iq2_xs * iq3_s before sllv * iq3_s * iq3_s small fix * iq3_s sllv can be safely replaced with sse multiply

vulkan: detect multiple devices by deviceUUID instead of deviceID (gg…

557b653

…ml-org#8022) * vulkan: detect multiple devices by deviceUUID instead of deviceID * vulkan: remove unneeded variables * vulkan: fix id query

fix ubatch, autoselect vulkan dgpu if possible

1339847

Update llama-quantize ppl/file size output from LLaMA-v1 to Llama-3 v…

5b48cd5

…alues (ggml-org#8058) Uses the values computed by @JohannesGaessler in PR ggml-org#7413

convert-hf : change assert to exception (ggml-org#8015)

3aa184a

add llava separator

12abc41

henk717 and others added 26 commits June 30, 2024 10:28

Chat Adapters (#956)

8421243

* Give the CI builds a recognizable AVX1 name * Chat Adapters

add tensor split gui input for vulkan

18df56b

edit sampler order warning

8a07ce3

Merge branch 'upstream' into concedo_experimental

b1f9c97

Resolve make: *** No rule to make target ggml-metal.m', needed by `…

7499a6b

…ggml-metal.o'. Stop` error on macOS Metal (#957)

Fix new line issue with chat template, disable template when in-prefi…

9ef0780

…x/suffix is set (ggml-org#8203) * preserve new line llama_chat_format_single * disable chat template if in-prefix/suffix is set * remove redundant change

flake.lock: Update (ggml-org#8218)

d0a7145

[SYCL] Update SYCL-Rope op and Refactor (ggml-org#8157)

197fe6c

* align with rope.cu and move sycl-op to a single file

Document BERT support. (ggml-org#8205)

694c59c

* Update README.md document BERT support * Update README.md

nix : remove OpenCL remnants (ggml-org#8235)

257f8e4

* nix : remove OpenCL remnants * minor : remove parentheses

nix : enable curl (ggml-org#8043)

3840b6f

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

readme : update tool list (ggml-org#8209)

0ddeff1

* Added gppm to Tool list in README * Update README.md --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

readme: add Paddler to the list of projects (ggml-org#8239)

dae57a1

CUDA: refactor and optimize IQ MMVQ (ggml-org#8215)

cb5fad4

* CUDA: refactor and optimize IQ MMVQ * uint -> uint32_t * __dp4a -> ggml_cuda_dp4a * remove MIN_CC_DP4A checks * change default * try CI fix

Fix gemma2 tokenizer convert (ggml-org#8244)

5fac350

* fix gemma2 tokenizer convert * remove scores * improve code, fix new line issue

[SYCL] Fix the sub group size of Intel (ggml-org#8106)

d08c20e

* use warp_size macro for all sycl kernels * fix mask of permute_sub_group_by_xor * fix rms_norm with correct warp number * fix rms_norm_f32/group_norm_f32 * move norm to norm.cpp file * fix quantize bug * fix mmvq's batch size

[SYCL] Fix win build conflict of math library (ggml-org#8230)

a9f3b10

* fix win build conflict of math library * fix the condition: !(win32 & SYCL) * revert warp_size=16

cuda : update supports_op for matrix multiplication (ggml-org#8245)

0e0590a

updated lite, add gemma 2 template

82202ae

Merge branch 'upstream' into concedo_experimental

0fc18d2

# Conflicts: # .devops/nix/package.nix # CMakePresets.json # README.md # flake.lock # ggml/src/CMakeLists.txt # tests/test-backend-ops.cpp # tests/test-chat-template.cpp

copy the metal file to root dir as well

06d9068

add target for oldcpu cuda

ecec9fb

(cherry picked from commit 572aba8)

pi6am merged commit 63c4437 into pi6am:concedo Jul 9, 2024

pi6am pushed a commit that referenced this pull request Jul 27, 2024

Streamline with fstrings (LostRuins#1006)

ce971a0

* fstring #1 * fstring #2

pi6am pushed a commit that referenced this pull request Jul 27, 2024

Streamline with dictionaries (LostRuins#1005)

7de1ebf

* dictionary #1 * dictionary #2

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Merge upstream #1

Merge upstream #1

pi6am commented Jul 9, 2024

Merge upstream #1

Merge upstream #1

Conversation

pi6am commented Jul 9, 2024