Skip to content

Pull requests: ggml-org/llama.cpp

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Add sm70 FlashAttention config case for DKQ=256, DV=256, ncols=64 CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#27997 opened Aug 30, 2026 by mistrjirka Draft
kv cache : optimize restoring non-contiguous cells testing Everything test related
#27991 opened Aug 29, 2026 by itsnotoger Loading…
ggml-cpu : add mirror NUMA strategy (replicate weights on each node) CUDA Related to the CUDA backend documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning server
#27986 opened Aug 29, 2026 by matteoscalabrini Loading…
On ggml-cpu: tweak ARM_NATIVE_FLAG native baseline arch when extensions need it ggml changes relating to the ggml tensor library for machine learning
#27984 opened Aug 29, 2026 by cameronelliott Loading…
quantize: add IQ2_NL and IQ3_NL types (CPU + Metal + CUDA + Vulkan) Apple Metal https://en.wikipedia.org/wiki/Metal_(API) conversion CUDA Related to the CUDA backend examples ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language testing Everything test related Vulkan Issues specific to the Vulkan backend
#27983 opened Aug 29, 2026 by EAddario Contributor Draft
CUDA: let any expert count use the fast mm_ids_helper path CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#27978 opened Aug 29, 2026 by ServeurpersoCom Contributor Loading…
qwen4exp: reduce the generation slowdown as context grows model Model specific
#27977 opened Aug 29, 2026 by ServeurpersoCom Contributor Loading…
vulkan: fuse GATED_DELTA_NET state write into recurrent cache ggml changes relating to the ggml tensor library for machine learning Vulkan Issues specific to the Vulkan backend
#27973 opened Aug 29, 2026 by PrajwalMukatti Draft
CUDA + ggml: add sparse-fa for DSV4/GLM CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning model Model specific testing Everything test related
#27970 opened Aug 29, 2026 by am17an Contributor Loading…
common: rename --tensor-read-lazy to --lazy-mode, add -lzm shorthand documentation Improvements or additions to documentation server
#27969 opened Aug 29, 2026 by ggerganov Member Loading…
[SYCL] Enhance to get the free memory of Intel GPU documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language
#27968 opened Aug 29, 2026 by arthw Contributor Loading…
memory : copy Hadamard matrix to k_rot tensor only if it has buffer assigned merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge.
#27967 opened Aug 29, 2026 by fairydreaming Contributor Loading…
metal : Add fa-vec tunings for M3 Pro Apple Metal https://en.wikipedia.org/wiki/Metal_(API) ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge.
#27963 opened Aug 29, 2026 by addianto Loading…
HIP : optimize IQ2/IQ3 (__vsub4 __vcmpne4) using SWAR CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#27962 opened Aug 29, 2026 by yanjs Loading…
ggml-cpu : conditionally add SpacemiT IME kernel sources ggml changes relating to the ggml tensor library for machine learning
#27961 opened Aug 29, 2026 by alanhc Draft
1 task done
ggml : add ggml_backend_op_alloc_size_may_expand, use it in RPC ggml changes relating to the ggml tensor library for machine learning
#27960 opened Aug 29, 2026 by ggerganov Member Loading…
ui : add model download pipeline server/ui
#27959 opened Aug 29, 2026 by allozaur Contributor 5/5 Draft
ui : add model compatibility estimation server/ui
#27957 opened Aug 29, 2026 by allozaur Contributor 4/5 Draft
CUDA : fix divergent FlashAttention barrier CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#27955 opened Aug 29, 2026 by mistrjirka Loading…
vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 ggml changes relating to the ggml tensor library for machine learning Vulkan Issues specific to the Vulkan backend
#27952 opened Aug 29, 2026 by 0cc4m Contributor Loading…
ProTip! Add no:assignee to see everything that’s not assigned.