lakis_lazulli

1.7.0c Silverpine LLM Error.

May 16th, 2026
45
0
290 days
Not a member of Pastebin yet? Sign Up, it unlocks many cool features!
text 16.88 KB | None | 0 0
  1. ***
  2. Welcome to KoboldCpp - Version 1.111.2
  3. Loading Chat Completions Adapter: /tmp/_MEINgQyNs/kcpp_adapters/AutoGuess.json
  4. Chat Completions Adapter Loaded
  5. Auto Recommended GPU Layers: 15
  6. GPU layers is default: Will enable AutoFit for increased estimation accuracy.
  7. System: Linux #1 SMP PREEMPT_DYNAMIC Debian 6.12.88-1 (2026-05-15) x86_64
  8. Detected Available GPU Memory: 16368 MB
  9. Detected Available RAM: 55199 MB
  10. Initializing dynamic library: koboldcpp_vulkan.so
  11. ==========
  12. Namespace(admin=False, admindir='', adminpassword=None, adminunloadtimeout=0, analyze='', autofit=True, autofitpadding=1024, autoswapmode=False, batchsize=512, benchmark=None, blasthreads=0, chatcompletionsadapter='AutoGuess', cli=False, config=None, contextsize=4096, debugmode=0, defaultgenamt=1024, device='', downloaddir='', draftamount=8, draftgpulayers=999, draftgpusplit=None, draftmodel='', embeddingsgpu=False, embeddingsmaxctx=0, embeddingsmodel='', enableguidance=False, exportconfig='', exporttemplate='', failsafe=False, flashattention=False, forceversion=False, foreground=False, gendefaults='', gendefaultsoverwrite=False, genlimit=0, gpulayers=15, highpriority=False, hordeconfig=None, hordegenlen=0, hordekey='', hordemaxctx=0, hordemodelname='', hordeworkername='', host='', ignoremissing=False, jinja=False, jinja_kwargs='', jinja_tools=False, launch=False, lora=None, loramult=1.0, lowvram=False, maingpu=-1, maxrequestsize=32, mcpfile='', mmproj='', mmprojcpu=False, model=['/home/lakis/Documents/redacted/Silverpine_1.7.0c_Linux/Silverpine_Data/StreamingAssets/KoboldCPP/Gemma-4-Sparse.gguf'], model_param='/home/lakis/Documents/redacted/Silverpine_1.7.0c_Linux/Silverpine_Data/StreamingAssets/KoboldCPP/Gemma-4-Sparse.gguf', moecpu=0, moeexperts=-1, multiplayer=False, multiuser=100, musicdiffusion='', musicembeddings='', musicllm='', musiclowvram=False, musicvae='', noavx2=False, noblas=False, nobostoken=False, nocertify=False, nofastforward=True, noflashattention=False, nommap=False, nomodel=False, nopipelineparallel=False, noshift=True, onready='', overridekv='', overridenativecontext=0, overridetensors='', password=None, pipelineparallel=False, port=5003, port_param=5001, preloadstory='', prompt='', proxy_port=None, quantkv=0, quiet=True, ratelimit=0, remotetunnel=False, ropeconfig=[0.0, 10000.0], routermode=False, savedatafile='', sdclamped=0, sdclampedsoft=0, sdclip1='', sdclip2='', sdclipgpu=False, sdconfig=None, sdconvdirect='off', sdflashattention=False, sdgendefaults=False, sdlora=[], sdloramult=[1.0], sdmaingpu=-1, sdmodel='', sdnotile=False, sdoffloadcpu=False, sdphotomaker='', sdquant=0, sdt5xxl='', sdthreads=0, sdtiledvae=768, sdupscaler='', sdvae='', sdvaeauto=False, sdvaecpu=False, showgui=False, singleinstance=False, skiplauncher=True, smartcache=0, smartcontext=False, ssl=None, tensor_split=None, testmemory=False, threads=15, ttsdir='', ttsgpu=False, ttsmaxlen=4096, ttsmodel='', ttsthreads=0, ttswavtokenizer='', unpack='', usecpu=False, usecuda=None, usemlock=False, usemmap=False, useswa=True, usevulkan=[], version=False, visionmaxres=1024, websearch=False, whispermodel='')
  13. ==========
  14. Loading Text Model: /home/lakis/Documents/redacted/Silverpine_1.7.0c_Linux/Silverpine_Data/StreamingAssets/KoboldCPP/Gemma-4-Sparse.gguf
  15. ggml_vulkan: Found 2 Vulkan devices:
  16. ggml_vulkan: 0 = AMD Radeon RX 6900 XT (RADV NAVI21) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 32 | shared memory: 65536 | int dot: 1 | matrix cores: none
  17. ggml_vulkan: 1 = AMD Radeon Graphics (RADV RAPHAEL_MENDOCINO) (radv) | uma: 1 | fp16: 1 | bf16: 0 | warp size: 32 | shared memory: 65536 | int dot: 1 | matrix cores: none
  18. llama_model_load_from_file_impl: using device Vulkan0 (AMD Radeon RX 6900 XT (RADV NAVI21)) (0000:03:00.0) - 14947 MiB free
  19. llama_model_loader: loaded meta data with 54 key-value pairs and 658 tensors from /home/lakis/Documents/redacted/Silverpine_1.7.0c_Linux/Silverpine_Data/StreamingAssets/KoboldCPP/Gemma-4-Sparse.gguf (version GGUF V3 (latest))
  20. print_info: file format = GGUF V3 (latest)
  21. print_info: file size = 15.85 GiB (5.40 BPW)
  22. init_tokenizer: initializing tokenizer for type 2
  23. load: 0 unused tokens
  24. load: control-looking token: 212 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden
  25. load: control-looking token: 50 '<|tool_response>' was not control-type; this is probably a bug in the model. its type will be overridden
  26. load: printing all EOG tokens:
  27. load: - 1 ('<eos>')
  28. load: - 50 ('<|tool_response>')
  29. load: - 106 ('<turn|>')
  30. load: - 212 ('</s>')
  31. load: special tokens cache size = 25
  32. load: token to piece cache size = 1.9445 MB
  33. print_info: arch = gemma4
  34. print_info: vocab_only = 0
  35. print_info: no_alloc = 0
  36. print_info: n_ctx_train = 262144
  37. print_info: n_embd = 2816
  38. print_info: n_embd_inp = 2816
  39. print_info: n_layer = 30
  40. print_info: n_head = 16
  41. print_info: n_head_kv = [8, 8, 8, 8, 8, 2, 8, 8, 8, 8, 8, 2, 8, 8, 8, 8, 8, 2, 8, 8, 8, 8, 8, 2, 8, 8, 8, 8, 8, 2]
  42. print_info: n_rot = 512
  43. print_info: n_swa = 1024
  44. print_info: is_swa_any = 1
  45. print_info: n_embd_head_k = 512
  46. print_info: n_embd_head_v = 512
  47. print_info: n_gqa = [2, 2, 2, 2, 2, 8, 2, 2, 2, 2, 2, 8, 2, 2, 2, 2, 2, 8, 2, 2, 2, 2, 2, 8, 2, 2, 2, 2, 2, 8]
  48. print_info: n_embd_k_gqa = [2048, 2048, 2048, 2048, 2048, 1024, 2048, 2048, 2048, 2048, 2048, 1024, 2048, 2048, 2048, 2048, 2048, 1024, 2048, 2048, 2048, 2048, 2048, 1024, 2048, 2048, 2048, 2048, 2048, 1024]
  49. print_info: n_embd_v_gqa = [2048, 2048, 2048, 2048, 2048, 1024, 2048, 2048, 2048, 2048, 2048, 1024, 2048, 2048, 2048, 2048, 2048, 1024, 2048, 2048, 2048, 2048, 2048, 1024, 2048, 2048, 2048, 2048, 2048, 1024]
  50. print_info: f_norm_eps = 0.0e+00
  51. print_info: f_norm_rms_eps = 1.0e-06
  52. print_info: f_clamp_kqv = 0.0e+00
  53. print_info: f_max_alibi_bias = 0.0e+00
  54. print_info: f_logit_scale = 0.0e+00
  55. print_info: f_attn_scale = 1.0e+00
  56. print_info: n_ff = 2112
  57. print_info: n_expert = 128
  58. print_info: n_expert_used = 8
  59. print_info: n_expert_groups = 0
  60. print_info: n_group_used = 0
  61. print_info: causal attn = 1
  62. print_info: pooling type = -1
  63. print_info: rope type = 2
  64. print_info: rope scaling = linear
  65. print_info: freq_base_train = 1000000.0
  66. print_info: freq_scale_train = 1
  67. print_info: freq_base_swa = 10000.0
  68. print_info: freq_scale_swa = 1
  69. print_info: n_embd_head_k_swa = 256
  70. print_info: n_embd_head_v_swa = 256
  71. print_info: n_rot_swa = 256
  72. print_info: n_ctx_orig_yarn = 262144
  73. print_info: rope_yarn_log_mul = 0.0000
  74. print_info: rope_finetuned = unknown
  75. print_info: model type = ?B
  76. print_info: model params = 25.23 B
  77. print_info: general.name = Gemma 4 26B A4B It
  78. print_info: vocab type = BPE
  79. print_info: n_vocab = 262144
  80. print_info: n_merges = 514906
  81. print_info: BOS token = 2 '<bos>'
  82. print_info: EOS token = 1 '<eos>'
  83. print_info: UNK token = 3 '<unk>'
  84. print_info: PAD token = 0 '<pad>'
  85. print_info: MASK token = 4 '<mask>'
  86. print_info: LF token = 107 '
  87. '
  88. print_info: EOG token = 1 '<eos>'
  89. print_info: EOG token = 50 '<|tool_response>'
  90. print_info: EOG token = 106 '<turn|>'
  91. print_info: EOG token = 212 '</s>'
  92. print_info: max token length = 93
  93. load_tensors: loading model tensors, this can take a while... (mmap = false, direct_io = false)
  94. str: cannot properly format tensor name output with suffix=weight bid=-1 xid=-1
  95. tensor blk.23.ffn_gate.weight (3 MiB q4_K) buffer type overridden to Vulkan_Host
  96. tensor blk.23.ffn_down.weight (6 MiB q8_0) buffer type overridden to Vulkan_Host
  97. tensor blk.23.ffn_gate_inp.weight (1 MiB f32) buffer type overridden to Vulkan_Host
  98. tensor blk.23.ffn_gate_inp.scale (0 MiB f32) buffer type overridden to Vulkan_Host
  99. tensor blk.23.ffn_gate_up_exps.weight (272 MiB q4_K) buffer type overridden to Vulkan_Host
  100. tensor blk.23.ffn_down_exps.weight (166 MiB q5_0) buffer type overridden to Vulkan_Host
  101. tensor blk.24.ffn_gate_up_exps.weight (272 MiB q4_K) buffer type overridden to Vulkan_Host
  102. tensor blk.24.ffn_down_exps.weight (257 MiB q8_0) buffer type overridden to Vulkan_Host
  103. tensor blk.25.ffn_gate_up_exps.weight (272 MiB q4_K) buffer type overridden to Vulkan_Host
  104. tensor blk.25.ffn_down_exps.weight (166 MiB q5_0) buffer type overridden to Vulkan_Host
  105. tensor blk.26.ffn_gate_up_exps.weight (272 MiB q4_K) buffer type overridden to Vulkan_Host
  106. tensor blk.26.ffn_down_exps.weight (166 MiB q5_0) buffer type overridden to Vulkan_Host
  107. tensor blk.27.ffn_gate_up_exps.weight (272 MiB q4_K) buffer type overridden to Vulkan_Host
  108. tensor blk.27.ffn_down_exps.weight (257 MiB q8_0) buffer type overridden to Vulkan_Host
  109. tensor blk.28.ffn_gate_up_exps.weight (272 MiB q4_K) buffer type overridden to Vulkan_Host
  110. tensor blk.28.ffn_down_exps.weight (166 MiB q5_0) buffer type overridden to Vulkan_Host
  111. tensor blk.29.ffn_gate_up_exps.weight (272 MiB q4_K) buffer type overridden to Vulkan_Host
  112. tensor blk.29.ffn_down_exps.weight (166 MiB q5_0) buffer type overridden to Vulkan_Host
  113. tensor blk.23.ffn_down_exps.scale (0 MiB f32) buffer type overridden to Vulkan_Host
  114. tensor blk.24.ffn_down_exps.scale (0 MiB f32) buffer type overridden to Vulkan_Host
  115. tensor blk.25.ffn_down_exps.scale (0 MiB f32) buffer type overridden to Vulkan_Host
  116. tensor blk.26.ffn_down_exps.scale (0 MiB f32) buffer type overridden to Vulkan_Host
  117. tensor blk.27.ffn_down_exps.scale (0 MiB f32) buffer type overridden to Vulkan_Host
  118. tensor blk.28.ffn_down_exps.scale (0 MiB f32) buffer type overridden to Vulkan_Host
  119. tensor blk.29.ffn_down_exps.scale (0 MiB f32) buffer type overridden to Vulkan_Host
  120. done_getting_tensors: tensor 'blk.23.ffn_gate.weight' (q4_K) (and 24 others) moved from Vulkan0, using Vulkan_Host instead
  121. load_tensors: offloading output layer to GPU
  122. load_tensors: offloading 29 repeating layers to GPU
  123. load_tensors: offloaded 31/31 layers to GPU
  124. load_tensors: Vulkan0 model buffer size = 12968.31 MiB
  125. load_tensors: Vulkan_Host model buffer size = 3840.01 MiB
  126. ...................................................................
  127. llama_context: constructing llama_context
  128. llama_context: n_seq_max = 1
  129. llama_context: n_ctx = 4352
  130. llama_context: n_ctx_seq = 4352
  131. llama_context: n_batch = 1024
  132. llama_context: n_ubatch = 512
  133. llama_context: causal_attn = 1
  134. llama_context: flash_attn = enabled
  135. llama_context: kv_unified = true
  136. llama_context: freq_base = 1000000.0
  137. llama_context: freq_scale = 1
  138. llama_context: n_ctx_seq (4352) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
  139. set_abort_callback: call
  140. llama_context: Vulkan_Host output buffer size = 1.00 MiB
  141. llama_kv_cache_iswa: creating non-SWA KV cache, size = 4352 cells
  142. llama_kv_cache: reusing layers:
  143. llama_kv_cache: - layer 0: no reuse
  144. llama_kv_cache: - layer 1: no reuse
  145. llama_kv_cache: - layer 2: no reuse
  146. llama_kv_cache: - layer 3: no reuse
  147. llama_kv_cache: - layer 4: no reuse
  148. llama_kv_cache: - layer 5: no reuse
  149. llama_kv_cache: - layer 6: no reuse
  150. llama_kv_cache: - layer 7: no reuse
  151. llama_kv_cache: - layer 8: no reuse
  152. llama_kv_cache: - layer 9: no reuse
  153. llama_kv_cache: - layer 10: no reuse
  154. llama_kv_cache: - layer 11: no reuse
  155. llama_kv_cache: - layer 12: no reuse
  156. llama_kv_cache: - layer 13: no reuse
  157. llama_kv_cache: - layer 14: no reuse
  158. llama_kv_cache: - layer 15: no reuse
  159. llama_kv_cache: - layer 16: no reuse
  160. llama_kv_cache: - layer 17: no reuse
  161. llama_kv_cache: - layer 18: no reuse
  162. llama_kv_cache: - layer 19: no reuse
  163. llama_kv_cache: - layer 20: no reuse
  164. llama_kv_cache: - layer 21: no reuse
  165. llama_kv_cache: - layer 22: no reuse
  166. llama_kv_cache: - layer 23: no reuse
  167. llama_kv_cache: - layer 24: no reuse
  168. llama_kv_cache: - layer 25: no reuse
  169. llama_kv_cache: - layer 26: no reuse
  170. llama_kv_cache: - layer 27: no reuse
  171. llama_kv_cache: - layer 28: no reuse
  172. llama_kv_cache: - layer 29: no reuse
  173. llama_kv_cache: Vulkan0 KV buffer size = 85.00 MiB
  174. llama_kv_cache: size = 85.00 MiB ( 4352 cells, 5 layers, 1/1 seqs), K (f16): 42.50 MiB, V (f16): 42.50 MiB
  175. llama_kv_cache: attn_rot_k = 0
  176. llama_kv_cache: attn_rot_v = 0
  177. llama_kv_cache_iswa: creating SWA KV cache, size = 1664 cells
  178. llama_kv_cache: reusing layers:
  179. llama_kv_cache: - layer 0: no reuse
  180. llama_kv_cache: - layer 1: no reuse
  181. llama_kv_cache: - layer 2: no reuse
  182. llama_kv_cache: - layer 3: no reuse
  183. llama_kv_cache: - layer 4: no reuse
  184. llama_kv_cache: - layer 5: no reuse
  185. llama_kv_cache: - layer 6: no reuse
  186. llama_kv_cache: - layer 7: no reuse
  187. llama_kv_cache: - layer 8: no reuse
  188. llama_kv_cache: - layer 9: no reuse
  189. llama_kv_cache: - layer 10: no reuse
  190. llama_kv_cache: - layer 11: no reuse
  191. llama_kv_cache: - layer 12: no reuse
  192. llama_kv_cache: - layer 13: no reuse
  193. llama_kv_cache: - layer 14: no reuse
  194. llama_kv_cache: - layer 15: no reuse
  195. llama_kv_cache: - layer 16: no reuse
  196. llama_kv_cache: - layer 17: no reuse
  197. llama_kv_cache: - layer 18: no reuse
  198. llama_kv_cache: - layer 19: no reuse
  199. llama_kv_cache: - layer 20: no reuse
  200. llama_kv_cache: - layer 21: no reuse
  201. llama_kv_cache: - layer 22: no reuse
  202. llama_kv_cache: - layer 23: no reuse
  203. llama_kv_cache: - layer 24: no reuse
  204. llama_kv_cache: - layer 25: no reuse
  205. llama_kv_cache: - layer 26: no reuse
  206. llama_kv_cache: - layer 27: no reuse
  207. llama_kv_cache: - layer 28: no reuse
  208. llama_kv_cache: - layer 29: no reuse
  209. llama_kv_cache: Vulkan0 KV buffer size = 325.00 MiB
  210. llama_kv_cache: size = 325.00 MiB ( 1664 cells, 25 layers, 1/1 seqs), K (f16): 162.50 MiB, V (f16): 162.50 MiB
  211. llama_kv_cache: attn_rot_k = 0
  212. llama_kv_cache: attn_rot_v = 0
  213. llama_context: enumerating backends
  214. llama_context: backend_ptrs.size() = 2
  215. sched_reserve: reserving ...
  216. sched_reserve: max_nodes = 5272
  217. sched_reserve: reserving full memory module
  218. sched_reserve: worst-case: n_tokens = 512, n_seqs = 1, n_outputs = 1
  219. sched_reserve: resolving fused Gated Delta Net support:
  220. sched_reserve: fused Gated Delta Net (autoregressive) enabled
  221. sched_reserve: fused Gated Delta Net (chunked) enabled
  222. sched_reserve: Vulkan0 compute buffer size = 521.62 MiB
  223. sched_reserve: Vulkan_Host compute buffer size = 22.78 MiB
  224. sched_reserve: graph nodes = 2647
  225. sched_reserve: graph splits = 20 (with bs=512), 22 (with bs=1)
  226. sched_reserve: reserve took 6.30 ms, sched copies = 1
  227. attach_threadpool: call
  228. Load Text Model OK: True
  229. Chat completion heuristic: Google Gemma 4 (26B and 31B)
  230. Embedded KoboldAI Lite loaded.
  231. Embedded API docs loaded.
  232. Llama.cpp UI loaded.
  233. ======
  234. Active Modules: TextGeneration
  235. Inactive Modules: ImageGeneration VoiceRecognition MultimodalVision MultimodalAudio NetworkMultiplayer ApiKeyPassword WebSearchProxy TextToSpeech VectorEmbeddings AdminControl MCPBridge MusicGen RouterMode
  236. Enabled APIs: KoboldCppApi OpenAiApi OllamaApi AnthropicApi
  237. Note: For third party Ollama API Emulation, you should set the port to 11434.
  238. Starting Kobold API on port 5003 at http://localhost:5003/api/
  239. Starting OpenAI Compatible API on port 5003 at http://localhost:5003/v1/
  240. Starting llama.cpp secondary WebUI at http://localhost:5003/lcpp/
  241. ======
  242. Please connect to custom endpoint at http://localhost:5003
  243.  
  244. The reported GGUF Arch is: gemma4
  245. Arch Category: 49
  246.  
  247. ---
  248. Identified as GGUF model.
  249. Attempting to Load...
  250. ---
  251.  
  252. SWA Mode IS ENABLED!
  253. Note that using SWA Mode cannot be used with Context Shifting
  254. Using automatic RoPE scaling for GGUF. If the model has custom RoPE settings, they'll be used directly instead!
  255. System Info: AVX = 1 | AVX_VNNI = 0 | AVX2 = 1 | AVX512 = 0 | AVX512_VBMI = 0 | AVX512_VNNI = 0 | AVX512_BF16 = 0 | AMX_INT8 = 0 | FMA = 1 | NEON = 0 | SVE = 0 | ARM_FMA = 0 | F16C = 1 | FP16_VA = 0 | RISCV_VECT = 0 | WASM_SIMD = 0 | SSE3 = 1 | SSSE3 = 1 | VSX = 0 | MATMUL_INT8 = 0 | LLAMAFILE = 1 |
  256.  
  257. Attempting to use llama.cpp's automating fitting code. This will override all your layer configs, may or may not work!
  258. Autofit Reserve Space: 1024 MB
  259. Autofit Success: 1, Autofit Result: -c 4224 -ngl 31 -ot blk\.23\.ffn_(gate|gate_up|down).*=CPU,blk\.24\.ffn_(up|down|gate_up|gate)_(ch|)exps=CPU,blk\.25\.ffn_(up|down|gate_up|gate)_(ch|)exps=CPU,blk\.26\.ffn_(up|down|gate_up|gate)_(ch|)exps=CPU,blk\.27\.ffn_(up|down|gate_up|gate)_(ch|)exps=CPU,blk\.28\.ffn_(up|down|gate_up|gate)_(ch|)exps=CPU,blk\.29\.ffn_(up|down|gate_up|gate)_(ch|)exps=CPU,blk\.30\.ffn_(up|down|gate_up|gate)_(ch|)exps=CPU
  260. Automatic RoPE Scaling: Using model internal value.
  261. Threadpool set to 15 threads and 15 blasthreads...
  262. Starting model warm up, please wait a moment...
  263.  
  264. [22:23:47] CtxLimit:21/4096, Amt:2/512, Init:0.00s, Process:0.12s (157.02T/s), Generate:0.04s (46.51T/s), Total:0.16sfree(): invalid pointer
Advertisement
Add Comment
Please, Sign In to add comment