Guest User

silverpine 1.7.3 multiple npc log

a guest
Jul 26th, 2026
24
0
Never
Not a member of Pastebin yet? Sign Up, it unlocks many cool features!
text 15.79 KB | None | 0 0
  1. ***
  2. Welcome to KoboldCpp - Version 1.111.2
  3. Loading Chat Completions Adapter: /tmp/_MEI0ykTEs/kcpp_adapters/AutoGuess.json
  4. Chat Completions Adapter Loaded
  5. System: Linux #1 SMP PREEMPT_DYNAMIC Wed, 08 Jul 2026 18:34:01 +0000 x86_64
  6. Detected Available GPU Memory: 16376 MB
  7. Detected Available RAM: 24498 MB
  8. Initializing dynamic library: koboldcpp_cublas.so
  9. ==========
  10. Namespace(admin=False, admindir='', adminpassword=None, adminunloadtimeout=0, analyze='', autofit=False, autofitpadding=1024, autoswapmode=False, batchsize=512, benchmark=None, blasthreads=0, chatcompletionsadapter='AutoGuess', cli=False, config=None, contextsize=4096, debugmode=0, defaultgenamt=1024, device='', downloaddir='', draftamount=8, draftgpulayers=999, draftgpusplit=None, draftmodel='', embeddingsgpu=False, embeddingsmaxctx=0, embeddingsmodel='', enableguidance=False, exportconfig='', exporttemplate='', failsafe=False, flashattention=False, forceversion=False, foreground=False, gendefaults='', gendefaultsoverwrite=False, genlimit=0, gpulayers=999, highpriority=False, hordeconfig=None, hordegenlen=0, hordekey='', hordemaxctx=0, hordemodelname='', hordeworkername='', host='', ignoremissing=False, jinja=False, jinja_kwargs='', jinja_tools=False, launch=False, lora=None, loramult=1.0, lowvram=False, maingpu=-1, maxrequestsize=32, mcpfile='', mmproj='', mmprojcpu=False, model=['/run/media/haristaan/sas/pobrane/Silverpine 1.7.3 Linux/Silverpine_Data/StreamingAssets/KoboldCPP/Gemma-4-Dense-Small.gguf'], model_param='/run/media/haristaan/sas/pobrane/Silverpine 1.7.3 Linux/Silverpine_Data/StreamingAssets/KoboldCPP/Gemma-4-Dense-Small.gguf', moecpu=0, moeexperts=-1, multiplayer=False, multiuser=100, musicdiffusion='', musicembeddings='', musicllm='', musiclowvram=False, musicvae='', noavx2=False, noblas=False, nobostoken=False, nocertify=False, nofastforward=True, noflashattention=False, nommap=False, nomodel=False, nopipelineparallel=False, noshift=True, onready='', overridekv='', overridenativecontext=0, overridetensors='', password=None, pipelineparallel=False, port=5003, port_param=5001, preloadstory='', prompt='', proxy_port=None, quantkv=0, quiet=True, ratelimit=0, remotetunnel=False, ropeconfig=[0.0, 10000.0], routermode=False, savedatafile='', sdclamped=0, sdclampedsoft=0, sdclip1='', sdclip2='', sdclipgpu=False, sdconfig=None, sdconvdirect='off', sdflashattention=False, sdgendefaults=False, sdlora=[], sdloramult=[1.0], sdmaingpu=-1, sdmodel='', sdnotile=False, sdoffloadcpu=False, sdphotomaker='', sdquant=0, sdt5xxl='', sdthreads=0, sdtiledvae=768, sdupscaler='', sdvae='', sdvaeauto=False, sdvaecpu=False, showgui=False, singleinstance=False, skiplauncher=True, smartcache=0, smartcontext=False, ssl=None, tensor_split=None, testmemory=False, threads=7, ttsdir='', ttsgpu=False, ttsmaxlen=4096, ttsmodel='', ttsthreads=0, ttswavtokenizer='', unpack='', usecpu=False, usecuda=['mmq'], usemlock=False, usemmap=False, useswa=True, usevulkan=None, version=False, visionmaxres=1024, websearch=False, whispermodel='')
  11. ==========
  12. Loading Text Model: /run/media/haristaan/sas/pobrane/Silverpine 1.7.3 Linux/Silverpine_Data/StreamingAssets/KoboldCPP/Gemma-4-Dense-Small.gguf
  13. ggml_cuda_init: found 1 CUDA devices (Total VRAM: 15949 MiB):
  14. Device 0: NVIDIA GeForce RTX 4070 Ti SUPER, compute capability 8.9, VMM: yes, VRAM: 15949 MiB
  15. llama_model_load_from_file_impl: using device CUDA0 (NVIDIA GeForce RTX 4070 Ti SUPER) (0000:26:00.0) - 14667 MiB free
  16. llama_model_loader: loaded meta data with 56 key-value pairs and 667 tensors from /run/media/haristaan/sas/pobrane/Silverpine 1.7.3 Linux/Silverpine_Data/StreamingAssets/KoboldCPP/Gemma-4-Dense-Small.gguf (version GGUF V3 (latest))
  17. print_info: file format = GGUF V3 (latest)
  18. print_info: file size = 11.78 GiB (8.50 BPW)
  19. init_tokenizer: initializing tokenizer for type 2
  20. load: 0 unused tokens
  21. load: control-looking token: 212 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden
  22. load: control-looking token: 50 '<|tool_response>' was not control-type; this is probably a bug in the model. its type will be overridden
  23. load: printing all EOG tokens:
  24. load: - 1 ('<eos>')
  25. load: - 50 ('<|tool_response>')
  26. load: - 106 ('<turn|>')
  27. load: - 212 ('</s>')
  28. load: special tokens cache size = 25
  29. load: token to piece cache size = 1.9445 MB
  30. print_info: arch = gemma4
  31. print_info: vocab_only = 0
  32. print_info: no_alloc = 0
  33. print_info: n_ctx_train = 131072
  34. print_info: n_embd = 3840
  35. print_info: n_embd_inp = 3840
  36. print_info: n_layer = 48
  37. print_info: n_head = 16
  38. print_info: n_head_kv = [8, 8, 8, 8, 8, 1, 8, 8, 8, 8, 8, 1, 8, 8, 8, 8, 8, 1, 8, 8, 8, 8, 8, 1, 8, 8, 8, 8, 8, 1, 8, 8, 8, 8, 8, 1, 8, 8, 8, 8, 8, 1, 8, 8, 8, 8, 8, 1]
  39. print_info: n_rot = 512
  40. print_info: n_swa = 1024
  41. print_info: is_swa_any = 1
  42. print_info: n_embd_head_k = 512
  43. print_info: n_embd_head_v = 512
  44. print_info: n_gqa = [2, 2, 2, 2, 2, 16, 2, 2, 2, 2, 2, 16, 2, 2, 2, 2, 2, 16, 2, 2, 2, 2, 2, 16, 2, 2, 2, 2, 2, 16, 2, 2, 2, 2, 2, 16, 2, 2, 2, 2, 2, 16, 2, 2, 2, 2, 2, 16]
  45. print_info: n_embd_k_gqa = [2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512]
  46. print_info: n_embd_v_gqa = [2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512]
  47. print_info: f_norm_eps = 0.0e+00
  48. print_info: f_norm_rms_eps = 1.0e-06
  49. print_info: f_clamp_kqv = 0.0e+00
  50. print_info: f_max_alibi_bias = 0.0e+00
  51. print_info: f_logit_scale = 0.0e+00
  52. print_info: f_attn_scale = 1.0e+00
  53. print_info: n_ff = 15360
  54. print_info: n_expert = 0
  55. print_info: n_expert_used = 0
  56. print_info: n_expert_groups = 0
  57. print_info: n_group_used = 0
  58. print_info: causal attn = 1
  59. print_info: pooling type = -1
  60. print_info: rope type = 2
  61. print_info: rope scaling = linear
  62. print_info: freq_base_train = 1000000.0
  63. print_info: freq_scale_train = 1
  64. print_info: freq_base_swa = 10000.0
  65. print_info: freq_scale_swa = 1
  66. print_info: n_embd_head_k_swa = 256
  67. print_info: n_embd_head_v_swa = 256
  68. print_info: n_rot_swa = 256
  69. print_info: n_ctx_orig_yarn = 131072
  70. print_info: rope_yarn_log_mul = 0.0000
  71. print_info: rope_finetuned = unknown
  72. print_info: model type = ?B
  73. print_info: model params = 11.91 B
  74. print_info: general.name = Gemma 4 12B It
  75. print_info: vocab type = BPE
  76. print_info: n_vocab = 262144
  77. print_info: n_merges = 514906
  78. print_info: BOS token = 2 '<bos>'
  79. print_info: EOS token = 1 '<eos>'
  80. print_info: UNK token = 3 '<unk>'
  81. print_info: PAD token = 0 '<pad>'
  82. print_info: MASK token = 4 '<mask>'
  83. print_info: LF token = 107 '
  84. '
  85. print_info: EOG token = 1 '<eos>'
  86. print_info: EOG token = 50 '<|tool_response>'
  87. print_info: EOG token = 106 '<turn|>'
  88. print_info: EOG token = 212 '</s>'
  89. print_info: max token length = 93
  90. load_tensors: loading model tensors, this can take a while... (mmap = false, direct_io = false)
  91. str: cannot properly format tensor name output with suffix=weight bid=-1 xid=-1
  92. load_tensors: offloading output layer to GPU
  93. load_tensors: offloading 47 repeating layers to GPU
  94. load_tensors: offloaded 49/49 layers to GPU
  95. load_tensors: CUDA0 model buffer size = 12067.72 MiB
  96. load_tensors: CUDA_Host model buffer size = 1020.00 MiB
  97. load_all_data: using async uploads for device CUDA0, buffer type CUDA0, backend CUDA0
  98. .......................................................................................
  99. llama_context: constructing llama_context
  100. llama_context: n_seq_max = 1
  101. llama_context: n_ctx = 4352
  102. llama_context: n_ctx_seq = 4352
  103. llama_context: n_batch = 1024
  104. llama_context: n_ubatch = 512
  105. llama_context: causal_attn = 1
  106. llama_context: flash_attn = enabled
  107. llama_context: kv_unified = true
  108. llama_context: freq_base = 1000000.0
  109. llama_context: freq_scale = 1
  110. llama_context: n_ctx_seq (4352) < n_ctx_train (131072) -- the full capacity of the model will not be utilized
  111. set_abort_callback: call
  112. llama_context: CUDA_Host output buffer size = 1.00 MiB
  113. llama_kv_cache_iswa: creating non-SWA KV cache, size = 4352 cells
  114. llama_kv_cache: reusing layers:
  115. llama_kv_cache: - layer 0: no reuse
  116. llama_kv_cache: - layer 1: no reuse
  117. llama_kv_cache: - layer 2: no reuse
  118. llama_kv_cache: - layer 3: no reuse
  119. llama_kv_cache: - layer 4: no reuse
  120. llama_kv_cache: - layer 5: no reuse
  121. llama_kv_cache: - layer 6: no reuse
  122. llama_kv_cache: - layer 7: no reuse
  123. llama_kv_cache: - layer 8: no reuse
  124. llama_kv_cache: - layer 9: no reuse
  125. llama_kv_cache: - layer 10: no reuse
  126. llama_kv_cache: - layer 11: no reuse
  127. llama_kv_cache: - layer 12: no reuse
  128. llama_kv_cache: - layer 13: no reuse
  129. llama_kv_cache: - layer 14: no reuse
  130. llama_kv_cache: - layer 15: no reuse
  131. llama_kv_cache: - layer 16: no reuse
  132. llama_kv_cache: - layer 17: no reuse
  133. llama_kv_cache: - layer 18: no reuse
  134. llama_kv_cache: - layer 19: no reuse
  135. llama_kv_cache: - layer 20: no reuse
  136. llama_kv_cache: - layer 21: no reuse
  137. llama_kv_cache: - layer 22: no reuse
  138. llama_kv_cache: - layer 23: no reuse
  139. llama_kv_cache: - layer 24: no reuse
  140. llama_kv_cache: - layer 25: no reuse
  141. llama_kv_cache: - layer 26: no reuse
  142. llama_kv_cache: - layer 27: no reuse
  143. llama_kv_cache: - layer 28: no reuse
  144. llama_kv_cache: - layer 29: no reuse
  145. llama_kv_cache: - layer 30: no reuse
  146. llama_kv_cache: - layer 31: no reuse
  147. llama_kv_cache: - layer 32: no reuse
  148. llama_kv_cache: - layer 33: no reuse
  149. llama_kv_cache: - layer 34: no reuse
  150. llama_kv_cache: - layer 35: no reuse
  151. llama_kv_cache: - layer 36: no reuse
  152. llama_kv_cache: - layer 37: no reuse
  153. llama_kv_cache: - layer 38: no reuse
  154. llama_kv_cache: - layer 39: no reuse
  155. llama_kv_cache: - layer 40: no reuse
  156. llama_kv_cache: - layer 41: no reuse
  157. llama_kv_cache: - layer 42: no reuse
  158. llama_kv_cache: - layer 43: no reuse
  159. llama_kv_cache: - layer 44: no reuse
  160. llama_kv_cache: - layer 45: no reuse
  161. llama_kv_cache: - layer 46: no reuse
  162. llama_kv_cache: - layer 47: no reuse
  163. llama_kv_cache: CUDA0 KV buffer size = 68.00 MiB
  164. llama_kv_cache: size = 68.00 MiB ( 4352 cells, 8 layers, 1/1 seqs), K (f16): 34.00 MiB, V (f16): 34.00 MiB
  165. llama_kv_cache: attn_rot_k = 0
  166. llama_kv_cache: attn_rot_v = 0
  167. llama_kv_cache_iswa: creating SWA KV cache, size = 1664 cells
  168. llama_kv_cache: reusing layers:
  169. llama_kv_cache: - layer 0: no reuse
  170. llama_kv_cache: - layer 1: no reuse
  171. llama_kv_cache: - layer 2: no reuse
  172. llama_kv_cache: - layer 3: no reuse
  173. llama_kv_cache: - layer 4: no reuse
  174. llama_kv_cache: - layer 5: no reuse
  175. llama_kv_cache: - layer 6: no reuse
  176. llama_kv_cache: - layer 7: no reuse
  177. llama_kv_cache: - layer 8: no reuse
  178. llama_kv_cache: - layer 9: no reuse
  179. llama_kv_cache: - layer 10: no reuse
  180. llama_kv_cache: - layer 11: no reuse
  181. llama_kv_cache: - layer 12: no reuse
  182. llama_kv_cache: - layer 13: no reuse
  183. llama_kv_cache: - layer 14: no reuse
  184. llama_kv_cache: - layer 15: no reuse
  185. llama_kv_cache: - layer 16: no reuse
  186. llama_kv_cache: - layer 17: no reuse
  187. llama_kv_cache: - layer 18: no reuse
  188. llama_kv_cache: - layer 19: no reuse
  189. llama_kv_cache: - layer 20: no reuse
  190. llama_kv_cache: - layer 21: no reuse
  191. llama_kv_cache: - layer 22: no reuse
  192. llama_kv_cache: - layer 23: no reuse
  193. llama_kv_cache: - layer 24: no reuse
  194. llama_kv_cache: - layer 25: no reuse
  195. llama_kv_cache: - layer 26: no reuse
  196. llama_kv_cache: - layer 27: no reuse
  197. llama_kv_cache: - layer 28: no reuse
  198. llama_kv_cache: - layer 29: no reuse
  199. llama_kv_cache: - layer 30: no reuse
  200. llama_kv_cache: - layer 31: no reuse
  201. llama_kv_cache: - layer 32: no reuse
  202. llama_kv_cache: - layer 33: no reuse
  203. llama_kv_cache: - layer 34: no reuse
  204. llama_kv_cache: - layer 35: no reuse
  205. llama_kv_cache: - layer 36: no reuse
  206. llama_kv_cache: - layer 37: no reuse
  207. llama_kv_cache: - layer 38: no reuse
  208. llama_kv_cache: - layer 39: no reuse
  209. llama_kv_cache: - layer 40: no reuse
  210. llama_kv_cache: - layer 41: no reuse
  211. llama_kv_cache: - layer 42: no reuse
  212. llama_kv_cache: - layer 43: no reuse
  213. llama_kv_cache: - layer 44: no reuse
  214. llama_kv_cache: - layer 45: no reuse
  215. llama_kv_cache: - layer 46: no reuse
  216. llama_kv_cache: - layer 47: no reuse
  217. llama_kv_cache: CUDA0 KV buffer size = 520.00 MiB
  218. llama_kv_cache: size = 520.00 MiB ( 1664 cells, 40 layers, 1/1 seqs), K (f16): 260.00 MiB, V (f16): 260.00 MiB
  219. llama_kv_cache: attn_rot_k = 0
  220. llama_kv_cache: attn_rot_v = 0
  221. llama_context: enumerating backends
  222. llama_context: backend_ptrs.size() = 2
  223. sched_reserve: reserving ...
  224. sched_reserve: max_nodes = 5344
  225. sched_reserve: reserving full memory module
  226. sched_reserve: worst-case: n_tokens = 512, n_seqs = 1, n_outputs = 1
  227. sched_reserve: resolving fused Gated Delta Net support:
  228. sched_reserve: fused Gated Delta Net (autoregressive) enabled
  229. sched_reserve: fused Gated Delta Net (chunked) enabled
  230. sched_reserve: CUDA0 compute buffer size = 519.50 MiB
  231. sched_reserve: CUDA_Host compute buffer size = 26.77 MiB
  232. sched_reserve: graph nodes = 1972
  233. sched_reserve: graph splits = 2
  234. sched_reserve: reserve took 10.98 ms, sched copies = 1
  235. attach_threadpool: call
  236. Load Text Model OK: True
  237. Chat completion heuristic: Google Gemma 4 (26B and 31B)
  238. Embedded KoboldAI Lite loaded.
  239. Embedded API docs loaded.
  240. Llama.cpp UI loaded.
  241. ======
  242. Active Modules: TextGeneration
  243. Inactive Modules: ImageGeneration VoiceRecognition MultimodalVision MultimodalAudio NetworkMultiplayer ApiKeyPassword WebSearchProxy TextToSpeech VectorEmbeddings AdminControl MCPBridge MusicGen RouterMode
  244. Enabled APIs: KoboldCppApi OpenAiApi OllamaApi AnthropicApi
  245. Note: For third party Ollama API Emulation, you should set the port to 11434.
  246. Starting Kobold API on port 5003 at http://localhost:5003/api/
  247. Starting OpenAI Compatible API on port 5003 at http://localhost:5003/v1/
  248. Starting llama.cpp secondary WebUI at http://localhost:5003/lcpp/
  249. ======
  250. Please connect to custom endpoint at http://localhost:5003
  251.  
  252. The reported GGUF Arch is: gemma4
  253. Arch Category: 49
  254.  
  255. ---
  256. Identified as GGUF model.
  257. Attempting to Load...
  258. ---
  259.  
  260. SWA Mode IS ENABLED!
  261. Note that using SWA Mode cannot be used with Context Shifting
  262. Using automatic RoPE scaling for GGUF. If the model has custom RoPE settings, they'll be used directly instead!
  263. System Info: AVX = 1 | AVX_VNNI = 0 | AVX2 = 1 | AVX512 = 0 | AVX512_VBMI = 0 | AVX512_VNNI = 0 | AVX512_BF16 = 0 | AMX_INT8 = 0 | FMA = 1 | NEON = 0 | SVE = 0 | ARM_FMA = 0 | F16C = 1 | FP16_VA = 0 | RISCV_VECT = 0 | WASM_SIMD = 0 | SSE3 = 1 | SSSE3 = 1 | VSX = 0 | MATMUL_INT8 = 0 | LLAMAFILE = 1 |
  264. CUDA MMQ: True
  265. ---
  266. Initializing CUDA/HIP, please wait, the following step may take a few minutes (only for first launch)...
  267. ---
  268. Automatic RoPE Scaling: Using model internal value.
  269. Threadpool set to 7 threads and 7 blasthreads...
  270. Starting model warm up, please wait a moment...
  271.  
  272. [22:44:28] CtxLimit:21/4096, Amt:2/512, Init:0.00s, Process:0.01s (1357.14T/s), Generate:0.06s (35.09T/s), Total:0.07s
  273. [22:45:05] CtxLimit:998/4096, Amt:64/512, Init:0.00s, Process:0.26s (3578.54T/s), Generate:3.27s (19.58T/s), Total:3.53s[PYI-11081:WARNING] Failed to remove temporary directory: /tmp/_MEI0ykTEs
  274.  
Advertisement
Add Comment
Please, Sign In to add comment