Not a member of Pastebin yet?
Sign Up,
it unlocks many cool features!
- ***
- Welcome to KoboldCpp - Version 1.111.2
- Loading Chat Completions Adapter: /tmp/_MEINgQyNs/kcpp_adapters/AutoGuess.json
- Chat Completions Adapter Loaded
- Auto Recommended GPU Layers: 15
- GPU layers is default: Will enable AutoFit for increased estimation accuracy.
- System: Linux #1 SMP PREEMPT_DYNAMIC Debian 6.12.88-1 (2026-05-15) x86_64
- Detected Available GPU Memory: 16368 MB
- Detected Available RAM: 55199 MB
- Initializing dynamic library: koboldcpp_vulkan.so
- ==========
- Namespace(admin=False, admindir='', adminpassword=None, adminunloadtimeout=0, analyze='', autofit=True, autofitpadding=1024, autoswapmode=False, batchsize=512, benchmark=None, blasthreads=0, chatcompletionsadapter='AutoGuess', cli=False, config=None, contextsize=4096, debugmode=0, defaultgenamt=1024, device='', downloaddir='', draftamount=8, draftgpulayers=999, draftgpusplit=None, draftmodel='', embeddingsgpu=False, embeddingsmaxctx=0, embeddingsmodel='', enableguidance=False, exportconfig='', exporttemplate='', failsafe=False, flashattention=False, forceversion=False, foreground=False, gendefaults='', gendefaultsoverwrite=False, genlimit=0, gpulayers=15, highpriority=False, hordeconfig=None, hordegenlen=0, hordekey='', hordemaxctx=0, hordemodelname='', hordeworkername='', host='', ignoremissing=False, jinja=False, jinja_kwargs='', jinja_tools=False, launch=False, lora=None, loramult=1.0, lowvram=False, maingpu=-1, maxrequestsize=32, mcpfile='', mmproj='', mmprojcpu=False, model=['/home/lakis/Documents/redacted/Silverpine_1.7.0c_Linux/Silverpine_Data/StreamingAssets/KoboldCPP/Gemma-4-Sparse.gguf'], model_param='/home/lakis/Documents/redacted/Silverpine_1.7.0c_Linux/Silverpine_Data/StreamingAssets/KoboldCPP/Gemma-4-Sparse.gguf', moecpu=0, moeexperts=-1, multiplayer=False, multiuser=100, musicdiffusion='', musicembeddings='', musicllm='', musiclowvram=False, musicvae='', noavx2=False, noblas=False, nobostoken=False, nocertify=False, nofastforward=True, noflashattention=False, nommap=False, nomodel=False, nopipelineparallel=False, noshift=True, onready='', overridekv='', overridenativecontext=0, overridetensors='', password=None, pipelineparallel=False, port=5003, port_param=5001, preloadstory='', prompt='', proxy_port=None, quantkv=0, quiet=True, ratelimit=0, remotetunnel=False, ropeconfig=[0.0, 10000.0], routermode=False, savedatafile='', sdclamped=0, sdclampedsoft=0, sdclip1='', sdclip2='', sdclipgpu=False, sdconfig=None, sdconvdirect='off', sdflashattention=False, sdgendefaults=False, sdlora=[], sdloramult=[1.0], sdmaingpu=-1, sdmodel='', sdnotile=False, sdoffloadcpu=False, sdphotomaker='', sdquant=0, sdt5xxl='', sdthreads=0, sdtiledvae=768, sdupscaler='', sdvae='', sdvaeauto=False, sdvaecpu=False, showgui=False, singleinstance=False, skiplauncher=True, smartcache=0, smartcontext=False, ssl=None, tensor_split=None, testmemory=False, threads=15, ttsdir='', ttsgpu=False, ttsmaxlen=4096, ttsmodel='', ttsthreads=0, ttswavtokenizer='', unpack='', usecpu=False, usecuda=None, usemlock=False, usemmap=False, useswa=True, usevulkan=[], version=False, visionmaxres=1024, websearch=False, whispermodel='')
- ==========
- Loading Text Model: /home/lakis/Documents/redacted/Silverpine_1.7.0c_Linux/Silverpine_Data/StreamingAssets/KoboldCPP/Gemma-4-Sparse.gguf
- ggml_vulkan: Found 2 Vulkan devices:
- ggml_vulkan: 0 = AMD Radeon RX 6900 XT (RADV NAVI21) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 32 | shared memory: 65536 | int dot: 1 | matrix cores: none
- ggml_vulkan: 1 = AMD Radeon Graphics (RADV RAPHAEL_MENDOCINO) (radv) | uma: 1 | fp16: 1 | bf16: 0 | warp size: 32 | shared memory: 65536 | int dot: 1 | matrix cores: none
- llama_model_load_from_file_impl: using device Vulkan0 (AMD Radeon RX 6900 XT (RADV NAVI21)) (0000:03:00.0) - 14947 MiB free
- llama_model_loader: loaded meta data with 54 key-value pairs and 658 tensors from /home/lakis/Documents/redacted/Silverpine_1.7.0c_Linux/Silverpine_Data/StreamingAssets/KoboldCPP/Gemma-4-Sparse.gguf (version GGUF V3 (latest))
- print_info: file format = GGUF V3 (latest)
- print_info: file size = 15.85 GiB (5.40 BPW)
- init_tokenizer: initializing tokenizer for type 2
- load: 0 unused tokens
- load: control-looking token: 212 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden
- load: control-looking token: 50 '<|tool_response>' was not control-type; this is probably a bug in the model. its type will be overridden
- load: printing all EOG tokens:
- load: - 1 ('<eos>')
- load: - 50 ('<|tool_response>')
- load: - 106 ('<turn|>')
- load: - 212 ('</s>')
- load: special tokens cache size = 25
- load: token to piece cache size = 1.9445 MB
- print_info: arch = gemma4
- print_info: vocab_only = 0
- print_info: no_alloc = 0
- print_info: n_ctx_train = 262144
- print_info: n_embd = 2816
- print_info: n_embd_inp = 2816
- print_info: n_layer = 30
- print_info: n_head = 16
- print_info: n_head_kv = [8, 8, 8, 8, 8, 2, 8, 8, 8, 8, 8, 2, 8, 8, 8, 8, 8, 2, 8, 8, 8, 8, 8, 2, 8, 8, 8, 8, 8, 2]
- print_info: n_rot = 512
- print_info: n_swa = 1024
- print_info: is_swa_any = 1
- print_info: n_embd_head_k = 512
- print_info: n_embd_head_v = 512
- print_info: n_gqa = [2, 2, 2, 2, 2, 8, 2, 2, 2, 2, 2, 8, 2, 2, 2, 2, 2, 8, 2, 2, 2, 2, 2, 8, 2, 2, 2, 2, 2, 8]
- print_info: n_embd_k_gqa = [2048, 2048, 2048, 2048, 2048, 1024, 2048, 2048, 2048, 2048, 2048, 1024, 2048, 2048, 2048, 2048, 2048, 1024, 2048, 2048, 2048, 2048, 2048, 1024, 2048, 2048, 2048, 2048, 2048, 1024]
- print_info: n_embd_v_gqa = [2048, 2048, 2048, 2048, 2048, 1024, 2048, 2048, 2048, 2048, 2048, 1024, 2048, 2048, 2048, 2048, 2048, 1024, 2048, 2048, 2048, 2048, 2048, 1024, 2048, 2048, 2048, 2048, 2048, 1024]
- print_info: f_norm_eps = 0.0e+00
- print_info: f_norm_rms_eps = 1.0e-06
- print_info: f_clamp_kqv = 0.0e+00
- print_info: f_max_alibi_bias = 0.0e+00
- print_info: f_logit_scale = 0.0e+00
- print_info: f_attn_scale = 1.0e+00
- print_info: n_ff = 2112
- print_info: n_expert = 128
- print_info: n_expert_used = 8
- print_info: n_expert_groups = 0
- print_info: n_group_used = 0
- print_info: causal attn = 1
- print_info: pooling type = -1
- print_info: rope type = 2
- print_info: rope scaling = linear
- print_info: freq_base_train = 1000000.0
- print_info: freq_scale_train = 1
- print_info: freq_base_swa = 10000.0
- print_info: freq_scale_swa = 1
- print_info: n_embd_head_k_swa = 256
- print_info: n_embd_head_v_swa = 256
- print_info: n_rot_swa = 256
- print_info: n_ctx_orig_yarn = 262144
- print_info: rope_yarn_log_mul = 0.0000
- print_info: rope_finetuned = unknown
- print_info: model type = ?B
- print_info: model params = 25.23 B
- print_info: general.name = Gemma 4 26B A4B It
- print_info: vocab type = BPE
- print_info: n_vocab = 262144
- print_info: n_merges = 514906
- print_info: BOS token = 2 '<bos>'
- print_info: EOS token = 1 '<eos>'
- print_info: UNK token = 3 '<unk>'
- print_info: PAD token = 0 '<pad>'
- print_info: MASK token = 4 '<mask>'
- print_info: LF token = 107 '
- '
- print_info: EOG token = 1 '<eos>'
- print_info: EOG token = 50 '<|tool_response>'
- print_info: EOG token = 106 '<turn|>'
- print_info: EOG token = 212 '</s>'
- print_info: max token length = 93
- load_tensors: loading model tensors, this can take a while... (mmap = false, direct_io = false)
- str: cannot properly format tensor name output with suffix=weight bid=-1 xid=-1
- tensor blk.23.ffn_gate.weight (3 MiB q4_K) buffer type overridden to Vulkan_Host
- tensor blk.23.ffn_down.weight (6 MiB q8_0) buffer type overridden to Vulkan_Host
- tensor blk.23.ffn_gate_inp.weight (1 MiB f32) buffer type overridden to Vulkan_Host
- tensor blk.23.ffn_gate_inp.scale (0 MiB f32) buffer type overridden to Vulkan_Host
- tensor blk.23.ffn_gate_up_exps.weight (272 MiB q4_K) buffer type overridden to Vulkan_Host
- tensor blk.23.ffn_down_exps.weight (166 MiB q5_0) buffer type overridden to Vulkan_Host
- tensor blk.24.ffn_gate_up_exps.weight (272 MiB q4_K) buffer type overridden to Vulkan_Host
- tensor blk.24.ffn_down_exps.weight (257 MiB q8_0) buffer type overridden to Vulkan_Host
- tensor blk.25.ffn_gate_up_exps.weight (272 MiB q4_K) buffer type overridden to Vulkan_Host
- tensor blk.25.ffn_down_exps.weight (166 MiB q5_0) buffer type overridden to Vulkan_Host
- tensor blk.26.ffn_gate_up_exps.weight (272 MiB q4_K) buffer type overridden to Vulkan_Host
- tensor blk.26.ffn_down_exps.weight (166 MiB q5_0) buffer type overridden to Vulkan_Host
- tensor blk.27.ffn_gate_up_exps.weight (272 MiB q4_K) buffer type overridden to Vulkan_Host
- tensor blk.27.ffn_down_exps.weight (257 MiB q8_0) buffer type overridden to Vulkan_Host
- tensor blk.28.ffn_gate_up_exps.weight (272 MiB q4_K) buffer type overridden to Vulkan_Host
- tensor blk.28.ffn_down_exps.weight (166 MiB q5_0) buffer type overridden to Vulkan_Host
- tensor blk.29.ffn_gate_up_exps.weight (272 MiB q4_K) buffer type overridden to Vulkan_Host
- tensor blk.29.ffn_down_exps.weight (166 MiB q5_0) buffer type overridden to Vulkan_Host
- tensor blk.23.ffn_down_exps.scale (0 MiB f32) buffer type overridden to Vulkan_Host
- tensor blk.24.ffn_down_exps.scale (0 MiB f32) buffer type overridden to Vulkan_Host
- tensor blk.25.ffn_down_exps.scale (0 MiB f32) buffer type overridden to Vulkan_Host
- tensor blk.26.ffn_down_exps.scale (0 MiB f32) buffer type overridden to Vulkan_Host
- tensor blk.27.ffn_down_exps.scale (0 MiB f32) buffer type overridden to Vulkan_Host
- tensor blk.28.ffn_down_exps.scale (0 MiB f32) buffer type overridden to Vulkan_Host
- tensor blk.29.ffn_down_exps.scale (0 MiB f32) buffer type overridden to Vulkan_Host
- done_getting_tensors: tensor 'blk.23.ffn_gate.weight' (q4_K) (and 24 others) moved from Vulkan0, using Vulkan_Host instead
- load_tensors: offloading output layer to GPU
- load_tensors: offloading 29 repeating layers to GPU
- load_tensors: offloaded 31/31 layers to GPU
- load_tensors: Vulkan0 model buffer size = 12968.31 MiB
- load_tensors: Vulkan_Host model buffer size = 3840.01 MiB
- ...................................................................
- llama_context: constructing llama_context
- llama_context: n_seq_max = 1
- llama_context: n_ctx = 4352
- llama_context: n_ctx_seq = 4352
- llama_context: n_batch = 1024
- llama_context: n_ubatch = 512
- llama_context: causal_attn = 1
- llama_context: flash_attn = enabled
- llama_context: kv_unified = true
- llama_context: freq_base = 1000000.0
- llama_context: freq_scale = 1
- llama_context: n_ctx_seq (4352) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
- set_abort_callback: call
- llama_context: Vulkan_Host output buffer size = 1.00 MiB
- llama_kv_cache_iswa: creating non-SWA KV cache, size = 4352 cells
- llama_kv_cache: reusing layers:
- llama_kv_cache: - layer 0: no reuse
- llama_kv_cache: - layer 1: no reuse
- llama_kv_cache: - layer 2: no reuse
- llama_kv_cache: - layer 3: no reuse
- llama_kv_cache: - layer 4: no reuse
- llama_kv_cache: - layer 5: no reuse
- llama_kv_cache: - layer 6: no reuse
- llama_kv_cache: - layer 7: no reuse
- llama_kv_cache: - layer 8: no reuse
- llama_kv_cache: - layer 9: no reuse
- llama_kv_cache: - layer 10: no reuse
- llama_kv_cache: - layer 11: no reuse
- llama_kv_cache: - layer 12: no reuse
- llama_kv_cache: - layer 13: no reuse
- llama_kv_cache: - layer 14: no reuse
- llama_kv_cache: - layer 15: no reuse
- llama_kv_cache: - layer 16: no reuse
- llama_kv_cache: - layer 17: no reuse
- llama_kv_cache: - layer 18: no reuse
- llama_kv_cache: - layer 19: no reuse
- llama_kv_cache: - layer 20: no reuse
- llama_kv_cache: - layer 21: no reuse
- llama_kv_cache: - layer 22: no reuse
- llama_kv_cache: - layer 23: no reuse
- llama_kv_cache: - layer 24: no reuse
- llama_kv_cache: - layer 25: no reuse
- llama_kv_cache: - layer 26: no reuse
- llama_kv_cache: - layer 27: no reuse
- llama_kv_cache: - layer 28: no reuse
- llama_kv_cache: - layer 29: no reuse
- llama_kv_cache: Vulkan0 KV buffer size = 85.00 MiB
- llama_kv_cache: size = 85.00 MiB ( 4352 cells, 5 layers, 1/1 seqs), K (f16): 42.50 MiB, V (f16): 42.50 MiB
- llama_kv_cache: attn_rot_k = 0
- llama_kv_cache: attn_rot_v = 0
- llama_kv_cache_iswa: creating SWA KV cache, size = 1664 cells
- llama_kv_cache: reusing layers:
- llama_kv_cache: - layer 0: no reuse
- llama_kv_cache: - layer 1: no reuse
- llama_kv_cache: - layer 2: no reuse
- llama_kv_cache: - layer 3: no reuse
- llama_kv_cache: - layer 4: no reuse
- llama_kv_cache: - layer 5: no reuse
- llama_kv_cache: - layer 6: no reuse
- llama_kv_cache: - layer 7: no reuse
- llama_kv_cache: - layer 8: no reuse
- llama_kv_cache: - layer 9: no reuse
- llama_kv_cache: - layer 10: no reuse
- llama_kv_cache: - layer 11: no reuse
- llama_kv_cache: - layer 12: no reuse
- llama_kv_cache: - layer 13: no reuse
- llama_kv_cache: - layer 14: no reuse
- llama_kv_cache: - layer 15: no reuse
- llama_kv_cache: - layer 16: no reuse
- llama_kv_cache: - layer 17: no reuse
- llama_kv_cache: - layer 18: no reuse
- llama_kv_cache: - layer 19: no reuse
- llama_kv_cache: - layer 20: no reuse
- llama_kv_cache: - layer 21: no reuse
- llama_kv_cache: - layer 22: no reuse
- llama_kv_cache: - layer 23: no reuse
- llama_kv_cache: - layer 24: no reuse
- llama_kv_cache: - layer 25: no reuse
- llama_kv_cache: - layer 26: no reuse
- llama_kv_cache: - layer 27: no reuse
- llama_kv_cache: - layer 28: no reuse
- llama_kv_cache: - layer 29: no reuse
- llama_kv_cache: Vulkan0 KV buffer size = 325.00 MiB
- llama_kv_cache: size = 325.00 MiB ( 1664 cells, 25 layers, 1/1 seqs), K (f16): 162.50 MiB, V (f16): 162.50 MiB
- llama_kv_cache: attn_rot_k = 0
- llama_kv_cache: attn_rot_v = 0
- llama_context: enumerating backends
- llama_context: backend_ptrs.size() = 2
- sched_reserve: reserving ...
- sched_reserve: max_nodes = 5272
- sched_reserve: reserving full memory module
- sched_reserve: worst-case: n_tokens = 512, n_seqs = 1, n_outputs = 1
- sched_reserve: resolving fused Gated Delta Net support:
- sched_reserve: fused Gated Delta Net (autoregressive) enabled
- sched_reserve: fused Gated Delta Net (chunked) enabled
- sched_reserve: Vulkan0 compute buffer size = 521.62 MiB
- sched_reserve: Vulkan_Host compute buffer size = 22.78 MiB
- sched_reserve: graph nodes = 2647
- sched_reserve: graph splits = 20 (with bs=512), 22 (with bs=1)
- sched_reserve: reserve took 6.30 ms, sched copies = 1
- attach_threadpool: call
- Load Text Model OK: True
- Chat completion heuristic: Google Gemma 4 (26B and 31B)
- Embedded KoboldAI Lite loaded.
- Embedded API docs loaded.
- Llama.cpp UI loaded.
- ======
- Active Modules: TextGeneration
- Inactive Modules: ImageGeneration VoiceRecognition MultimodalVision MultimodalAudio NetworkMultiplayer ApiKeyPassword WebSearchProxy TextToSpeech VectorEmbeddings AdminControl MCPBridge MusicGen RouterMode
- Enabled APIs: KoboldCppApi OpenAiApi OllamaApi AnthropicApi
- Note: For third party Ollama API Emulation, you should set the port to 11434.
- Starting Kobold API on port 5003 at http://localhost:5003/api/
- Starting OpenAI Compatible API on port 5003 at http://localhost:5003/v1/
- Starting llama.cpp secondary WebUI at http://localhost:5003/lcpp/
- ======
- Please connect to custom endpoint at http://localhost:5003
- The reported GGUF Arch is: gemma4
- Arch Category: 49
- ---
- Identified as GGUF model.
- Attempting to Load...
- ---
- SWA Mode IS ENABLED!
- Note that using SWA Mode cannot be used with Context Shifting
- Using automatic RoPE scaling for GGUF. If the model has custom RoPE settings, they'll be used directly instead!
- System Info: AVX = 1 | AVX_VNNI = 0 | AVX2 = 1 | AVX512 = 0 | AVX512_VBMI = 0 | AVX512_VNNI = 0 | AVX512_BF16 = 0 | AMX_INT8 = 0 | FMA = 1 | NEON = 0 | SVE = 0 | ARM_FMA = 0 | F16C = 1 | FP16_VA = 0 | RISCV_VECT = 0 | WASM_SIMD = 0 | SSE3 = 1 | SSSE3 = 1 | VSX = 0 | MATMUL_INT8 = 0 | LLAMAFILE = 1 |
- Attempting to use llama.cpp's automating fitting code. This will override all your layer configs, may or may not work!
- Autofit Reserve Space: 1024 MB
- Autofit Success: 1, Autofit Result: -c 4224 -ngl 31 -ot blk\.23\.ffn_(gate|gate_up|down).*=CPU,blk\.24\.ffn_(up|down|gate_up|gate)_(ch|)exps=CPU,blk\.25\.ffn_(up|down|gate_up|gate)_(ch|)exps=CPU,blk\.26\.ffn_(up|down|gate_up|gate)_(ch|)exps=CPU,blk\.27\.ffn_(up|down|gate_up|gate)_(ch|)exps=CPU,blk\.28\.ffn_(up|down|gate_up|gate)_(ch|)exps=CPU,blk\.29\.ffn_(up|down|gate_up|gate)_(ch|)exps=CPU,blk\.30\.ffn_(up|down|gate_up|gate)_(ch|)exps=CPU
- Automatic RoPE Scaling: Using model internal value.
- Threadpool set to 15 threads and 15 blasthreads...
- Starting model warm up, please wait a moment...
- [22:23:47] CtxLimit:21/4096, Amt:2/512, Init:0.00s, Process:0.12s (157.02T/s), Generate:0.04s (46.51T/s), Total:0.16sfree(): invalid pointer
Advertisement
Add Comment
Please, Sign In to add comment