Not a member of Pastebin yet?
Sign Up,
it unlocks many cool features!
- ***
- Welcome to KoboldCpp - Version 1.111.2
- Loading Chat Completions Adapter: /tmp/_MEI0ykTEs/kcpp_adapters/AutoGuess.json
- Chat Completions Adapter Loaded
- System: Linux #1 SMP PREEMPT_DYNAMIC Wed, 08 Jul 2026 18:34:01 +0000 x86_64
- Detected Available GPU Memory: 16376 MB
- Detected Available RAM: 24498 MB
- Initializing dynamic library: koboldcpp_cublas.so
- ==========
- Namespace(admin=False, admindir='', adminpassword=None, adminunloadtimeout=0, analyze='', autofit=False, autofitpadding=1024, autoswapmode=False, batchsize=512, benchmark=None, blasthreads=0, chatcompletionsadapter='AutoGuess', cli=False, config=None, contextsize=4096, debugmode=0, defaultgenamt=1024, device='', downloaddir='', draftamount=8, draftgpulayers=999, draftgpusplit=None, draftmodel='', embeddingsgpu=False, embeddingsmaxctx=0, embeddingsmodel='', enableguidance=False, exportconfig='', exporttemplate='', failsafe=False, flashattention=False, forceversion=False, foreground=False, gendefaults='', gendefaultsoverwrite=False, genlimit=0, gpulayers=999, highpriority=False, hordeconfig=None, hordegenlen=0, hordekey='', hordemaxctx=0, hordemodelname='', hordeworkername='', host='', ignoremissing=False, jinja=False, jinja_kwargs='', jinja_tools=False, launch=False, lora=None, loramult=1.0, lowvram=False, maingpu=-1, maxrequestsize=32, mcpfile='', mmproj='', mmprojcpu=False, model=['/run/media/haristaan/sas/pobrane/Silverpine 1.7.3 Linux/Silverpine_Data/StreamingAssets/KoboldCPP/Gemma-4-Dense-Small.gguf'], model_param='/run/media/haristaan/sas/pobrane/Silverpine 1.7.3 Linux/Silverpine_Data/StreamingAssets/KoboldCPP/Gemma-4-Dense-Small.gguf', moecpu=0, moeexperts=-1, multiplayer=False, multiuser=100, musicdiffusion='', musicembeddings='', musicllm='', musiclowvram=False, musicvae='', noavx2=False, noblas=False, nobostoken=False, nocertify=False, nofastforward=True, noflashattention=False, nommap=False, nomodel=False, nopipelineparallel=False, noshift=True, onready='', overridekv='', overridenativecontext=0, overridetensors='', password=None, pipelineparallel=False, port=5003, port_param=5001, preloadstory='', prompt='', proxy_port=None, quantkv=0, quiet=True, ratelimit=0, remotetunnel=False, ropeconfig=[0.0, 10000.0], routermode=False, savedatafile='', sdclamped=0, sdclampedsoft=0, sdclip1='', sdclip2='', sdclipgpu=False, sdconfig=None, sdconvdirect='off', sdflashattention=False, sdgendefaults=False, sdlora=[], sdloramult=[1.0], sdmaingpu=-1, sdmodel='', sdnotile=False, sdoffloadcpu=False, sdphotomaker='', sdquant=0, sdt5xxl='', sdthreads=0, sdtiledvae=768, sdupscaler='', sdvae='', sdvaeauto=False, sdvaecpu=False, showgui=False, singleinstance=False, skiplauncher=True, smartcache=0, smartcontext=False, ssl=None, tensor_split=None, testmemory=False, threads=7, ttsdir='', ttsgpu=False, ttsmaxlen=4096, ttsmodel='', ttsthreads=0, ttswavtokenizer='', unpack='', usecpu=False, usecuda=['mmq'], usemlock=False, usemmap=False, useswa=True, usevulkan=None, version=False, visionmaxres=1024, websearch=False, whispermodel='')
- ==========
- Loading Text Model: /run/media/haristaan/sas/pobrane/Silverpine 1.7.3 Linux/Silverpine_Data/StreamingAssets/KoboldCPP/Gemma-4-Dense-Small.gguf
- ggml_cuda_init: found 1 CUDA devices (Total VRAM: 15949 MiB):
- Device 0: NVIDIA GeForce RTX 4070 Ti SUPER, compute capability 8.9, VMM: yes, VRAM: 15949 MiB
- llama_model_load_from_file_impl: using device CUDA0 (NVIDIA GeForce RTX 4070 Ti SUPER) (0000:26:00.0) - 14667 MiB free
- llama_model_loader: loaded meta data with 56 key-value pairs and 667 tensors from /run/media/haristaan/sas/pobrane/Silverpine 1.7.3 Linux/Silverpine_Data/StreamingAssets/KoboldCPP/Gemma-4-Dense-Small.gguf (version GGUF V3 (latest))
- print_info: file format = GGUF V3 (latest)
- print_info: file size = 11.78 GiB (8.50 BPW)
- init_tokenizer: initializing tokenizer for type 2
- load: 0 unused tokens
- load: control-looking token: 212 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden
- load: control-looking token: 50 '<|tool_response>' was not control-type; this is probably a bug in the model. its type will be overridden
- load: printing all EOG tokens:
- load: - 1 ('<eos>')
- load: - 50 ('<|tool_response>')
- load: - 106 ('<turn|>')
- load: - 212 ('</s>')
- load: special tokens cache size = 25
- load: token to piece cache size = 1.9445 MB
- print_info: arch = gemma4
- print_info: vocab_only = 0
- print_info: no_alloc = 0
- print_info: n_ctx_train = 131072
- print_info: n_embd = 3840
- print_info: n_embd_inp = 3840
- print_info: n_layer = 48
- print_info: n_head = 16
- print_info: n_head_kv = [8, 8, 8, 8, 8, 1, 8, 8, 8, 8, 8, 1, 8, 8, 8, 8, 8, 1, 8, 8, 8, 8, 8, 1, 8, 8, 8, 8, 8, 1, 8, 8, 8, 8, 8, 1, 8, 8, 8, 8, 8, 1, 8, 8, 8, 8, 8, 1]
- print_info: n_rot = 512
- print_info: n_swa = 1024
- print_info: is_swa_any = 1
- print_info: n_embd_head_k = 512
- print_info: n_embd_head_v = 512
- print_info: n_gqa = [2, 2, 2, 2, 2, 16, 2, 2, 2, 2, 2, 16, 2, 2, 2, 2, 2, 16, 2, 2, 2, 2, 2, 16, 2, 2, 2, 2, 2, 16, 2, 2, 2, 2, 2, 16, 2, 2, 2, 2, 2, 16, 2, 2, 2, 2, 2, 16]
- print_info: n_embd_k_gqa = [2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512]
- print_info: n_embd_v_gqa = [2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512, 2048, 2048, 2048, 2048, 2048, 512]
- print_info: f_norm_eps = 0.0e+00
- print_info: f_norm_rms_eps = 1.0e-06
- print_info: f_clamp_kqv = 0.0e+00
- print_info: f_max_alibi_bias = 0.0e+00
- print_info: f_logit_scale = 0.0e+00
- print_info: f_attn_scale = 1.0e+00
- print_info: n_ff = 15360
- print_info: n_expert = 0
- print_info: n_expert_used = 0
- print_info: n_expert_groups = 0
- print_info: n_group_used = 0
- print_info: causal attn = 1
- print_info: pooling type = -1
- print_info: rope type = 2
- print_info: rope scaling = linear
- print_info: freq_base_train = 1000000.0
- print_info: freq_scale_train = 1
- print_info: freq_base_swa = 10000.0
- print_info: freq_scale_swa = 1
- print_info: n_embd_head_k_swa = 256
- print_info: n_embd_head_v_swa = 256
- print_info: n_rot_swa = 256
- print_info: n_ctx_orig_yarn = 131072
- print_info: rope_yarn_log_mul = 0.0000
- print_info: rope_finetuned = unknown
- print_info: model type = ?B
- print_info: model params = 11.91 B
- print_info: general.name = Gemma 4 12B It
- print_info: vocab type = BPE
- print_info: n_vocab = 262144
- print_info: n_merges = 514906
- print_info: BOS token = 2 '<bos>'
- print_info: EOS token = 1 '<eos>'
- print_info: UNK token = 3 '<unk>'
- print_info: PAD token = 0 '<pad>'
- print_info: MASK token = 4 '<mask>'
- print_info: LF token = 107 '
- '
- print_info: EOG token = 1 '<eos>'
- print_info: EOG token = 50 '<|tool_response>'
- print_info: EOG token = 106 '<turn|>'
- print_info: EOG token = 212 '</s>'
- print_info: max token length = 93
- load_tensors: loading model tensors, this can take a while... (mmap = false, direct_io = false)
- str: cannot properly format tensor name output with suffix=weight bid=-1 xid=-1
- load_tensors: offloading output layer to GPU
- load_tensors: offloading 47 repeating layers to GPU
- load_tensors: offloaded 49/49 layers to GPU
- load_tensors: CUDA0 model buffer size = 12067.72 MiB
- load_tensors: CUDA_Host model buffer size = 1020.00 MiB
- load_all_data: using async uploads for device CUDA0, buffer type CUDA0, backend CUDA0
- .......................................................................................
- llama_context: constructing llama_context
- llama_context: n_seq_max = 1
- llama_context: n_ctx = 4352
- llama_context: n_ctx_seq = 4352
- llama_context: n_batch = 1024
- llama_context: n_ubatch = 512
- llama_context: causal_attn = 1
- llama_context: flash_attn = enabled
- llama_context: kv_unified = true
- llama_context: freq_base = 1000000.0
- llama_context: freq_scale = 1
- llama_context: n_ctx_seq (4352) < n_ctx_train (131072) -- the full capacity of the model will not be utilized
- set_abort_callback: call
- llama_context: CUDA_Host output buffer size = 1.00 MiB
- llama_kv_cache_iswa: creating non-SWA KV cache, size = 4352 cells
- llama_kv_cache: reusing layers:
- llama_kv_cache: - layer 0: no reuse
- llama_kv_cache: - layer 1: no reuse
- llama_kv_cache: - layer 2: no reuse
- llama_kv_cache: - layer 3: no reuse
- llama_kv_cache: - layer 4: no reuse
- llama_kv_cache: - layer 5: no reuse
- llama_kv_cache: - layer 6: no reuse
- llama_kv_cache: - layer 7: no reuse
- llama_kv_cache: - layer 8: no reuse
- llama_kv_cache: - layer 9: no reuse
- llama_kv_cache: - layer 10: no reuse
- llama_kv_cache: - layer 11: no reuse
- llama_kv_cache: - layer 12: no reuse
- llama_kv_cache: - layer 13: no reuse
- llama_kv_cache: - layer 14: no reuse
- llama_kv_cache: - layer 15: no reuse
- llama_kv_cache: - layer 16: no reuse
- llama_kv_cache: - layer 17: no reuse
- llama_kv_cache: - layer 18: no reuse
- llama_kv_cache: - layer 19: no reuse
- llama_kv_cache: - layer 20: no reuse
- llama_kv_cache: - layer 21: no reuse
- llama_kv_cache: - layer 22: no reuse
- llama_kv_cache: - layer 23: no reuse
- llama_kv_cache: - layer 24: no reuse
- llama_kv_cache: - layer 25: no reuse
- llama_kv_cache: - layer 26: no reuse
- llama_kv_cache: - layer 27: no reuse
- llama_kv_cache: - layer 28: no reuse
- llama_kv_cache: - layer 29: no reuse
- llama_kv_cache: - layer 30: no reuse
- llama_kv_cache: - layer 31: no reuse
- llama_kv_cache: - layer 32: no reuse
- llama_kv_cache: - layer 33: no reuse
- llama_kv_cache: - layer 34: no reuse
- llama_kv_cache: - layer 35: no reuse
- llama_kv_cache: - layer 36: no reuse
- llama_kv_cache: - layer 37: no reuse
- llama_kv_cache: - layer 38: no reuse
- llama_kv_cache: - layer 39: no reuse
- llama_kv_cache: - layer 40: no reuse
- llama_kv_cache: - layer 41: no reuse
- llama_kv_cache: - layer 42: no reuse
- llama_kv_cache: - layer 43: no reuse
- llama_kv_cache: - layer 44: no reuse
- llama_kv_cache: - layer 45: no reuse
- llama_kv_cache: - layer 46: no reuse
- llama_kv_cache: - layer 47: no reuse
- llama_kv_cache: CUDA0 KV buffer size = 68.00 MiB
- llama_kv_cache: size = 68.00 MiB ( 4352 cells, 8 layers, 1/1 seqs), K (f16): 34.00 MiB, V (f16): 34.00 MiB
- llama_kv_cache: attn_rot_k = 0
- llama_kv_cache: attn_rot_v = 0
- llama_kv_cache_iswa: creating SWA KV cache, size = 1664 cells
- llama_kv_cache: reusing layers:
- llama_kv_cache: - layer 0: no reuse
- llama_kv_cache: - layer 1: no reuse
- llama_kv_cache: - layer 2: no reuse
- llama_kv_cache: - layer 3: no reuse
- llama_kv_cache: - layer 4: no reuse
- llama_kv_cache: - layer 5: no reuse
- llama_kv_cache: - layer 6: no reuse
- llama_kv_cache: - layer 7: no reuse
- llama_kv_cache: - layer 8: no reuse
- llama_kv_cache: - layer 9: no reuse
- llama_kv_cache: - layer 10: no reuse
- llama_kv_cache: - layer 11: no reuse
- llama_kv_cache: - layer 12: no reuse
- llama_kv_cache: - layer 13: no reuse
- llama_kv_cache: - layer 14: no reuse
- llama_kv_cache: - layer 15: no reuse
- llama_kv_cache: - layer 16: no reuse
- llama_kv_cache: - layer 17: no reuse
- llama_kv_cache: - layer 18: no reuse
- llama_kv_cache: - layer 19: no reuse
- llama_kv_cache: - layer 20: no reuse
- llama_kv_cache: - layer 21: no reuse
- llama_kv_cache: - layer 22: no reuse
- llama_kv_cache: - layer 23: no reuse
- llama_kv_cache: - layer 24: no reuse
- llama_kv_cache: - layer 25: no reuse
- llama_kv_cache: - layer 26: no reuse
- llama_kv_cache: - layer 27: no reuse
- llama_kv_cache: - layer 28: no reuse
- llama_kv_cache: - layer 29: no reuse
- llama_kv_cache: - layer 30: no reuse
- llama_kv_cache: - layer 31: no reuse
- llama_kv_cache: - layer 32: no reuse
- llama_kv_cache: - layer 33: no reuse
- llama_kv_cache: - layer 34: no reuse
- llama_kv_cache: - layer 35: no reuse
- llama_kv_cache: - layer 36: no reuse
- llama_kv_cache: - layer 37: no reuse
- llama_kv_cache: - layer 38: no reuse
- llama_kv_cache: - layer 39: no reuse
- llama_kv_cache: - layer 40: no reuse
- llama_kv_cache: - layer 41: no reuse
- llama_kv_cache: - layer 42: no reuse
- llama_kv_cache: - layer 43: no reuse
- llama_kv_cache: - layer 44: no reuse
- llama_kv_cache: - layer 45: no reuse
- llama_kv_cache: - layer 46: no reuse
- llama_kv_cache: - layer 47: no reuse
- llama_kv_cache: CUDA0 KV buffer size = 520.00 MiB
- llama_kv_cache: size = 520.00 MiB ( 1664 cells, 40 layers, 1/1 seqs), K (f16): 260.00 MiB, V (f16): 260.00 MiB
- llama_kv_cache: attn_rot_k = 0
- llama_kv_cache: attn_rot_v = 0
- llama_context: enumerating backends
- llama_context: backend_ptrs.size() = 2
- sched_reserve: reserving ...
- sched_reserve: max_nodes = 5344
- sched_reserve: reserving full memory module
- sched_reserve: worst-case: n_tokens = 512, n_seqs = 1, n_outputs = 1
- sched_reserve: resolving fused Gated Delta Net support:
- sched_reserve: fused Gated Delta Net (autoregressive) enabled
- sched_reserve: fused Gated Delta Net (chunked) enabled
- sched_reserve: CUDA0 compute buffer size = 519.50 MiB
- sched_reserve: CUDA_Host compute buffer size = 26.77 MiB
- sched_reserve: graph nodes = 1972
- sched_reserve: graph splits = 2
- sched_reserve: reserve took 10.98 ms, sched copies = 1
- attach_threadpool: call
- Load Text Model OK: True
- Chat completion heuristic: Google Gemma 4 (26B and 31B)
- Embedded KoboldAI Lite loaded.
- Embedded API docs loaded.
- Llama.cpp UI loaded.
- ======
- Active Modules: TextGeneration
- Inactive Modules: ImageGeneration VoiceRecognition MultimodalVision MultimodalAudio NetworkMultiplayer ApiKeyPassword WebSearchProxy TextToSpeech VectorEmbeddings AdminControl MCPBridge MusicGen RouterMode
- Enabled APIs: KoboldCppApi OpenAiApi OllamaApi AnthropicApi
- Note: For third party Ollama API Emulation, you should set the port to 11434.
- Starting Kobold API on port 5003 at http://localhost:5003/api/
- Starting OpenAI Compatible API on port 5003 at http://localhost:5003/v1/
- Starting llama.cpp secondary WebUI at http://localhost:5003/lcpp/
- ======
- Please connect to custom endpoint at http://localhost:5003
- The reported GGUF Arch is: gemma4
- Arch Category: 49
- ---
- Identified as GGUF model.
- Attempting to Load...
- ---
- SWA Mode IS ENABLED!
- Note that using SWA Mode cannot be used with Context Shifting
- Using automatic RoPE scaling for GGUF. If the model has custom RoPE settings, they'll be used directly instead!
- System Info: AVX = 1 | AVX_VNNI = 0 | AVX2 = 1 | AVX512 = 0 | AVX512_VBMI = 0 | AVX512_VNNI = 0 | AVX512_BF16 = 0 | AMX_INT8 = 0 | FMA = 1 | NEON = 0 | SVE = 0 | ARM_FMA = 0 | F16C = 1 | FP16_VA = 0 | RISCV_VECT = 0 | WASM_SIMD = 0 | SSE3 = 1 | SSSE3 = 1 | VSX = 0 | MATMUL_INT8 = 0 | LLAMAFILE = 1 |
- CUDA MMQ: True
- ---
- Initializing CUDA/HIP, please wait, the following step may take a few minutes (only for first launch)...
- ---
- Automatic RoPE Scaling: Using model internal value.
- Threadpool set to 7 threads and 7 blasthreads...
- Starting model warm up, please wait a moment...
- [22:44:28] CtxLimit:21/4096, Amt:2/512, Init:0.00s, Process:0.01s (1357.14T/s), Generate:0.06s (35.09T/s), Total:0.07s
- [22:45:05] CtxLimit:998/4096, Amt:64/512, Init:0.00s, Process:0.26s (3578.54T/s), Generate:3.27s (19.58T/s), Total:3.53s[PYI-11081:WARNING] Failed to remove temporary directory: /tmp/_MEI0ykTEs
Advertisement
Add Comment
Please, Sign In to add comment