Not a member of Pastebin yet?
Sign Up,
it unlocks many cool features!
- === Parallel Benchmark Results ===
- Model: gemma-3-4b-it-q4_0.gguf
- Date: Sun Oct 19 19:35:41 EDT 2025
- macOS Version: 26.0.1
- Hardware: Mac16,6
- Chip: Apple M4 Max
- Total RAM: 128.00 GB
- CPU Cores: 16
- GPU Layers: 999
- Flash Attention: 1
- === Benchmark Parameters ===
- Context Size: 300000
- Ubatch Size: 2048
- Prompt Tokens: 4096,8192
- Generation Tokens: 32
- Parallel Requests: 1,2,4,8,16,32
- === Results ===
- ggml_metal_library_init: using embedded metal library
- ggml_metal_library_init: loaded in 0.005 sec
- ggml_metal_device_init: GPU name: Apple M4 Max
- ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
- ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
- ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
- ggml_metal_device_init: simdgroup reduction = true
- ggml_metal_device_init: simdgroup matrix mul. = true
- ggml_metal_device_init: has unified memory = true
- ggml_metal_device_init: has bfloat = true
- ggml_metal_device_init: use residency sets = true
- ggml_metal_device_init: use shared buffers = true
- ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
- build: 6789 (3d4e86bb) with Apple clang version 17.0.0 (clang-1700.3.19.1) for arm64-apple-darwin25.0.0
- llama_model_load_from_file_impl: using device Metal (Apple M4 Max) (unknown id) - 110100 MiB free
- llama_model_loader: loaded meta data with 40 key-value pairs and 444 tensors from /Usersusermodels/dgx/gemma-3-4b-it-q4_0.gguf (version GGUF V3 (latest))
- llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
- llama_model_loader: - kv 0: general.architecture str = gemma3
- llama_model_loader: - kv 1: general.type str = model
- llama_model_loader: - kv 2: general.name str = Gemma-3-4B-It
- llama_model_loader: - kv 3: general.finetune str = it
- llama_model_loader: - kv 4: general.basename str = Gemma-3-4B-It
- llama_model_loader: - kv 5: general.quantized_by str = Unsloth
- llama_model_loader: - kv 6: general.size_label str = 4B
- llama_model_loader: - kv 7: general.repo_url str = https://huggingface.co/unsloth
- llama_model_loader: - kv 8: gemma3.context_length u32 = 131072
- llama_model_loader: - kv 9: gemma3.embedding_length u32 = 2560
- llama_model_loader: - kv 10: gemma3.block_count u32 = 34
- llama_model_loader: - kv 11: gemma3.feed_forward_length u32 = 10240
- llama_model_loader: - kv 12: gemma3.attention.head_count u32 = 8
- llama_model_loader: - kv 13: gemma3.attention.layer_norm_rms_epsilon f32 = 0.000001
- llama_model_loader: - kv 14: gemma3.attention.key_length u32 = 256
- llama_model_loader: - kv 15: gemma3.attention.value_length u32 = 256
- llama_model_loader: - kv 16: gemma3.rope.freq_base f32 = 1000000.000000
- llama_model_loader: - kv 17: gemma3.attention.sliding_window u32 = 1024
- llama_model_loader: - kv 18: gemma3.attention.head_count_kv u32 = 4
- llama_model_loader: - kv 19: gemma3.rope.scaling.type str = linear
- llama_model_loader: - kv 20: gemma3.rope.scaling.factor f32 = 8.000000
- llama_model_loader: - kv 21: tokenizer.ggml.model str = llama
- llama_model_loader: - kv 22: tokenizer.ggml.pre str = default
- llama_model_loader: - kv 23: tokenizer.ggml.tokens arr[str,262208] = ["<pad>", "<eos>", "<bos>", "<unk>", ...
- llama_model_loader: - kv 24: tokenizer.ggml.scores arr[f32,262208] = [-1000.000000, -1000.000000, -1000.00...
- llama_model_loader: - kv 25: tokenizer.ggml.token_type arr[i32,262208] = [3, 3, 3, 3, 3, 4, 3, 3, 3, 3, 3, 3, ...
- llama_model_loader: - kv 26: tokenizer.ggml.bos_token_id u32 = 2
- llama_model_loader: - kv 27: tokenizer.ggml.eos_token_id u32 = 106
- llama_model_loader: - kv 28: tokenizer.ggml.unknown_token_id u32 = 3
- llama_model_loader: - kv 29: tokenizer.ggml.padding_token_id u32 = 0
- llama_model_loader: - kv 30: tokenizer.ggml.add_bos_token bool = true
- llama_model_loader: - kv 31: tokenizer.ggml.add_eos_token bool = false
- llama_model_loader: - kv 32: tokenizer.chat_template str = {{ bos_token }}\n{%- if messages[0]['r...
- llama_model_loader: - kv 33: tokenizer.ggml.add_space_prefix bool = false
- llama_model_loader: - kv 34: general.quantization_version u32 = 2
- llama_model_loader: - kv 35: general.file_type u32 = 2
- llama_model_loader: - kv 36: quantize.imatrix.file str = gemma-3-4b-it-GGUF/imatrix_unsloth.dat
- llama_model_loader: - kv 37: quantize.imatrix.dataset str = unsloth_calibration_gemma-3-4b-it.txt
- llama_model_loader: - kv 38: quantize.imatrix.entries_count i32 = 238
- llama_model_loader: - kv 39: quantize.imatrix.chunks_count i32 = 663
- llama_model_loader: - type f32: 205 tensors
- llama_model_loader: - type q4_0: 234 tensors
- llama_model_loader: - type q4_1: 4 tensors
- llama_model_loader: - type q6_K: 1 tensors
- print_info: file format = GGUF V3 (latest)
- print_info: file type = Q4_0
- print_info: file size = 2.20 GiB (4.87 BPW)
- load: printing all EOG tokens:
- load: - 106 ('<end_of_turn>')
- load: special tokens cache size = 6415
- load: token to piece cache size = 1.9446 MB
- print_info: arch = gemma3
- print_info: vocab_only = 0
- print_info: n_ctx_train = 131072
- print_info: n_embd = 2560
- print_info: n_layer = 34
- print_info: n_head = 8
- print_info: n_head_kv = 4
- print_info: n_rot = 256
- print_info: n_swa = 1024
- print_info: is_swa_any = 1
- print_info: n_embd_head_k = 256
- print_info: n_embd_head_v = 256
- print_info: n_gqa = 2
- print_info: n_embd_k_gqa = 1024
- print_info: n_embd_v_gqa = 1024
- print_info: f_norm_eps = 0.0e+00
- print_info: f_norm_rms_eps = 1.0e-06
- print_info: f_clamp_kqv = 0.0e+00
- print_info: f_max_alibi_bias = 0.0e+00
- print_info: f_logit_scale = 0.0e+00
- print_info: f_attn_scale = 6.2e-02
- print_info: n_ff = 10240
- print_info: n_expert = 0
- print_info: n_expert_used = 0
- print_info: causal attn = 1
- print_info: pooling type = 0
- print_info: rope type = 2
- print_info: rope scaling = linear
- print_info: freq_base_train = 1000000.0
- print_info: freq_scale_train = 0.125
- print_info: n_ctx_orig_yarn = 131072
- print_info: rope_finetuned = unknown
- print_info: model type = 4B
- print_info: model params = 3.88 B
- print_info: general.name = Gemma-3-4B-It
- print_info: vocab type = SPM
- print_info: n_vocab = 262208
- print_info: n_merges = 0
- print_info: BOS token = 2 '<bos>'
- print_info: EOS token = 106 '<end_of_turn>'
- print_info: EOT token = 106 '<end_of_turn>'
- print_info: UNK token = 3 '<unk>'
- print_info: PAD token = 0 '<pad>'
- print_info: LF token = 248 '<0x0A>'
- print_info: EOG token = 106 '<end_of_turn>'
- print_info: max token length = 48
- load_tensors: loading model tensors, this can take a while... (mmap = true)
- load_tensors: offloading 34 repeating layers to GPU
- load_tensors: offloading output layer to GPU
- load_tensors: offloaded 35/35 layers to GPU
- load_tensors: Metal_Mapped model buffer size = 2254.03 MiB
- load_tensors: CPU_Mapped model buffer size = 525.13 MiB
- .................................................................
- llama_init_from_model: model default pooling_type is [0], but [-1] was specified
- llama_context: constructing llama_context
- llama_context: n_seq_max = 32
- llama_context: n_ctx = 300000
- llama_context: n_ctx_per_seq = 9375
- llama_context: n_batch = 2048
- llama_context: n_ubatch = 2048
- llama_context: causal_attn = 1
- llama_context: flash_attn = enabled
- llama_context: kv_unified = false
- llama_context: freq_base = 1000000.0
- llama_context: freq_scale = 0.125
- llama_context: n_ctx_per_seq (9375) < n_ctx_train (131072) -- the full capacity of the model will not be utilized
- ggml_metal_init: allocating
- ggml_metal_init: picking default device: Apple M4 Max
- ggml_metal_init: use bfloat = true
- ggml_metal_init: use fusion = true
- ggml_metal_init: use concurrency = true
- ggml_metal_init: use graph optimize = true
- llama_context: CPU output buffer size = 32.01 MiB
- llama_kv_cache_iswa: creating non-SWA KV cache, size = 9472 cells
- llama_kv_cache: Metal KV buffer size = 5920.00 MiB
- llama_kv_cache: size = 5920.00 MiB ( 9472 cells, 5 layers, 32/32 seqs), K (f16): 2960.00 MiB, V (f16): 2960.00 MiB
- llama_kv_cache_iswa: creating SWA KV cache, size = 3072 cells
- llama_kv_cache: Metal KV buffer size = 11136.00 MiB
- llama_kv_cache: size = 11136.00 MiB ( 3072 cells, 29 layers, 32/32 seqs), K (f16): 5568.00 MiB, V (f16): 5568.00 MiB
- llama_context: Metal compute buffer size = 2068.50 MiB
- llama_context: CPU compute buffer size = 118.09 MiB
- llama_context: graph nodes = 1437
- llama_context: graph splits = 2
- main: n_kv_max = 303104, n_batch = 2048, n_ubatch = 2048, flash_attn = 1, is_pp_shared = 0, n_gpu_layers = 999, n_threads = 12, n_threads_batch = 12
- | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s |
- |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
- | 4096 | 32 | 1 | 4128 | 2.468 | 1659.39 | 0.256 | 125.11 | 2.724 | 1515.34 |
- | 4096 | 32 | 2 | 8256 | 5.196 | 1576.59 | 0.417 | 153.41 | 5.613 | 1470.82 |
- | 4096 | 32 | 4 | 16512 | 10.522 | 1557.12 | 0.768 | 166.72 | 11.290 | 1462.56 |
- | 4096 | 32 | 8 | 33024 | 20.653 | 1586.57 | 1.445 | 177.19 | 22.098 | 1494.42 |
- | 4096 | 32 | 16 | 66048 | 40.388 | 1622.67 | 1.398 | 366.34 | 41.785 | 1580.65 |
- | 4096 | 32 | 32 | 132096 | 80.829 | 1621.61 | 1.767 | 579.39 | 82.596 | 1599.30 |
- | 8192 | 32 | 1 | 8224 | 5.408 | 1514.82 | 0.272 | 117.79 | 5.680 | 1448.00 |
- | 8192 | 32 | 2 | 16448 | 10.336 | 1585.11 | 0.443 | 144.52 | 10.779 | 1525.92 |
- | 8192 | 32 | 4 | 32896 | 20.636 | 1587.88 | 0.770 | 166.28 | 21.406 | 1536.76 |
- | 8192 | 32 | 8 | 65792 | 42.082 | 1557.33 | 1.505 | 170.05 | 43.588 | 1509.42 |
- | 8192 | 32 | 16 | 131584 | 83.043 | 1578.37 | 1.555 | 329.35 | 84.597 | 1555.41 |
- | 8192 | 32 | 32 | 263168 | 174.075 | 1505.92 | 2.071 | 494.53 | 176.146 | 1494.04 |
- llama_perf_context_print: load time = 1067.47 ms
- llama_perf_context_print: prompt eval time = 507816.42 ms / 778128 tokens ( 0.65 ms per token, 1532.30 tokens per second)
- llama_perf_context_print: eval time = 527.06 ms / 64 runs ( 8.24 ms per token, 121.43 tokens per second)
- llama_perf_context_print: total time = 509409.84 ms / 778192 tokens
- llama_perf_context_print: graphs reused = 372
- ggml_metal_free: deallocating
- === Parallel Benchmark Results ===
- Model: glm-4.5-air-q4_k_m-00001-of-00002.gguf
- Date: Sun Oct 19 20:50:36 EDT 2025
- macOS Version: 26.0.1
- Hardware: Mac16,6
- Chip: Apple M4 Max
- Total RAM: 128.00 GB
- CPU Cores: 16
- GPU Layers: 999
- Flash Attention: 1
- === Benchmark Parameters ===
- Context Size: 300000
- Ubatch Size: 2048
- Prompt Tokens: 4096,8192
- Generation Tokens: 32
- Parallel Requests: 1,2,4,8,16,32
- === Results ===
- ggml_metal_library_init: using embedded metal library
- ggml_metal_library_init: loaded in 0.007 sec
- ggml_metal_device_init: GPU name: Apple M4 Max
- ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
- ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
- ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
- ggml_metal_device_init: simdgroup reduction = true
- ggml_metal_device_init: simdgroup matrix mul. = true
- ggml_metal_device_init: has unified memory = true
- ggml_metal_device_init: has bfloat = true
- ggml_metal_device_init: use residency sets = true
- ggml_metal_device_init: use shared buffers = true
- ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
- build: 6789 (3d4e86bb) with Apple clang version 17.0.0 (clang-1700.3.19.1) for arm64-apple-darwin25.0.0
- llama_model_load_from_file_impl: using device Metal (Apple M4 Max) (unknown id) - 110100 MiB free
- llama_model_loader: additional 1 GGUFs metadata loaded.
- llama_model_loader: loaded meta data with 48 key-value pairs and 803 tensors from /Usersusermodels/dgx/glm-4.5-air-q4_k_m-00001-of-00002.gguf (version GGUF V3 (latest))
- llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
- llama_model_loader: - kv 0: general.architecture str = glm4moe
- llama_model_loader: - kv 1: general.type str = model
- llama_model_loader: - kv 2: general.name str = GLM 4.5 Air
- llama_model_loader: - kv 3: general.size_label str = 128x9.4B
- llama_model_loader: - kv 4: general.license str = mit
- llama_model_loader: - kv 5: general.tags arr[str,1] = ["text-generation"]
- llama_model_loader: - kv 6: general.languages arr[str,2] = ["en", "zh"]
- llama_model_loader: - kv 7: glm4moe.block_count u32 = 47
- llama_model_loader: - kv 8: glm4moe.context_length u32 = 131072
- llama_model_loader: - kv 9: glm4moe.embedding_length u32 = 4096
- llama_model_loader: - kv 10: glm4moe.feed_forward_length u32 = 10944
- llama_model_loader: - kv 11: glm4moe.attention.head_count u32 = 96
- llama_model_loader: - kv 12: glm4moe.attention.head_count_kv u32 = 8
- llama_model_loader: - kv 13: glm4moe.rope.freq_base f32 = 1000000.000000
- llama_model_loader: - kv 14: glm4moe.attention.layer_norm_rms_epsilon f32 = 0.000010
- llama_model_loader: - kv 15: glm4moe.expert_used_count u32 = 8
- llama_model_loader: - kv 16: glm4moe.attention.key_length u32 = 128
- llama_model_loader: - kv 17: glm4moe.attention.value_length u32 = 128
- llama_model_loader: - kv 18: glm4moe.rope.dimension_count u32 = 64
- llama_model_loader: - kv 19: glm4moe.expert_count u32 = 128
- llama_model_loader: - kv 20: glm4moe.expert_feed_forward_length u32 = 1408
- llama_model_loader: - kv 21: glm4moe.expert_shared_count u32 = 1
- llama_model_loader: - kv 22: glm4moe.leading_dense_block_count u32 = 1
- llama_model_loader: - kv 23: glm4moe.expert_gating_func u32 = 2
- llama_model_loader: - kv 24: glm4moe.expert_weights_scale f32 = 1.000000
- llama_model_loader: - kv 25: glm4moe.expert_weights_norm bool = true
- llama_model_loader: - kv 26: glm4moe.nextn_predict_layers u32 = 1
- llama_model_loader: - kv 27: tokenizer.ggml.model str = gpt2
- llama_model_loader: - kv 28: tokenizer.ggml.pre str = glm4
- llama_model_loader: - kv 29: tokenizer.ggml.tokens arr[str,151552] = ["!", "\"", "#", "$", "%", "&", "'", ...
- llama_model_loader: - kv 30: tokenizer.ggml.token_type arr[i32,151552] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
- llama_model_loader: - kv 31: tokenizer.ggml.merges arr[str,318088] = ["Ġ Ġ", "Ġ ĠĠĠ", "ĠĠ ĠĠ", "...
- llama_model_loader: - kv 32: tokenizer.ggml.eos_token_id u32 = 151329
- llama_model_loader: - kv 33: tokenizer.ggml.padding_token_id u32 = 151329
- llama_model_loader: - kv 34: tokenizer.ggml.bos_token_id u32 = 151331
- llama_model_loader: - kv 35: tokenizer.ggml.eot_token_id u32 = 151336
- llama_model_loader: - kv 36: tokenizer.ggml.unknown_token_id u32 = 151329
- llama_model_loader: - kv 37: tokenizer.ggml.eom_token_id u32 = 151338
- llama_model_loader: - kv 38: tokenizer.chat_template str = [gMASK]<sop>\n{%- if tools -%}\n<|syste...
- llama_model_loader: - kv 39: general.quantization_version u32 = 2
- llama_model_loader: - kv 40: general.file_type u32 = 15
- llama_model_loader: - kv 41: quantize.imatrix.file str = /models_out/GLM-4.5-Air-GGUF/zai-org_...
- llama_model_loader: - kv 42: quantize.imatrix.dataset str = /training_dir/calibration_datav3.txt
- llama_model_loader: - kv 43: quantize.imatrix.entries_count u32 = 502
- llama_model_loader: - kv 44: quantize.imatrix.chunks_count u32 = 485
- llama_model_loader: - kv 45: split.no u16 = 0
- llama_model_loader: - kv 46: split.tensors.count i32 = 803
- llama_model_loader: - kv 47: split.count u16 = 2
- llama_model_loader: - type f32: 331 tensors
- llama_model_loader: - type q5_0: 24 tensors
- llama_model_loader: - type q8_0: 160 tensors
- llama_model_loader: - type q4_K: 169 tensors
- llama_model_loader: - type q5_K: 47 tensors
- llama_model_loader: - type q6_K: 72 tensors
- print_info: file format = GGUF V3 (latest)
- print_info: file type = Q4_K - Medium
- print_info: file size = 68.45 GiB (5.32 BPW)
- load: special_eot_id is not in special_eog_ids - the tokenizer config may be incorrect
- load: special_eom_id is not in special_eog_ids - the tokenizer config may be incorrect
- load: printing all EOG tokens:
- load: - 151329 ('<|endoftext|>')
- load: - 151336 ('<|user|>')
- load: - 151338 ('<|observation|>')
- load: special tokens cache size = 36
- load: token to piece cache size = 0.9713 MB
- print_info: arch = glm4moe
- print_info: vocab_only = 0
- print_info: n_ctx_train = 131072
- print_info: n_embd = 4096
- print_info: n_layer = 47
- print_info: n_head = 96
- print_info: n_head_kv = 8
- print_info: n_rot = 64
- print_info: n_swa = 0
- print_info: is_swa_any = 0
- print_info: n_embd_head_k = 128
- print_info: n_embd_head_v = 128
- print_info: n_gqa = 12
- print_info: n_embd_k_gqa = 1024
- print_info: n_embd_v_gqa = 1024
- print_info: f_norm_eps = 0.0e+00
- print_info: f_norm_rms_eps = 1.0e-05
- print_info: f_clamp_kqv = 0.0e+00
- print_info: f_max_alibi_bias = 0.0e+00
- print_info: f_logit_scale = 0.0e+00
- print_info: f_attn_scale = 0.0e+00
- print_info: n_ff = 10944
- print_info: n_expert = 128
- print_info: n_expert_used = 8
- print_info: causal attn = 1
- print_info: pooling type = 0
- print_info: rope type = 2
- print_info: rope scaling = linear
- print_info: freq_base_train = 1000000.0
- print_info: freq_scale_train = 1
- print_info: n_ctx_orig_yarn = 131072
- print_info: rope_finetuned = unknown
- print_info: model type = 106B.A12B
- print_info: model params = 110.47 B
- print_info: general.name = GLM 4.5 Air
- print_info: vocab type = BPE
- print_info: n_vocab = 151552
- print_info: n_merges = 318088
- print_info: BOS token = 151331 '[gMASK]'
- print_info: EOS token = 151329 '<|endoftext|>'
- print_info: EOT token = 151336 '<|user|>'
- print_info: EOM token = 151338 '<|observation|>'
- print_info: UNK token = 151329 '<|endoftext|>'
- print_info: PAD token = 151329 '<|endoftext|>'
- print_info: LF token = 198 'Ċ'
- print_info: FIM PRE token = 151347 '<|code_prefix|>'
- print_info: FIM SUF token = 151349 '<|code_suffix|>'
- print_info: FIM MID token = 151348 '<|code_middle|>'
- print_info: EOG token = 151329 '<|endoftext|>'
- print_info: EOG token = 151336 '<|user|>'
- print_info: EOG token = 151338 '<|observation|>'
- print_info: max token length = 1024
- load_tensors: loading model tensors, this can take a while... (mmap = true)
- model has unused tensor blk.46.attn_norm.weight (size = 16384 bytes) -- ignoring
- model has unused tensor blk.46.attn_q.weight (size = 28311552 bytes) -- ignoring
- model has unused tensor blk.46.attn_k.weight (size = 4456448 bytes) -- ignoring
- model has unused tensor blk.46.attn_v.weight (size = 3440640 bytes) -- ignoring
- model has unused tensor blk.46.attn_q.bias (size = 49152 bytes) -- ignoring
- model has unused tensor blk.46.attn_k.bias (size = 4096 bytes) -- ignoring
- model has unused tensor blk.46.attn_v.bias (size = 4096 bytes) -- ignoring
- model has unused tensor blk.46.attn_output.weight (size = 34603008 bytes) -- ignoring
- model has unused tensor blk.46.post_attention_norm.weight (size = 16384 bytes) -- ignoring
- model has unused tensor blk.46.ffn_gate_inp.weight (size = 2097152 bytes) -- ignoring
- model has unused tensor blk.46.exp_probs_b.bias (size = 512 bytes) -- ignoring
- model has unused tensor blk.46.ffn_gate_exps.weight (size = 415236096 bytes) -- ignoring
- model has unused tensor blk.46.ffn_down_exps.weight (size = 784334848 bytes) -- ignoring
- model has unused tensor blk.46.ffn_up_exps.weight (size = 415236096 bytes) -- ignoring
- model has unused tensor blk.46.ffn_gate_shexp.weight (size = 6127616 bytes) -- ignoring
- model has unused tensor blk.46.ffn_down_shexp.weight (size = 6127616 bytes) -- ignoring
- model has unused tensor blk.46.ffn_up_shexp.weight (size = 6127616 bytes) -- ignoring
- model has unused tensor blk.46.nextn.eh_proj.weight (size = 18874368 bytes) -- ignoring
- model has unused tensor blk.46.nextn.enorm.weight (size = 16384 bytes) -- ignoring
- model has unused tensor blk.46.nextn.hnorm.weight (size = 16384 bytes) -- ignoring
- model has unused tensor blk.46.nextn.embed_tokens.weight (size = 349175808 bytes) -- ignoring
- model has unused tensor blk.46.nextn.shared_head_head.weight (size = 349175808 bytes) -- ignoring
- model has unused tensor blk.46.nextn.shared_head_norm.weight (size = 16384 bytes) -- ignoring
- load_tensors: offloading 47 repeating layers to GPU
- load_tensors: offloading output layer to GPU
- load_tensors: offloaded 48/48 layers to GPU
- load_tensors: Metal_Mapped model buffer size = 37977.33 MiB
- load_tensors: Metal_Mapped model buffer size = 29799.45 MiB
- load_tensors: CPU_Mapped model buffer size = 333.00 MiB
- ..................................................................................................
- llama_init_from_model: model default pooling_type is [0], but [-1] was specified
- llama_context: constructing llama_context
- llama_context: n_seq_max = 32
- llama_context: n_ctx = 300000
- llama_context: n_ctx_per_seq = 9375
- llama_context: n_batch = 2048
- llama_context: n_ubatch = 2048
- llama_context: causal_attn = 1
- llama_context: flash_attn = enabled
- llama_context: kv_unified = false
- llama_context: freq_base = 1000000.0
- llama_context: freq_scale = 1
- llama_context: n_ctx_per_seq (9375) < n_ctx_train (131072) -- the full capacity of the model will not be utilized
- ggml_metal_init: allocating
- ggml_metal_init: picking default device: Apple M4 Max
- ggml_metal_init: use bfloat = true
- ggml_metal_init: use fusion = true
- ggml_metal_init: use concurrency = true
- ggml_metal_init: use graph optimize = true
- llama_context: CPU output buffer size = 18.50 MiB
- llama_kv_cache: Metal KV buffer size = 54464.00 MiB
- llama_kv_cache: size = 54464.00 MiB ( 9472 cells, 46 layers, 32/32 seqs), K (f16): 27232.00 MiB, V (f16): 27232.00 MiB
- llama_context: Metal compute buffer size = 12387.07 MiB
- llama_context: CPU compute buffer size = 106.05 MiB
- llama_context: graph nodes = 3193
- llama_context: graph splits = 2
- main: n_kv_max = 303104, n_batch = 2048, n_ubatch = 2048, flash_attn = 1, is_pp_shared = 0, n_gpu_layers = 999, n_threads = 12, n_threads_batch = 12
- | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s |
- |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
- | 4096 | 32 | 1 | 4128 | 40.042 | 102.29 | 437.052 | 0.07 | 477.094 | 8.65 |
- | 4096 | 32 | 2 | 8256 | 77.299 | 105.98 | 501.545 | 0.13 | 578.844 | 14.26 |
- | 4096 | 32 | 4 | 16512 | 161.838 | 101.24 | 499.020 | 0.26 | 660.858 | 24.99 |
- | 4096 | 32 | 8 | 33024 | 326.307 | 100.42 | 517.406 | 0.49 | 843.713 | 39.14 |
- | 4096 | 32 | 16 | 66048 | 658.933 | 99.46 | 538.631 | 0.95 | 1197.564 | 55.15 |
- | 4096 | 32 | 32 | 132096 | 1352.850 | 96.89 | 585.793 | 1.75 | 1938.643 | 68.14 |
- | 8192 | 32 | 1 | 8224 | 89.265 | 91.77 | 583.327 | 0.05 | 672.591 | 12.23 |
- | 8192 | 32 | 2 | 16448 | 178.881 | 91.59 | 583.590 | 0.11 | 762.471 | 21.57 |
- | 8192 | 32 | 4 | 32896 | 363.187 | 90.22 | 597.051 | 0.21 | 960.238 | 34.26 |
- | 8192 | 32 | 8 | 65792 | 740.917 | 88.45 | 586.614 | 0.44 | 1327.531 | 49.56 |
- | 8192 | 32 | 16 | 131584 | 1457.437 | 89.93 | 598.005 | 0.86 | 2055.442 | 64.02 |
- | 8192 | 32 | 32 | 263168 | 3294.007 | 79.58 | 1375.630 | 0.74 | 4669.637 | 56.36 |
- llama_perf_context_print: load time = 16229.62 ms
- llama_perf_context_print: prompt eval time = 15132036.12 ms / 778128 tokens ( 19.45 ms per token, 51.42 tokens per second)
- llama_perf_context_print: eval time = 1020375.68 ms / 64 runs (15943.37 ms per token, 0.06 tokens per second)
- llama_perf_context_print: total time = 16160981.34 ms / 778192 tokens
- llama_perf_context_print: graphs reused = 372
- ggml_metal_free: deallocating
- === Parallel Benchmark Results ===
- Model: gpt-oss-120b-mxfp4-00001-of-00003.gguf
- Date: Sun Oct 19 17:59:54 EDT 2025
- macOS Version: 26.0.1
- Hardware: Mac16,6
- Chip: Apple M4 Max
- Total RAM: 128.00 GB
- CPU Cores: 16
- GPU Layers: 999
- Flash Attention: 1
- === Benchmark Parameters ===
- Context Size: 300000
- Ubatch Size: 2048
- Prompt Tokens: 4096,8192
- Generation Tokens: 32
- Parallel Requests: 1,2,4,8,16,32
- === Results ===
- ggml_metal_library_init: using embedded metal library
- ggml_metal_library_init: loaded in 0.006 sec
- ggml_metal_device_init: GPU name: Apple M4 Max
- ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
- ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
- ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
- ggml_metal_device_init: simdgroup reduction = true
- ggml_metal_device_init: simdgroup matrix mul. = true
- ggml_metal_device_init: has unified memory = true
- ggml_metal_device_init: has bfloat = true
- ggml_metal_device_init: use residency sets = true
- ggml_metal_device_init: use shared buffers = true
- ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
- build: 6789 (3d4e86bb) with Apple clang version 17.0.0 (clang-1700.3.19.1) for arm64-apple-darwin25.0.0
- llama_model_load_from_file_impl: using device Metal (Apple M4 Max) (unknown id) - 110100 MiB free
- llama_model_loader: additional 2 GGUFs metadata loaded.
- llama_model_loader: loaded meta data with 38 key-value pairs and 687 tensors from /Usersusermodels/dgx/gpt-oss-120b-mxfp4-00001-of-00003.gguf (version GGUF V3 (latest))
- llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
- llama_model_loader: - kv 0: general.architecture str = gpt-oss
- llama_model_loader: - kv 1: general.type str = model
- llama_model_loader: - kv 2: general.name str = Gpt Oss 120b
- llama_model_loader: - kv 3: general.basename str = gpt-oss
- llama_model_loader: - kv 4: general.size_label str = 120B
- llama_model_loader: - kv 5: general.license str = apache-2.0
- llama_model_loader: - kv 6: general.tags arr[str,2] = ["vllm", "text-generation"]
- llama_model_loader: - kv 7: gpt-oss.block_count u32 = 36
- llama_model_loader: - kv 8: gpt-oss.context_length u32 = 131072
- llama_model_loader: - kv 9: gpt-oss.embedding_length u32 = 2880
- llama_model_loader: - kv 10: gpt-oss.feed_forward_length u32 = 2880
- llama_model_loader: - kv 11: gpt-oss.attention.head_count u32 = 64
- llama_model_loader: - kv 12: gpt-oss.attention.head_count_kv u32 = 8
- llama_model_loader: - kv 13: gpt-oss.rope.freq_base f32 = 150000.000000
- llama_model_loader: - kv 14: gpt-oss.attention.layer_norm_rms_epsilon f32 = 0.000010
- llama_model_loader: - kv 15: gpt-oss.expert_count u32 = 128
- llama_model_loader: - kv 16: gpt-oss.expert_used_count u32 = 4
- llama_model_loader: - kv 17: gpt-oss.attention.key_length u32 = 64
- llama_model_loader: - kv 18: gpt-oss.attention.value_length u32 = 64
- llama_model_loader: - kv 19: gpt-oss.attention.sliding_window u32 = 128
- llama_model_loader: - kv 20: gpt-oss.expert_feed_forward_length u32 = 2880
- llama_model_loader: - kv 21: gpt-oss.rope.scaling.type str = yarn
- llama_model_loader: - kv 22: gpt-oss.rope.scaling.factor f32 = 32.000000
- llama_model_loader: - kv 23: gpt-oss.rope.scaling.original_context_length u32 = 4096
- llama_model_loader: - kv 24: tokenizer.ggml.model str = gpt2
- llama_model_loader: - kv 25: tokenizer.ggml.pre str = gpt-4o
- llama_model_loader: - kv 26: tokenizer.ggml.tokens arr[str,201088] = ["!", "\"", "#", "$", "%", "&", "'", ...
- llama_model_loader: - kv 27: tokenizer.ggml.token_type arr[i32,201088] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
- llama_model_loader: - kv 28: tokenizer.ggml.merges arr[str,446189] = ["Ġ Ġ", "Ġ ĠĠĠ", "ĠĠ ĠĠ", "...
- llama_model_loader: - kv 29: tokenizer.ggml.bos_token_id u32 = 199998
- llama_model_loader: - kv 30: tokenizer.ggml.eos_token_id u32 = 200002
- llama_model_loader: - kv 31: tokenizer.ggml.padding_token_id u32 = 199999
- llama_model_loader: - kv 32: tokenizer.chat_template str = {#-\n In addition to the normal input...
- llama_model_loader: - kv 33: general.quantization_version u32 = 2
- llama_model_loader: - kv 34: general.file_type u32 = 38
- llama_model_loader: - kv 35: split.no u16 = 0
- llama_model_loader: - kv 36: split.tensors.count i32 = 687
- llama_model_loader: - kv 37: split.count u16 = 3
- llama_model_loader: - type f32: 433 tensors
- llama_model_loader: - type q8_0: 146 tensors
- llama_model_loader: - type mxfp4: 108 tensors
- print_info: file format = GGUF V3 (latest)
- print_info: file type = MXFP4 MoE
- print_info: file size = 59.02 GiB (4.34 BPW)
- load: printing all EOG tokens:
- load: - 199999 ('<|endoftext|>')
- load: - 200002 ('<|return|>')
- load: - 200007 ('<|end|>')
- load: - 200012 ('<|call|>')
- load: special_eog_ids contains both '<|return|>' and '<|call|>' tokens, removing '<|end|>' token from EOG list
- load: special tokens cache size = 21
- load: token to piece cache size = 1.3332 MB
- print_info: arch = gpt-oss
- print_info: vocab_only = 0
- print_info: n_ctx_train = 131072
- print_info: n_embd = 2880
- print_info: n_layer = 36
- print_info: n_head = 64
- print_info: n_head_kv = 8
- print_info: n_rot = 64
- print_info: n_swa = 128
- print_info: is_swa_any = 1
- print_info: n_embd_head_k = 64
- print_info: n_embd_head_v = 64
- print_info: n_gqa = 8
- print_info: n_embd_k_gqa = 512
- print_info: n_embd_v_gqa = 512
- print_info: f_norm_eps = 0.0e+00
- print_info: f_norm_rms_eps = 1.0e-05
- print_info: f_clamp_kqv = 0.0e+00
- print_info: f_max_alibi_bias = 0.0e+00
- print_info: f_logit_scale = 0.0e+00
- print_info: f_attn_scale = 0.0e+00
- print_info: n_ff = 2880
- print_info: n_expert = 128
- print_info: n_expert_used = 4
- print_info: causal attn = 1
- print_info: pooling type = 0
- print_info: rope type = 2
- print_info: rope scaling = yarn
- print_info: freq_base_train = 150000.0
- print_info: freq_scale_train = 0.03125
- print_info: n_ctx_orig_yarn = 4096
- print_info: rope_finetuned = unknown
- print_info: model type = 120B
- print_info: model params = 116.83 B
- print_info: general.name = Gpt Oss 120b
- print_info: n_ff_exp = 2880
- print_info: vocab type = BPE
- print_info: n_vocab = 201088
- print_info: n_merges = 446189
- print_info: BOS token = 199998 '<|startoftext|>'
- print_info: EOS token = 200002 '<|return|>'
- print_info: EOT token = 200007 '<|end|>'
- print_info: PAD token = 199999 '<|endoftext|>'
- print_info: LF token = 198 'Ċ'
- print_info: EOG token = 199999 '<|endoftext|>'
- print_info: EOG token = 200002 '<|return|>'
- print_info: EOG token = 200012 '<|call|>'
- print_info: max token length = 256
- load_tensors: loading model tensors, this can take a while... (mmap = true)
- load_tensors: offloading 36 repeating layers to GPU
- load_tensors: offloading output layer to GPU
- load_tensors: offloaded 37/37 layers to GPU
- load_tensors: Metal_Mapped model buffer size = 30268.16 MiB
- load_tensors: Metal_Mapped model buffer size = 30170.30 MiB
- load_tensors: CPU_Mapped model buffer size = 586.82 MiB
- ....................................................................................................
- llama_init_from_model: model default pooling_type is [0], but [-1] was specified
- llama_context: constructing llama_context
- llama_context: n_seq_max = 32
- llama_context: n_ctx = 300000
- llama_context: n_ctx_per_seq = 9375
- llama_context: n_batch = 2048
- llama_context: n_ubatch = 2048
- llama_context: causal_attn = 1
- llama_context: flash_attn = enabled
- llama_context: kv_unified = false
- llama_context: freq_base = 150000.0
- llama_context: freq_scale = 0.03125
- llama_context: n_ctx_per_seq (9375) < n_ctx_train (131072) -- the full capacity of the model will not be utilized
- ggml_metal_init: allocating
- ggml_metal_init: picking default device: Apple M4 Max
- ggml_metal_init: use bfloat = true
- ggml_metal_init: use fusion = true
- ggml_metal_init: use concurrency = true
- ggml_metal_init: use graph optimize = true
- llama_context: CPU output buffer size = 24.55 MiB
- llama_kv_cache_iswa: creating non-SWA KV cache, size = 9472 cells
- llama_kv_cache: Metal KV buffer size = 10656.00 MiB
- llama_kv_cache: size = 10656.00 MiB ( 9472 cells, 18 layers, 32/32 seqs), K (f16): 5328.00 MiB, V (f16): 5328.00 MiB
- llama_kv_cache_iswa: creating SWA KV cache, size = 2304 cells
- llama_kv_cache: Metal KV buffer size = 2592.00 MiB
- llama_kv_cache: size = 2592.00 MiB ( 2304 cells, 18 layers, 32/32 seqs), K (f16): 1296.00 MiB, V (f16): 1296.00 MiB
- llama_context: Metal compute buffer size = 11211.14 MiB
- llama_context: CPU compute buffer size = 114.59 MiB
- llama_context: graph nodes = 2096
- llama_context: graph splits = 2
- main: n_kv_max = 303104, n_batch = 2048, n_ubatch = 2048, flash_attn = 1, is_pp_shared = 0, n_gpu_layers = 999, n_threads = 12, n_threads_batch = 12
- | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s |
- |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
- | 4096 | 32 | 1 | 4128 | 3.840 | 1066.55 | 0.414 | 77.31 | 4.254 | 970.30 |
- | 4096 | 32 | 2 | 8256 | 10.012 | 818.22 | 0.699 | 91.56 | 10.711 | 770.79 |
- | 4096 | 32 | 4 | 16512 | 19.070 | 859.15 | 1.223 | 104.64 | 20.293 | 813.67 |
- | 4096 | 32 | 8 | 33024 | 36.727 | 892.20 | 2.361 | 108.42 | 39.088 | 844.86 |
- | 4096 | 32 | 16 | 66048 | 73.762 | 888.48 | 4.075 | 125.65 | 77.837 | 848.54 |
- | 4096 | 32 | 32 | 132096 | 147.384 | 889.32 | 6.361 | 160.97 | 153.746 | 859.19 |
- | 8192 | 32 | 1 | 8224 | 10.347 | 791.72 | 0.460 | 69.49 | 10.808 | 760.95 |
- | 8192 | 32 | 2 | 16448 | 20.998 | 780.25 | 0.743 | 86.09 | 21.742 | 756.52 |
- | 8192 | 32 | 4 | 32896 | 43.610 | 751.39 | 1.393 | 91.90 | 45.003 | 730.98 |
- | 8192 | 32 | 8 | 65792 | 81.570 | 803.43 | 2.807 | 91.21 | 84.377 | 779.74 |
- | 8192 | 32 | 16 | 131584 | 159.933 | 819.54 | 4.209 | 121.63 | 164.143 | 801.64 |
- | 8192 | 32 | 32 | 263168 | 320.016 | 819.16 | 7.321 | 139.87 | 327.337 | 803.97 |
- llama_perf_context_print: load time = 3230.50 ms
- llama_perf_context_print: prompt eval time = 958569.92 ms / 778128 tokens ( 1.23 ms per token, 811.76 tokens per second)
- llama_perf_context_print: eval time = 873.87 ms / 64 runs ( 13.65 ms per token, 73.24 tokens per second)
- llama_perf_context_print: total time = 962608.89 ms / 778192 tokens
- llama_perf_context_print: graphs reused = 372
- ggml_metal_free: deallocating
- === Parallel Benchmark Results ===
- Model: gpt-oss-20b-mxfp4.gguf
- Date: Sun Oct 19 17:32:49 EDT 2025
- macOS Version: 26.0.1
- Hardware: Mac16,6
- Chip: Apple M4 Max
- Total RAM: 128.00 GB
- CPU Cores: 16
- GPU Layers: 999
- Flash Attention: 1
- === Benchmark Parameters ===
- Context Size: 300000
- Ubatch Size: 2048
- Prompt Tokens: 4096,8192
- Generation Tokens: 32
- Parallel Requests: 1,2,4,8,16,32
- === Results ===
- ggml_metal_library_init: using embedded metal library
- ggml_metal_library_init: loaded in 0.006 sec
- ggml_metal_device_init: GPU name: Apple M4 Max
- ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
- ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
- ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
- ggml_metal_device_init: simdgroup reduction = true
- ggml_metal_device_init: simdgroup matrix mul. = true
- ggml_metal_device_init: has unified memory = true
- ggml_metal_device_init: has bfloat = true
- ggml_metal_device_init: use residency sets = true
- ggml_metal_device_init: use shared buffers = true
- ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
- build: 6789 (3d4e86bb) with Apple clang version 17.0.0 (clang-1700.3.19.1) for arm64-apple-darwin25.0.0
- llama_model_load_from_file_impl: using device Metal (Apple M4 Max) (unknown id) - 110100 MiB free
- llama_model_loader: loaded meta data with 35 key-value pairs and 459 tensors from /Usersusermodels/dgx/gpt-oss-20b-mxfp4.gguf (version GGUF V3 (latest))
- llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
- llama_model_loader: - kv 0: general.architecture str = gpt-oss
- llama_model_loader: - kv 1: general.type str = model
- llama_model_loader: - kv 2: general.name str = Gpt Oss 20b
- llama_model_loader: - kv 3: general.basename str = gpt-oss
- llama_model_loader: - kv 4: general.size_label str = 20B
- llama_model_loader: - kv 5: general.license str = apache-2.0
- llama_model_loader: - kv 6: general.tags arr[str,2] = ["vllm", "text-generation"]
- llama_model_loader: - kv 7: gpt-oss.block_count u32 = 24
- llama_model_loader: - kv 8: gpt-oss.context_length u32 = 131072
- llama_model_loader: - kv 9: gpt-oss.embedding_length u32 = 2880
- llama_model_loader: - kv 10: gpt-oss.feed_forward_length u32 = 2880
- llama_model_loader: - kv 11: gpt-oss.attention.head_count u32 = 64
- llama_model_loader: - kv 12: gpt-oss.attention.head_count_kv u32 = 8
- llama_model_loader: - kv 13: gpt-oss.rope.freq_base f32 = 150000.000000
- llama_model_loader: - kv 14: gpt-oss.attention.layer_norm_rms_epsilon f32 = 0.000010
- llama_model_loader: - kv 15: gpt-oss.expert_count u32 = 32
- llama_model_loader: - kv 16: gpt-oss.expert_used_count u32 = 4
- llama_model_loader: - kv 17: gpt-oss.attention.key_length u32 = 64
- llama_model_loader: - kv 18: gpt-oss.attention.value_length u32 = 64
- llama_model_loader: - kv 19: gpt-oss.attention.sliding_window u32 = 128
- llama_model_loader: - kv 20: gpt-oss.expert_feed_forward_length u32 = 2880
- llama_model_loader: - kv 21: gpt-oss.rope.scaling.type str = yarn
- llama_model_loader: - kv 22: gpt-oss.rope.scaling.factor f32 = 32.000000
- llama_model_loader: - kv 23: gpt-oss.rope.scaling.original_context_length u32 = 4096
- llama_model_loader: - kv 24: tokenizer.ggml.model str = gpt2
- llama_model_loader: - kv 25: tokenizer.ggml.pre str = gpt-4o
- llama_model_loader: - kv 26: tokenizer.ggml.tokens arr[str,201088] = ["!", "\"", "#", "$", "%", "&", "'", ...
- llama_model_loader: - kv 27: tokenizer.ggml.token_type arr[i32,201088] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
- llama_model_loader: - kv 28: tokenizer.ggml.merges arr[str,446189] = ["Ġ Ġ", "Ġ ĠĠĠ", "ĠĠ ĠĠ", "...
- llama_model_loader: - kv 29: tokenizer.ggml.bos_token_id u32 = 199998
- llama_model_loader: - kv 30: tokenizer.ggml.eos_token_id u32 = 200002
- llama_model_loader: - kv 31: tokenizer.ggml.padding_token_id u32 = 199999
- llama_model_loader: - kv 32: tokenizer.chat_template str = {#-\n In addition to the normal input...
- llama_model_loader: - kv 33: general.quantization_version u32 = 2
- llama_model_loader: - kv 34: general.file_type u32 = 38
- llama_model_loader: - type f32: 289 tensors
- llama_model_loader: - type q8_0: 98 tensors
- llama_model_loader: - type mxfp4: 72 tensors
- print_info: file format = GGUF V3 (latest)
- print_info: file type = MXFP4 MoE
- print_info: file size = 11.27 GiB (4.63 BPW)
- load: printing all EOG tokens:
- load: - 199999 ('<|endoftext|>')
- load: - 200002 ('<|return|>')
- load: - 200007 ('<|end|>')
- load: - 200012 ('<|call|>')
- load: special_eog_ids contains both '<|return|>' and '<|call|>' tokens, removing '<|end|>' token from EOG list
- load: special tokens cache size = 21
- load: token to piece cache size = 1.3332 MB
- print_info: arch = gpt-oss
- print_info: vocab_only = 0
- print_info: n_ctx_train = 131072
- print_info: n_embd = 2880
- print_info: n_layer = 24
- print_info: n_head = 64
- print_info: n_head_kv = 8
- print_info: n_rot = 64
- print_info: n_swa = 128
- print_info: is_swa_any = 1
- print_info: n_embd_head_k = 64
- print_info: n_embd_head_v = 64
- print_info: n_gqa = 8
- print_info: n_embd_k_gqa = 512
- print_info: n_embd_v_gqa = 512
- print_info: f_norm_eps = 0.0e+00
- print_info: f_norm_rms_eps = 1.0e-05
- print_info: f_clamp_kqv = 0.0e+00
- print_info: f_max_alibi_bias = 0.0e+00
- print_info: f_logit_scale = 0.0e+00
- print_info: f_attn_scale = 0.0e+00
- print_info: n_ff = 2880
- print_info: n_expert = 32
- print_info: n_expert_used = 4
- print_info: causal attn = 1
- print_info: pooling type = 0
- print_info: rope type = 2
- print_info: rope scaling = yarn
- print_info: freq_base_train = 150000.0
- print_info: freq_scale_train = 0.03125
- print_info: n_ctx_orig_yarn = 4096
- print_info: rope_finetuned = unknown
- print_info: model type = 20B
- print_info: model params = 20.91 B
- print_info: general.name = Gpt Oss 20b
- print_info: n_ff_exp = 2880
- print_info: vocab type = BPE
- print_info: n_vocab = 201088
- print_info: n_merges = 446189
- print_info: BOS token = 199998 '<|startoftext|>'
- print_info: EOS token = 200002 '<|return|>'
- print_info: EOT token = 200007 '<|end|>'
- print_info: PAD token = 199999 '<|endoftext|>'
- print_info: LF token = 198 'Ċ'
- print_info: EOG token = 199999 '<|endoftext|>'
- print_info: EOG token = 200002 '<|return|>'
- print_info: EOG token = 200012 '<|call|>'
- print_info: max token length = 256
- load_tensors: loading model tensors, this can take a while... (mmap = true)
- load_tensors: offloading 24 repeating layers to GPU
- load_tensors: offloading output layer to GPU
- load_tensors: offloaded 25/25 layers to GPU
- load_tensors: Metal_Mapped model buffer size = 11536.18 MiB
- load_tensors: CPU_Mapped model buffer size = 586.82 MiB
- ................................................................................
- llama_init_from_model: model default pooling_type is [0], but [-1] was specified
- llama_context: constructing llama_context
- llama_context: n_seq_max = 32
- llama_context: n_ctx = 300000
- llama_context: n_ctx_per_seq = 9375
- llama_context: n_batch = 2048
- llama_context: n_ubatch = 2048
- llama_context: causal_attn = 1
- llama_context: flash_attn = enabled
- llama_context: kv_unified = false
- llama_context: freq_base = 150000.0
- llama_context: freq_scale = 0.03125
- llama_context: n_ctx_per_seq (9375) < n_ctx_train (131072) -- the full capacity of the model will not be utilized
- ggml_metal_init: allocating
- ggml_metal_init: picking default device: Apple M4 Max
- ggml_metal_init: use bfloat = true
- ggml_metal_init: use fusion = true
- ggml_metal_init: use concurrency = true
- ggml_metal_init: use graph optimize = true
- llama_context: CPU output buffer size = 24.55 MiB
- llama_kv_cache_iswa: creating non-SWA KV cache, size = 9472 cells
- llama_kv_cache: Metal KV buffer size = 7104.00 MiB
- llama_kv_cache: size = 7104.00 MiB ( 9472 cells, 12 layers, 32/32 seqs), K (f16): 3552.00 MiB, V (f16): 3552.00 MiB
- llama_kv_cache_iswa: creating SWA KV cache, size = 2304 cells
- llama_kv_cache: Metal KV buffer size = 1728.00 MiB
- llama_kv_cache: size = 1728.00 MiB ( 2304 cells, 12 layers, 32/32 seqs), K (f16): 864.00 MiB, V (f16): 864.00 MiB
- llama_context: Metal compute buffer size = 7885.59 MiB
- llama_context: CPU compute buffer size = 114.59 MiB
- llama_context: graph nodes = 1400
- llama_context: graph splits = 2
- main: n_kv_max = 303104, n_batch = 2048, n_ubatch = 2048, flash_attn = 1, is_pp_shared = 0, n_gpu_layers = 999, n_threads = 12, n_threads_batch = 12
- | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s |
- |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
- | 4096 | 32 | 1 | 4128 | 2.655 | 1542.84 | 0.310 | 103.25 | 2.965 | 1392.35 |
- | 4096 | 32 | 2 | 8256 | 6.764 | 1211.13 | 0.566 | 113.12 | 7.330 | 1126.38 |
- | 4096 | 32 | 4 | 16512 | 12.218 | 1341.02 | 0.919 | 139.35 | 13.136 | 1257.00 |
- | 4096 | 32 | 8 | 33024 | 24.281 | 1349.52 | 1.751 | 146.24 | 26.032 | 1268.60 |
- | 4096 | 32 | 16 | 66048 | 47.249 | 1387.04 | 2.898 | 176.67 | 50.147 | 1317.09 |
- | 4096 | 32 | 32 | 132096 | 90.107 | 1454.62 | 3.180 | 322.00 | 93.287 | 1416.01 |
- | 8192 | 32 | 1 | 8224 | 5.998 | 1365.71 | 0.304 | 105.40 | 6.302 | 1304.99 |
- | 8192 | 32 | 2 | 16448 | 11.846 | 1383.03 | 0.492 | 129.97 | 12.339 | 1333.02 |
- | 8192 | 32 | 4 | 32896 | 23.743 | 1380.08 | 0.877 | 146.01 | 24.620 | 1336.14 |
- | 8192 | 32 | 8 | 65792 | 48.095 | 1362.65 | 1.690 | 151.48 | 49.785 | 1321.53 |
- | 8192 | 32 | 16 | 131584 | 96.484 | 1358.49 | 2.780 | 184.17 | 99.264 | 1325.60 |
- | 8192 | 32 | 32 | 263168 | 192.008 | 1365.28 | 3.552 | 288.28 | 195.560 | 1345.72 |
- llama_perf_context_print: load time = 1143.70 ms
- llama_perf_context_print: prompt eval time = 580230.14 ms / 778128 tokens ( 0.75 ms per token, 1341.07 tokens per second)
- llama_perf_context_print: eval time = 613.09 ms / 64 runs ( 9.58 ms per token, 104.39 tokens per second)
- llama_perf_context_print: total time = 581948.28 ms / 778192 tokens
- llama_perf_context_print: graphs reused = 372
- ggml_metal_free: deallocating
- === Parallel Benchmark Results ===
- Model: qwen2.5-coder-7b-instruct-q8_0.gguf
- Date: Sun Oct 19 19:08:54 EDT 2025
- macOS Version: 26.0.1
- Hardware: Mac16,6
- Chip: Apple M4 Max
- Total RAM: 128.00 GB
- CPU Cores: 16
- GPU Layers: 999
- Flash Attention: 1
- === Benchmark Parameters ===
- Context Size: 300000
- Ubatch Size: 2048
- Prompt Tokens: 4096,8192
- Generation Tokens: 32
- Parallel Requests: 1,2,4,8,16,32
- === Results ===
- ggml_metal_library_init: using embedded metal library
- ggml_metal_library_init: loaded in 0.005 sec
- ggml_metal_device_init: GPU name: Apple M4 Max
- ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
- ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
- ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
- ggml_metal_device_init: simdgroup reduction = true
- ggml_metal_device_init: simdgroup matrix mul. = true
- ggml_metal_device_init: has unified memory = true
- ggml_metal_device_init: has bfloat = true
- ggml_metal_device_init: use residency sets = true
- ggml_metal_device_init: use shared buffers = true
- ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
- build: 6789 (3d4e86bb) with Apple clang version 17.0.0 (clang-1700.3.19.1) for arm64-apple-darwin25.0.0
- llama_model_load_from_file_impl: using device Metal (Apple M4 Max) (unknown id) - 110100 MiB free
- llama_model_loader: loaded meta data with 29 key-value pairs and 339 tensors from /Usersusermodels/dgx/qwen2.5-coder-7b-instruct-q8_0.gguf (version GGUF V3 (latest))
- llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
- llama_model_loader: - kv 0: general.architecture str = qwen2
- llama_model_loader: - kv 1: general.type str = model
- llama_model_loader: - kv 2: general.name str = Qwen2.5 Coder 7B Instruct GGUF
- llama_model_loader: - kv 3: general.finetune str = Instruct-GGUF
- llama_model_loader: - kv 4: general.basename str = Qwen2.5-Coder
- llama_model_loader: - kv 5: general.size_label str = 7B
- llama_model_loader: - kv 6: qwen2.block_count u32 = 28
- llama_model_loader: - kv 7: qwen2.context_length u32 = 131072
- llama_model_loader: - kv 8: qwen2.embedding_length u32 = 3584
- llama_model_loader: - kv 9: qwen2.feed_forward_length u32 = 18944
- llama_model_loader: - kv 10: qwen2.attention.head_count u32 = 28
- llama_model_loader: - kv 11: qwen2.attention.head_count_kv u32 = 4
- llama_model_loader: - kv 12: qwen2.rope.freq_base f32 = 1000000.000000
- llama_model_loader: - kv 13: qwen2.attention.layer_norm_rms_epsilon f32 = 0.000001
- llama_model_loader: - kv 14: general.file_type u32 = 7
- llama_model_loader: - kv 15: tokenizer.ggml.model str = gpt2
- llama_model_loader: - kv 16: tokenizer.ggml.pre str = qwen2
- llama_model_loader: - kv 17: tokenizer.ggml.tokens arr[str,152064] = ["!", "\"", "#", "$", "%", "&", "'", ...
- llama_model_loader: - kv 18: tokenizer.ggml.token_type arr[i32,152064] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
- llama_model_loader: - kv 19: tokenizer.ggml.merges arr[str,151387] = ["Ġ Ġ", "ĠĠ ĠĠ", "i n", "Ġ t",...
- llama_model_loader: - kv 20: tokenizer.ggml.eos_token_id u32 = 151645
- llama_model_loader: - kv 21: tokenizer.ggml.padding_token_id u32 = 151643
- llama_model_loader: - kv 22: tokenizer.ggml.bos_token_id u32 = 151643
- llama_model_loader: - kv 23: tokenizer.ggml.add_bos_token bool = false
- llama_model_loader: - kv 24: tokenizer.chat_template str = {%- if tools %}\n {{- '<|im_start|>...
- llama_model_loader: - kv 25: general.quantization_version u32 = 2
- llama_model_loader: - kv 26: split.no u16 = 0
- llama_model_loader: - kv 27: split.count u16 = 0
- llama_model_loader: - kv 28: split.tensors.count i32 = 339
- llama_model_loader: - type f32: 141 tensors
- llama_model_loader: - type q8_0: 198 tensors
- print_info: file format = GGUF V3 (latest)
- print_info: file type = Q8_0
- print_info: file size = 7.54 GiB (8.50 BPW)
- load: printing all EOG tokens:
- load: - 151643 ('<|endoftext|>')
- load: - 151645 ('<|im_end|>')
- load: - 151662 ('<|fim_pad|>')
- load: - 151663 ('<|repo_name|>')
- load: - 151664 ('<|file_sep|>')
- load: special tokens cache size = 22
- load: token to piece cache size = 0.9310 MB
- print_info: arch = qwen2
- print_info: vocab_only = 0
- print_info: n_ctx_train = 131072
- print_info: n_embd = 3584
- print_info: n_layer = 28
- print_info: n_head = 28
- print_info: n_head_kv = 4
- print_info: n_rot = 128
- print_info: n_swa = 0
- print_info: is_swa_any = 0
- print_info: n_embd_head_k = 128
- print_info: n_embd_head_v = 128
- print_info: n_gqa = 7
- print_info: n_embd_k_gqa = 512
- print_info: n_embd_v_gqa = 512
- print_info: f_norm_eps = 0.0e+00
- print_info: f_norm_rms_eps = 1.0e-06
- print_info: f_clamp_kqv = 0.0e+00
- print_info: f_max_alibi_bias = 0.0e+00
- print_info: f_logit_scale = 0.0e+00
- print_info: f_attn_scale = 0.0e+00
- print_info: n_ff = 18944
- print_info: n_expert = 0
- print_info: n_expert_used = 0
- print_info: causal attn = 1
- print_info: pooling type = -1
- print_info: rope type = 2
- print_info: rope scaling = linear
- print_info: freq_base_train = 1000000.0
- print_info: freq_scale_train = 1
- print_info: n_ctx_orig_yarn = 131072
- print_info: rope_finetuned = unknown
- print_info: model type = 7B
- print_info: model params = 7.62 B
- print_info: general.name = Qwen2.5 Coder 7B Instruct GGUF
- print_info: vocab type = BPE
- print_info: n_vocab = 152064
- print_info: n_merges = 151387
- print_info: BOS token = 151643 '<|endoftext|>'
- print_info: EOS token = 151645 '<|im_end|>'
- print_info: EOT token = 151645 '<|im_end|>'
- print_info: PAD token = 151643 '<|endoftext|>'
- print_info: LF token = 198 'Ċ'
- print_info: FIM PRE token = 151659 '<|fim_prefix|>'
- print_info: FIM SUF token = 151661 '<|fim_suffix|>'
- print_info: FIM MID token = 151660 '<|fim_middle|>'
- print_info: FIM PAD token = 151662 '<|fim_pad|>'
- print_info: FIM REP token = 151663 '<|repo_name|>'
- print_info: FIM SEP token = 151664 '<|file_sep|>'
- print_info: EOG token = 151643 '<|endoftext|>'
- print_info: EOG token = 151645 '<|im_end|>'
- print_info: EOG token = 151662 '<|fim_pad|>'
- print_info: EOG token = 151663 '<|repo_name|>'
- print_info: EOG token = 151664 '<|file_sep|>'
- print_info: max token length = 256
- load_tensors: loading model tensors, this can take a while... (mmap = true)
- load_tensors: offloading 28 repeating layers to GPU
- load_tensors: offloading output layer to GPU
- load_tensors: offloaded 29/29 layers to GPU
- load_tensors: Metal_Mapped model buffer size = 7165.44 MiB
- load_tensors: CPU_Mapped model buffer size = 552.23 MiB
- .......................................................................................
- llama_context: constructing llama_context
- llama_context: n_seq_max = 32
- llama_context: n_ctx = 300000
- llama_context: n_ctx_per_seq = 9375
- llama_context: n_batch = 2048
- llama_context: n_ubatch = 2048
- llama_context: causal_attn = 1
- llama_context: flash_attn = enabled
- llama_context: kv_unified = false
- llama_context: freq_base = 1000000.0
- llama_context: freq_scale = 1
- llama_context: n_ctx_per_seq (9375) < n_ctx_train (131072) -- the full capacity of the model will not be utilized
- ggml_metal_init: allocating
- ggml_metal_init: picking default device: Apple M4 Max
- ggml_metal_init: use bfloat = true
- ggml_metal_init: use fusion = true
- ggml_metal_init: use concurrency = true
- ggml_metal_init: use graph optimize = true
- llama_context: CPU output buffer size = 18.56 MiB
- llama_kv_cache: Metal KV buffer size = 16576.00 MiB
- llama_kv_cache: size = 16576.00 MiB ( 9472 cells, 28 layers, 32/32 seqs), K (f16): 8288.00 MiB, V (f16): 8288.00 MiB
- llama_context: Metal compute buffer size = 1216.00 MiB
- llama_context: CPU compute buffer size = 102.05 MiB
- llama_context: graph nodes = 1015
- llama_context: graph splits = 2
- main: n_kv_max = 303104, n_batch = 2048, n_ubatch = 2048, flash_attn = 1, is_pp_shared = 0, n_gpu_layers = 999, n_threads = 12, n_threads_batch = 12
- | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s |
- |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
- | 4096 | 32 | 1 | 4128 | 5.597 | 731.85 | 0.573 | 55.85 | 6.170 | 669.08 |
- | 4096 | 32 | 2 | 8256 | 11.508 | 711.88 | 0.641 | 99.90 | 12.148 | 679.61 |
- | 4096 | 32 | 4 | 16512 | 22.183 | 738.57 | 1.226 | 104.40 | 23.409 | 705.35 |
- | 4096 | 32 | 8 | 33024 | 47.485 | 690.07 | 2.780 | 92.08 | 50.265 | 656.99 |
- | 4096 | 32 | 16 | 66048 | 95.468 | 686.47 | 2.453 | 208.69 | 97.921 | 674.50 |
- | 4096 | 32 | 32 | 132096 | 177.133 | 739.96 | 2.953 | 346.74 | 180.086 | 733.52 |
- | 8192 | 32 | 1 | 8224 | 11.835 | 692.19 | 0.594 | 53.90 | 12.429 | 661.70 |
- | 8192 | 32 | 2 | 16448 | 23.895 | 685.68 | 0.702 | 91.13 | 24.597 | 668.70 |
- | 8192 | 32 | 4 | 32896 | 47.030 | 696.75 | 1.344 | 95.26 | 48.374 | 680.04 |
- | 8192 | 32 | 8 | 65792 | 95.530 | 686.03 | 2.633 | 97.23 | 98.163 | 670.23 |
- | 8192 | 32 | 16 | 131584 | 191.551 | 684.27 | 2.880 | 177.76 | 194.431 | 676.76 |
- | 8192 | 32 | 32 | 263168 | 382.786 | 684.83 | 4.050 | 252.83 | 386.837 | 680.31 |
- llama_perf_context_print: load time = 1139.37 ms
- llama_perf_context_print: prompt eval time = 1133723.39 ms / 778128 tokens ( 1.46 ms per token, 686.35 tokens per second)
- llama_perf_context_print: eval time = 1166.11 ms / 64 runs ( 18.22 ms per token, 54.88 tokens per second)
- llama_perf_context_print: total time = 1135998.75 ms / 778192 tokens
- llama_perf_context_print: graphs reused = 372
- ggml_metal_free: deallocating
- === Parallel Benchmark Results ===
- Model: qwen3-30b-a3b-q8_0.gguf
- Date: Sun Oct 19 18:35:23 EDT 2025
- macOS Version: 26.0.1
- Hardware: Mac16,6
- Chip: Apple M4 Max
- Total RAM: 128.00 GB
- CPU Cores: 16
- GPU Layers: 999
- Flash Attention: 1
- === Benchmark Parameters ===
- Context Size: 300000
- Ubatch Size: 2048
- Prompt Tokens: 4096,8192
- Generation Tokens: 32
- Parallel Requests: 1,2,4,8,16,32
- === Results ===
- ggml_metal_library_init: using embedded metal library
- ggml_metal_library_init: loaded in 0.006 sec
- ggml_metal_device_init: GPU name: Apple M4 Max
- ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
- ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
- ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
- ggml_metal_device_init: simdgroup reduction = true
- ggml_metal_device_init: simdgroup matrix mul. = true
- ggml_metal_device_init: has unified memory = true
- ggml_metal_device_init: has bfloat = true
- ggml_metal_device_init: use residency sets = true
- ggml_metal_device_init: use shared buffers = true
- ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
- build: 6789 (3d4e86bb) with Apple clang version 17.0.0 (clang-1700.3.19.1) for arm64-apple-darwin25.0.0
- llama_model_load_from_file_impl: using device Metal (Apple M4 Max) (unknown id) - 110100 MiB free
- llama_model_loader: loaded meta data with 31 key-value pairs and 579 tensors from /Usersusermodels/dgx/qwen3-30b-a3b-q8_0.gguf (version GGUF V3 (latest))
- llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
- llama_model_loader: - kv 0: general.architecture str = qwen3moe
- llama_model_loader: - kv 1: general.type str = model
- llama_model_loader: - kv 2: general.name str = Qwen3 30Ba3 Instruct
- llama_model_loader: - kv 3: general.finetune str = 30Ba3-Instruct
- llama_model_loader: - kv 4: general.basename str = Qwen3
- llama_model_loader: - kv 5: general.size_label str = 128x1.8B
- llama_model_loader: - kv 6: qwen3moe.block_count u32 = 48
- llama_model_loader: - kv 7: qwen3moe.context_length u32 = 40960
- llama_model_loader: - kv 8: qwen3moe.embedding_length u32 = 2048
- llama_model_loader: - kv 9: qwen3moe.feed_forward_length u32 = 6144
- llama_model_loader: - kv 10: qwen3moe.attention.head_count u32 = 32
- llama_model_loader: - kv 11: qwen3moe.attention.head_count_kv u32 = 4
- llama_model_loader: - kv 12: qwen3moe.rope.freq_base f32 = 1000000.000000
- llama_model_loader: - kv 13: qwen3moe.attention.layer_norm_rms_epsilon f32 = 0.000001
- llama_model_loader: - kv 14: qwen3moe.expert_used_count u32 = 8
- llama_model_loader: - kv 15: qwen3moe.attention.key_length u32 = 128
- llama_model_loader: - kv 16: qwen3moe.attention.value_length u32 = 128
- llama_model_loader: - kv 17: qwen3moe.expert_count u32 = 128
- llama_model_loader: - kv 18: qwen3moe.expert_feed_forward_length u32 = 768
- llama_model_loader: - kv 19: tokenizer.ggml.model str = gpt2
- llama_model_loader: - kv 20: tokenizer.ggml.pre str = qwen2
- llama_model_loader: - kv 21: tokenizer.ggml.tokens arr[str,151936] = ["!", "\"", "#", "$", "%", "&", "'", ...
- llama_model_loader: - kv 22: tokenizer.ggml.token_type arr[i32,151936] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
- llama_model_loader: - kv 23: tokenizer.ggml.merges arr[str,151387] = ["Ġ Ġ", "ĠĠ ĠĠ", "i n", "Ġ t",...
- llama_model_loader: - kv 24: tokenizer.ggml.eos_token_id u32 = 151645
- llama_model_loader: - kv 25: tokenizer.ggml.padding_token_id u32 = 151643
- llama_model_loader: - kv 26: tokenizer.ggml.bos_token_id u32 = 151643
- llama_model_loader: - kv 27: tokenizer.ggml.add_bos_token bool = false
- llama_model_loader: - kv 28: tokenizer.chat_template str = {%- if tools %}\n {{- '<|im_start|>...
- llama_model_loader: - kv 29: general.quantization_version u32 = 2
- llama_model_loader: - kv 30: general.file_type u32 = 7
- llama_model_loader: - type f32: 241 tensors
- llama_model_loader: - type q8_0: 338 tensors
- print_info: file format = GGUF V3 (latest)
- print_info: file type = Q8_0
- print_info: file size = 30.25 GiB (8.51 BPW)
- load: printing all EOG tokens:
- load: - 151643 ('<|endoftext|>')
- load: - 151645 ('<|im_end|>')
- load: - 151662 ('<|fim_pad|>')
- load: - 151663 ('<|repo_name|>')
- load: - 151664 ('<|file_sep|>')
- load: special tokens cache size = 26
- load: token to piece cache size = 0.9311 MB
- print_info: arch = qwen3moe
- print_info: vocab_only = 0
- print_info: n_ctx_train = 40960
- print_info: n_embd = 2048
- print_info: n_layer = 48
- print_info: n_head = 32
- print_info: n_head_kv = 4
- print_info: n_rot = 128
- print_info: n_swa = 0
- print_info: is_swa_any = 0
- print_info: n_embd_head_k = 128
- print_info: n_embd_head_v = 128
- print_info: n_gqa = 8
- print_info: n_embd_k_gqa = 512
- print_info: n_embd_v_gqa = 512
- print_info: f_norm_eps = 0.0e+00
- print_info: f_norm_rms_eps = 1.0e-06
- print_info: f_clamp_kqv = 0.0e+00
- print_info: f_max_alibi_bias = 0.0e+00
- print_info: f_logit_scale = 0.0e+00
- print_info: f_attn_scale = 0.0e+00
- print_info: n_ff = 6144
- print_info: n_expert = 128
- print_info: n_expert_used = 8
- print_info: causal attn = 1
- print_info: pooling type = 0
- print_info: rope type = 2
- print_info: rope scaling = linear
- print_info: freq_base_train = 1000000.0
- print_info: freq_scale_train = 1
- print_info: n_ctx_orig_yarn = 40960
- print_info: rope_finetuned = unknown
- print_info: model type = 30B.A3B
- print_info: model params = 30.53 B
- print_info: general.name = Qwen3 30Ba3 Instruct
- print_info: n_ff_exp = 768
- print_info: vocab type = BPE
- print_info: n_vocab = 151936
- print_info: n_merges = 151387
- print_info: BOS token = 151643 '<|endoftext|>'
- print_info: EOS token = 151645 '<|im_end|>'
- print_info: EOT token = 151645 '<|im_end|>'
- print_info: PAD token = 151643 '<|endoftext|>'
- print_info: LF token = 198 'Ċ'
- print_info: FIM PRE token = 151659 '<|fim_prefix|>'
- print_info: FIM SUF token = 151661 '<|fim_suffix|>'
- print_info: FIM MID token = 151660 '<|fim_middle|>'
- print_info: FIM PAD token = 151662 '<|fim_pad|>'
- print_info: FIM REP token = 151663 '<|repo_name|>'
- print_info: FIM SEP token = 151664 '<|file_sep|>'
- print_info: EOG token = 151643 '<|endoftext|>'
- print_info: EOG token = 151645 '<|im_end|>'
- print_info: EOG token = 151662 '<|fim_pad|>'
- print_info: EOG token = 151663 '<|repo_name|>'
- print_info: EOG token = 151664 '<|file_sep|>'
- print_info: max token length = 256
- load_tensors: loading model tensors, this can take a while... (mmap = true)
- load_tensors: offloading 48 repeating layers to GPU
- load_tensors: offloading output layer to GPU
- load_tensors: offloaded 49/49 layers to GPU
- load_tensors: Metal_Mapped model buffer size = 30973.40 MiB
- load_tensors: CPU_Mapped model buffer size = 315.30 MiB
- ...................................................................................................
- llama_init_from_model: model default pooling_type is [0], but [-1] was specified
- llama_context: constructing llama_context
- llama_context: n_seq_max = 32
- llama_context: n_ctx = 300000
- llama_context: n_ctx_per_seq = 9375
- llama_context: n_batch = 2048
- llama_context: n_ubatch = 2048
- llama_context: causal_attn = 1
- llama_context: flash_attn = enabled
- llama_context: kv_unified = false
- llama_context: freq_base = 1000000.0
- llama_context: freq_scale = 1
- llama_context: n_ctx_per_seq (9375) < n_ctx_train (40960) -- the full capacity of the model will not be utilized
- ggml_metal_init: allocating
- ggml_metal_init: picking default device: Apple M4 Max
- ggml_metal_init: use bfloat = true
- ggml_metal_init: use fusion = true
- ggml_metal_init: use concurrency = true
- ggml_metal_init: use graph optimize = true
- llama_context: CPU output buffer size = 18.55 MiB
- llama_kv_cache: Metal KV buffer size = 28416.00 MiB
- llama_kv_cache: size = 28416.00 MiB ( 9472 cells, 48 layers, 32/32 seqs), K (f16): 14208.00 MiB, V (f16): 14208.00 MiB
- llama_context: Metal compute buffer size = 7265.07 MiB
- llama_context: CPU compute buffer size = 90.05 MiB
- llama_context: graph nodes = 3079
- llama_context: graph splits = 2
- main: n_kv_max = 303104, n_batch = 2048, n_ubatch = 2048, flash_attn = 1, is_pp_shared = 0, n_gpu_layers = 999, n_threads = 12, n_threads_batch = 12
- | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s |
- |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
- | 4096 | 32 | 1 | 4128 | 2.635 | 1554.20 | 0.433 | 73.87 | 3.069 | 1345.22 |
- | 4096 | 32 | 2 | 8256 | 6.712 | 1220.48 | 0.645 | 99.18 | 7.357 | 1122.13 |
- | 4096 | 32 | 4 | 16512 | 13.668 | 1198.74 | 1.109 | 115.43 | 14.777 | 1117.44 |
- | 4096 | 32 | 8 | 33024 | 26.887 | 1218.74 | 2.115 | 121.02 | 29.002 | 1138.67 |
- | 4096 | 32 | 16 | 66048 | 53.227 | 1231.25 | 4.077 | 125.58 | 57.304 | 1152.58 |
- | 4096 | 32 | 32 | 132096 | 114.297 | 1146.77 | 5.387 | 190.10 | 119.683 | 1103.71 |
- | 8192 | 32 | 1 | 8224 | 8.294 | 987.73 | 0.484 | 66.11 | 8.778 | 936.91 |
- | 8192 | 32 | 2 | 16448 | 16.465 | 995.09 | 0.751 | 85.24 | 17.216 | 955.41 |
- | 8192 | 32 | 4 | 32896 | 32.029 | 1023.07 | 1.373 | 93.22 | 33.402 | 984.85 |
- | 8192 | 32 | 8 | 65792 | 64.867 | 1010.32 | 2.433 | 105.22 | 67.300 | 977.60 |
- | 8192 | 32 | 16 | 131584 | 129.668 | 1010.83 | 4.935 | 103.76 | 134.603 | 977.57 |
- | 8192 | 32 | 32 | 263168 | 260.440 | 1006.54 | 6.647 | 154.05 | 267.087 | 985.33 |
- llama_perf_context_print: load time = 2699.68 ms
- llama_perf_context_print: prompt eval time = 758745.76 ms / 778128 tokens ( 0.98 ms per token, 1025.55 tokens per second)
- llama_perf_context_print: eval time = 916.61 ms / 64 runs ( 14.32 ms per token, 69.82 tokens per second)
- llama_perf_context_print: total time = 762309.19 ms / 778192 tokens
- llama_perf_context_print: graphs reused = 372
- ggml_metal_free: deallocating
- === Sequential Benchmark Results ===
- Model: gemma-3-4b-it-q4_0.gguf
- Date: Sun Oct 19 19:27:50 EDT 2025
- macOS Version: 26.0.1
- Hardware: Mac16,6
- Chip: Apple M4 Max
- Total RAM: 128.00 GB
- CPU Cores: 16
- GPU Layers: 999
- Flash Attention: 1
- === Benchmark Parameters ===
- Prompt Sizes: 2048
- Generation Sizes: 32
- Batch Size:
- === Results ===
- ggml_metal_library_init: using embedded metal library
- ggml_metal_library_init: loaded in 0.007 sec
- ggml_metal_device_init: GPU name: Apple M4 Max
- ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
- ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
- ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
- ggml_metal_device_init: simdgroup reduction = true
- ggml_metal_device_init: simdgroup matrix mul. = true
- ggml_metal_device_init: has unified memory = true
- ggml_metal_device_init: has bfloat = true
- ggml_metal_device_init: use residency sets = true
- ggml_metal_device_init: use shared buffers = true
- ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
- | model | size | params | backend | threads | n_ubatch | fa | test | t/s |
- | ------------------------------ | ---------: | ---------: | ---------- | ------: | -------: | -: | --------------: | -------------------: |
- | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 | 1639.55 ± 49.17 |
- | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | tg32 | 132.69 ± 1.35 |
- | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d4096 | 1536.50 ± 94.10 |
- | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d4096 | 121.45 ± 1.36 |
- | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d8192 | 1498.66 ± 19.88 |
- | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d8192 | 118.23 ± 1.59 |
- | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d16384 | 1385.92 ± 23.74 |
- | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d16384 | 111.51 ± 0.26 |
- | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d32768 | 1166.95 ± 16.65 |
- | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d32768 | 99.56 ± 0.54 |
- build: 3d4e86bb (6789)
- === Sequential Benchmark Results ===
- Model: glm-4.5-air-q4_k_m-00001-of-00002.gguf
- Date: Sun Oct 19 19:44:11 EDT 2025
- macOS Version: 26.0.1
- Hardware: Mac16,6
- Chip: Apple M4 Max
- Total RAM: 128.00 GB
- CPU Cores: 16
- GPU Layers: 999
- Flash Attention: 1
- === Benchmark Parameters ===
- Prompt Sizes: 2048
- Generation Sizes: 32
- Batch Size:
- === Results ===
- ggml_metal_library_init: using embedded metal library
- ggml_metal_library_init: loaded in 0.005 sec
- ggml_metal_device_init: GPU name: Apple M4 Max
- ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
- ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
- ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
- ggml_metal_device_init: simdgroup reduction = true
- ggml_metal_device_init: simdgroup matrix mul. = true
- ggml_metal_device_init: has unified memory = true
- ggml_metal_device_init: has bfloat = true
- ggml_metal_device_init: use residency sets = true
- ggml_metal_device_init: use shared buffers = true
- ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
- | model | size | params | backend | threads | n_ubatch | fa | test | t/s |
- | ------------------------------ | ---------: | ---------: | ---------- | ------: | -------: | -: | --------------: | -------------------: |
- | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 | 312.02 ± 12.39 |
- | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | tg32 | 32.06 ± 2.94 |
- | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d4096 | 235.64 ± 2.57 |
- | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d4096 | 29.56 ± 0.88 |
- | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d8192 | 184.17 ± 6.39 |
- | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d8192 | 28.33 ± 0.21 |
- | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d16384 | 139.00 ± 1.75 |
- | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d16384 | 23.06 ± 0.40 |
- | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d32768 | 85.37 ± 0.75 |
- | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d32768 | 16.23 ± 0.73 |
- build: 3d4e86bb (6789)
- === Sequential Benchmark Results ===
- Model: gpt-oss-120b-mxfp4-00001-of-00003.gguf
- Date: Sun Oct 19 17:42:31 EDT 2025
- macOS Version: 26.0.1
- Hardware: Mac16,6
- Chip: Apple M4 Max
- Total RAM: 128.00 GB
- CPU Cores: 16
- GPU Layers: 999
- Flash Attention: 1
- === Benchmark Parameters ===
- Prompt Sizes: 2048
- Generation Sizes: 32
- Batch Size:
- === Results ===
- ggml_metal_library_init: using embedded metal library
- ggml_metal_library_init: loaded in 0.006 sec
- ggml_metal_device_init: GPU name: Apple M4 Max
- ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
- ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
- ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
- ggml_metal_device_init: simdgroup reduction = true
- ggml_metal_device_init: simdgroup matrix mul. = true
- ggml_metal_device_init: has unified memory = true
- ggml_metal_device_init: has bfloat = true
- ggml_metal_device_init: use residency sets = true
- ggml_metal_device_init: use shared buffers = true
- ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
- | model | size | params | backend | threads | n_ubatch | fa | test | t/s |
- | ------------------------------ | ---------: | ---------: | ---------- | ------: | -------: | -: | --------------: | -------------------: |
- | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 | 949.11 ± 87.73 |
- | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | tg32 | 81.68 ± 1.36 |
- | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d4096 | 812.53 ± 13.49 |
- | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d4096 | 74.77 ± 0.96 |
- | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d8192 | 716.74 ± 4.71 |
- | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d8192 | 73.03 ± 0.54 |
- | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d16384 | 580.13 ± 6.30 |
- | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d16384 | 66.01 ± 2.89 |
- | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d32768 | 411.04 ± 3.79 |
- | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d32768 | 57.12 ± 0.99 |
- build: 3d4e86bb (6789)
- === Sequential Benchmark Results ===
- Model: gpt-oss-20b-mxfp4.gguf
- Date: Sun Oct 19 17:22:09 EDT 2025
- macOS Version: 26.0.1
- Hardware: Mac16,6
- Chip: Apple M4 Max
- Total RAM: 128.00 GB
- CPU Cores: 16
- GPU Layers: 999
- Flash Attention: 1
- === Benchmark Parameters ===
- Prompt Sizes: 2048
- Generation Sizes: 32
- Batch Size:
- === Results ===
- ggml_metal_library_init: using embedded metal library
- ggml_metal_library_init: loaded in 0.004 sec
- ggml_metal_device_init: GPU name: Apple M4 Max
- ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
- ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
- ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
- ggml_metal_device_init: simdgroup reduction = true
- ggml_metal_device_init: simdgroup matrix mul. = true
- ggml_metal_device_init: has unified memory = true
- ggml_metal_device_init: has bfloat = true
- ggml_metal_device_init: use residency sets = true
- ggml_metal_device_init: use shared buffers = true
- ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
- | model | size | params | backend | threads | n_ubatch | fa | test | t/s |
- | ------------------------------ | ---------: | ---------: | ---------- | ------: | -------: | -: | --------------: | -------------------: |
- | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 | 1832.03 ± 22.74 |
- | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | tg32 | 122.27 ± 3.11 |
- | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d4096 | 1370.87 ± 44.31 |
- | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d4096 | 106.12 ± 1.90 |
- | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d8192 | 1174.47 ± 19.20 |
- | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d8192 | 100.58 ± 3.19 |
- | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d16384 | 918.14 ± 13.86 |
- | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d16384 | 92.41 ± 1.14 |
- | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d32768 | 634.84 ± 8.48 |
- | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d32768 | 76.55 ± 4.77 |
- build: 3d4e86bb (6789)
- === Sequential Benchmark Results ===
- Model: qwen2.5-coder-7b-instruct-q8_0.gguf
- Date: Sun Oct 19 18:48:06 EDT 2025
- macOS Version: 26.0.1
- Hardware: Mac16,6
- Chip: Apple M4 Max
- Total RAM: 128.00 GB
- CPU Cores: 16
- GPU Layers: 999
- Flash Attention: 1
- === Benchmark Parameters ===
- Prompt Sizes: 2048
- Generation Sizes: 32
- Batch Size:
- === Results ===
- ggml_metal_library_init: using embedded metal library
- ggml_metal_library_init: loaded in 0.006 sec
- ggml_metal_device_init: GPU name: Apple M4 Max
- ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
- ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
- ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
- ggml_metal_device_init: simdgroup reduction = true
- ggml_metal_device_init: simdgroup matrix mul. = true
- ggml_metal_device_init: has unified memory = true
- ggml_metal_device_init: has bfloat = true
- ggml_metal_device_init: use residency sets = true
- ggml_metal_device_init: use shared buffers = true
- ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
- | model | size | params | backend | threads | n_ubatch | fa | test | t/s |
- | ------------------------------ | ---------: | ---------: | ---------- | ------: | -------: | -: | --------------: | -------------------: |
- | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 | 752.70 ± 62.73 |
- | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | tg32 | 50.23 ± 7.60 |
- | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d4096 | 671.11 ± 14.22 |
- | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d4096 | 56.64 ± 0.28 |
- | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d8192 | 581.00 ± 20.79 |
- | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d8192 | 53.73 ± 0.79 |
- | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d16384 | 477.27 ± 12.32 |
- | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d16384 | 49.62 ± 0.23 |
- | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d32768 | 340.93 ± 1.76 |
- | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d32768 | 42.93 ± 0.69 |
- build: 3d4e86bb (6789)
- === Sequential Benchmark Results ===
- Model: qwen3-30b-a3b-q8_0.gguf
- Date: Sun Oct 19 18:15:57 EDT 2025
- macOS Version: 26.0.1
- Hardware: Mac16,6
- Chip: Apple M4 Max
- Total RAM: 128.00 GB
- CPU Cores: 16
- GPU Layers: 999
- Flash Attention: 1
- === Benchmark Parameters ===
- Prompt Sizes: 2048
- Generation Sizes: 32
- Batch Size:
- === Results ===
- ggml_metal_library_init: using embedded metal library
- ggml_metal_library_init: loaded in 0.006 sec
- ggml_metal_device_init: GPU name: Apple M4 Max
- ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
- ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
- ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
- ggml_metal_device_init: simdgroup reduction = true
- ggml_metal_device_init: simdgroup matrix mul. = true
- ggml_metal_device_init: has unified memory = true
- ggml_metal_device_init: has bfloat = true
- ggml_metal_device_init: use residency sets = true
- ggml_metal_device_init: use shared buffers = true
- ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
- | model | size | params | backend | threads | n_ubatch | fa | test | t/s |
- | ------------------------------ | ---------: | ---------: | ---------- | ------: | -------: | -: | --------------: | -------------------: |
- | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 | 1745.62 ± 29.06 |
- | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | tg32 | 84.96 ± 1.19 |
- | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d4096 | 918.50 ± 73.61 |
- | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d4096 | 70.22 ± 1.13 |
- | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d8192 | 698.42 ± 21.62 |
- | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d8192 | 64.64 ± 0.63 |
- | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d16384 | 449.50 ± 16.10 |
- | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d16384 | 53.08 ± 1.34 |
- | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d32768 | 253.05 ± 9.51 |
- | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d32768 | 40.53 ± 2.98 |
- build: 3d4e86bb (6789)
Advertisement
Add Comment
Please, Sign In to add comment