Guest User

Untitled

a guest
Dec 22nd, 2025
482
0
Never
Not a member of Pastebin yet? Sign Up, it unlocks many cool features!
text 89.39 KB | None | 0 0
  1. === Parallel Benchmark Results ===
  2. Model: gemma-3-4b-it-q4_0.gguf
  3. Date: Sun Oct 19 19:35:41 EDT 2025
  4. macOS Version: 26.0.1
  5. Hardware: Mac16,6
  6. Chip: Apple M4 Max
  7. Total RAM: 128.00 GB
  8. CPU Cores: 16
  9. GPU Layers: 999
  10. Flash Attention: 1
  11.  
  12. === Benchmark Parameters ===
  13. Context Size: 300000
  14. Ubatch Size: 2048
  15. Prompt Tokens: 4096,8192
  16. Generation Tokens: 32
  17. Parallel Requests: 1,2,4,8,16,32
  18.  
  19. === Results ===
  20. ggml_metal_library_init: using embedded metal library
  21. ggml_metal_library_init: loaded in 0.005 sec
  22. ggml_metal_device_init: GPU name: Apple M4 Max
  23. ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
  24. ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
  25. ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
  26. ggml_metal_device_init: simdgroup reduction = true
  27. ggml_metal_device_init: simdgroup matrix mul. = true
  28. ggml_metal_device_init: has unified memory = true
  29. ggml_metal_device_init: has bfloat = true
  30. ggml_metal_device_init: use residency sets = true
  31. ggml_metal_device_init: use shared buffers = true
  32. ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
  33. build: 6789 (3d4e86bb) with Apple clang version 17.0.0 (clang-1700.3.19.1) for arm64-apple-darwin25.0.0
  34. llama_model_load_from_file_impl: using device Metal (Apple M4 Max) (unknown id) - 110100 MiB free
  35. llama_model_loader: loaded meta data with 40 key-value pairs and 444 tensors from /Usersusermodels/dgx/gemma-3-4b-it-q4_0.gguf (version GGUF V3 (latest))
  36. llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
  37. llama_model_loader: - kv 0: general.architecture str = gemma3
  38. llama_model_loader: - kv 1: general.type str = model
  39. llama_model_loader: - kv 2: general.name str = Gemma-3-4B-It
  40. llama_model_loader: - kv 3: general.finetune str = it
  41. llama_model_loader: - kv 4: general.basename str = Gemma-3-4B-It
  42. llama_model_loader: - kv 5: general.quantized_by str = Unsloth
  43. llama_model_loader: - kv 6: general.size_label str = 4B
  44. llama_model_loader: - kv 7: general.repo_url str = https://huggingface.co/unsloth
  45. llama_model_loader: - kv 8: gemma3.context_length u32 = 131072
  46. llama_model_loader: - kv 9: gemma3.embedding_length u32 = 2560
  47. llama_model_loader: - kv 10: gemma3.block_count u32 = 34
  48. llama_model_loader: - kv 11: gemma3.feed_forward_length u32 = 10240
  49. llama_model_loader: - kv 12: gemma3.attention.head_count u32 = 8
  50. llama_model_loader: - kv 13: gemma3.attention.layer_norm_rms_epsilon f32 = 0.000001
  51. llama_model_loader: - kv 14: gemma3.attention.key_length u32 = 256
  52. llama_model_loader: - kv 15: gemma3.attention.value_length u32 = 256
  53. llama_model_loader: - kv 16: gemma3.rope.freq_base f32 = 1000000.000000
  54. llama_model_loader: - kv 17: gemma3.attention.sliding_window u32 = 1024
  55. llama_model_loader: - kv 18: gemma3.attention.head_count_kv u32 = 4
  56. llama_model_loader: - kv 19: gemma3.rope.scaling.type str = linear
  57. llama_model_loader: - kv 20: gemma3.rope.scaling.factor f32 = 8.000000
  58. llama_model_loader: - kv 21: tokenizer.ggml.model str = llama
  59. llama_model_loader: - kv 22: tokenizer.ggml.pre str = default
  60. llama_model_loader: - kv 23: tokenizer.ggml.tokens arr[str,262208] = ["<pad>", "<eos>", "<bos>", "<unk>", ...
  61. llama_model_loader: - kv 24: tokenizer.ggml.scores arr[f32,262208] = [-1000.000000, -1000.000000, -1000.00...
  62. llama_model_loader: - kv 25: tokenizer.ggml.token_type arr[i32,262208] = [3, 3, 3, 3, 3, 4, 3, 3, 3, 3, 3, 3, ...
  63. llama_model_loader: - kv 26: tokenizer.ggml.bos_token_id u32 = 2
  64. llama_model_loader: - kv 27: tokenizer.ggml.eos_token_id u32 = 106
  65. llama_model_loader: - kv 28: tokenizer.ggml.unknown_token_id u32 = 3
  66. llama_model_loader: - kv 29: tokenizer.ggml.padding_token_id u32 = 0
  67. llama_model_loader: - kv 30: tokenizer.ggml.add_bos_token bool = true
  68. llama_model_loader: - kv 31: tokenizer.ggml.add_eos_token bool = false
  69. llama_model_loader: - kv 32: tokenizer.chat_template str = {{ bos_token }}\n{%- if messages[0]['r...
  70. llama_model_loader: - kv 33: tokenizer.ggml.add_space_prefix bool = false
  71. llama_model_loader: - kv 34: general.quantization_version u32 = 2
  72. llama_model_loader: - kv 35: general.file_type u32 = 2
  73. llama_model_loader: - kv 36: quantize.imatrix.file str = gemma-3-4b-it-GGUF/imatrix_unsloth.dat
  74. llama_model_loader: - kv 37: quantize.imatrix.dataset str = unsloth_calibration_gemma-3-4b-it.txt
  75. llama_model_loader: - kv 38: quantize.imatrix.entries_count i32 = 238
  76. llama_model_loader: - kv 39: quantize.imatrix.chunks_count i32 = 663
  77. llama_model_loader: - type f32: 205 tensors
  78. llama_model_loader: - type q4_0: 234 tensors
  79. llama_model_loader: - type q4_1: 4 tensors
  80. llama_model_loader: - type q6_K: 1 tensors
  81. print_info: file format = GGUF V3 (latest)
  82. print_info: file type = Q4_0
  83. print_info: file size = 2.20 GiB (4.87 BPW)
  84. load: printing all EOG tokens:
  85. load: - 106 ('<end_of_turn>')
  86. load: special tokens cache size = 6415
  87. load: token to piece cache size = 1.9446 MB
  88. print_info: arch = gemma3
  89. print_info: vocab_only = 0
  90. print_info: n_ctx_train = 131072
  91. print_info: n_embd = 2560
  92. print_info: n_layer = 34
  93. print_info: n_head = 8
  94. print_info: n_head_kv = 4
  95. print_info: n_rot = 256
  96. print_info: n_swa = 1024
  97. print_info: is_swa_any = 1
  98. print_info: n_embd_head_k = 256
  99. print_info: n_embd_head_v = 256
  100. print_info: n_gqa = 2
  101. print_info: n_embd_k_gqa = 1024
  102. print_info: n_embd_v_gqa = 1024
  103. print_info: f_norm_eps = 0.0e+00
  104. print_info: f_norm_rms_eps = 1.0e-06
  105. print_info: f_clamp_kqv = 0.0e+00
  106. print_info: f_max_alibi_bias = 0.0e+00
  107. print_info: f_logit_scale = 0.0e+00
  108. print_info: f_attn_scale = 6.2e-02
  109. print_info: n_ff = 10240
  110. print_info: n_expert = 0
  111. print_info: n_expert_used = 0
  112. print_info: causal attn = 1
  113. print_info: pooling type = 0
  114. print_info: rope type = 2
  115. print_info: rope scaling = linear
  116. print_info: freq_base_train = 1000000.0
  117. print_info: freq_scale_train = 0.125
  118. print_info: n_ctx_orig_yarn = 131072
  119. print_info: rope_finetuned = unknown
  120. print_info: model type = 4B
  121. print_info: model params = 3.88 B
  122. print_info: general.name = Gemma-3-4B-It
  123. print_info: vocab type = SPM
  124. print_info: n_vocab = 262208
  125. print_info: n_merges = 0
  126. print_info: BOS token = 2 '<bos>'
  127. print_info: EOS token = 106 '<end_of_turn>'
  128. print_info: EOT token = 106 '<end_of_turn>'
  129. print_info: UNK token = 3 '<unk>'
  130. print_info: PAD token = 0 '<pad>'
  131. print_info: LF token = 248 '<0x0A>'
  132. print_info: EOG token = 106 '<end_of_turn>'
  133. print_info: max token length = 48
  134. load_tensors: loading model tensors, this can take a while... (mmap = true)
  135. load_tensors: offloading 34 repeating layers to GPU
  136. load_tensors: offloading output layer to GPU
  137. load_tensors: offloaded 35/35 layers to GPU
  138. load_tensors: Metal_Mapped model buffer size = 2254.03 MiB
  139. load_tensors: CPU_Mapped model buffer size = 525.13 MiB
  140. .................................................................
  141. llama_init_from_model: model default pooling_type is [0], but [-1] was specified
  142. llama_context: constructing llama_context
  143. llama_context: n_seq_max = 32
  144. llama_context: n_ctx = 300000
  145. llama_context: n_ctx_per_seq = 9375
  146. llama_context: n_batch = 2048
  147. llama_context: n_ubatch = 2048
  148. llama_context: causal_attn = 1
  149. llama_context: flash_attn = enabled
  150. llama_context: kv_unified = false
  151. llama_context: freq_base = 1000000.0
  152. llama_context: freq_scale = 0.125
  153. llama_context: n_ctx_per_seq (9375) < n_ctx_train (131072) -- the full capacity of the model will not be utilized
  154. ggml_metal_init: allocating
  155. ggml_metal_init: picking default device: Apple M4 Max
  156. ggml_metal_init: use bfloat = true
  157. ggml_metal_init: use fusion = true
  158. ggml_metal_init: use concurrency = true
  159. ggml_metal_init: use graph optimize = true
  160. llama_context: CPU output buffer size = 32.01 MiB
  161. llama_kv_cache_iswa: creating non-SWA KV cache, size = 9472 cells
  162. llama_kv_cache: Metal KV buffer size = 5920.00 MiB
  163. llama_kv_cache: size = 5920.00 MiB ( 9472 cells, 5 layers, 32/32 seqs), K (f16): 2960.00 MiB, V (f16): 2960.00 MiB
  164. llama_kv_cache_iswa: creating SWA KV cache, size = 3072 cells
  165. llama_kv_cache: Metal KV buffer size = 11136.00 MiB
  166. llama_kv_cache: size = 11136.00 MiB ( 3072 cells, 29 layers, 32/32 seqs), K (f16): 5568.00 MiB, V (f16): 5568.00 MiB
  167. llama_context: Metal compute buffer size = 2068.50 MiB
  168. llama_context: CPU compute buffer size = 118.09 MiB
  169. llama_context: graph nodes = 1437
  170. llama_context: graph splits = 2
  171.  
  172. main: n_kv_max = 303104, n_batch = 2048, n_ubatch = 2048, flash_attn = 1, is_pp_shared = 0, n_gpu_layers = 999, n_threads = 12, n_threads_batch = 12
  173.  
  174. | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s |
  175. |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
  176. | 4096 | 32 | 1 | 4128 | 2.468 | 1659.39 | 0.256 | 125.11 | 2.724 | 1515.34 |
  177. | 4096 | 32 | 2 | 8256 | 5.196 | 1576.59 | 0.417 | 153.41 | 5.613 | 1470.82 |
  178. | 4096 | 32 | 4 | 16512 | 10.522 | 1557.12 | 0.768 | 166.72 | 11.290 | 1462.56 |
  179. | 4096 | 32 | 8 | 33024 | 20.653 | 1586.57 | 1.445 | 177.19 | 22.098 | 1494.42 |
  180. | 4096 | 32 | 16 | 66048 | 40.388 | 1622.67 | 1.398 | 366.34 | 41.785 | 1580.65 |
  181. | 4096 | 32 | 32 | 132096 | 80.829 | 1621.61 | 1.767 | 579.39 | 82.596 | 1599.30 |
  182. | 8192 | 32 | 1 | 8224 | 5.408 | 1514.82 | 0.272 | 117.79 | 5.680 | 1448.00 |
  183. | 8192 | 32 | 2 | 16448 | 10.336 | 1585.11 | 0.443 | 144.52 | 10.779 | 1525.92 |
  184. | 8192 | 32 | 4 | 32896 | 20.636 | 1587.88 | 0.770 | 166.28 | 21.406 | 1536.76 |
  185. | 8192 | 32 | 8 | 65792 | 42.082 | 1557.33 | 1.505 | 170.05 | 43.588 | 1509.42 |
  186. | 8192 | 32 | 16 | 131584 | 83.043 | 1578.37 | 1.555 | 329.35 | 84.597 | 1555.41 |
  187. | 8192 | 32 | 32 | 263168 | 174.075 | 1505.92 | 2.071 | 494.53 | 176.146 | 1494.04 |
  188.  
  189. llama_perf_context_print: load time = 1067.47 ms
  190. llama_perf_context_print: prompt eval time = 507816.42 ms / 778128 tokens ( 0.65 ms per token, 1532.30 tokens per second)
  191. llama_perf_context_print: eval time = 527.06 ms / 64 runs ( 8.24 ms per token, 121.43 tokens per second)
  192. llama_perf_context_print: total time = 509409.84 ms / 778192 tokens
  193. llama_perf_context_print: graphs reused = 372
  194. ggml_metal_free: deallocating
  195.  
  196.  
  197. === Parallel Benchmark Results ===
  198. Model: glm-4.5-air-q4_k_m-00001-of-00002.gguf
  199. Date: Sun Oct 19 20:50:36 EDT 2025
  200. macOS Version: 26.0.1
  201. Hardware: Mac16,6
  202. Chip: Apple M4 Max
  203. Total RAM: 128.00 GB
  204. CPU Cores: 16
  205. GPU Layers: 999
  206. Flash Attention: 1
  207.  
  208. === Benchmark Parameters ===
  209. Context Size: 300000
  210. Ubatch Size: 2048
  211. Prompt Tokens: 4096,8192
  212. Generation Tokens: 32
  213. Parallel Requests: 1,2,4,8,16,32
  214.  
  215. === Results ===
  216. ggml_metal_library_init: using embedded metal library
  217. ggml_metal_library_init: loaded in 0.007 sec
  218. ggml_metal_device_init: GPU name: Apple M4 Max
  219. ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
  220. ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
  221. ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
  222. ggml_metal_device_init: simdgroup reduction = true
  223. ggml_metal_device_init: simdgroup matrix mul. = true
  224. ggml_metal_device_init: has unified memory = true
  225. ggml_metal_device_init: has bfloat = true
  226. ggml_metal_device_init: use residency sets = true
  227. ggml_metal_device_init: use shared buffers = true
  228. ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
  229. build: 6789 (3d4e86bb) with Apple clang version 17.0.0 (clang-1700.3.19.1) for arm64-apple-darwin25.0.0
  230. llama_model_load_from_file_impl: using device Metal (Apple M4 Max) (unknown id) - 110100 MiB free
  231. llama_model_loader: additional 1 GGUFs metadata loaded.
  232. llama_model_loader: loaded meta data with 48 key-value pairs and 803 tensors from /Usersusermodels/dgx/glm-4.5-air-q4_k_m-00001-of-00002.gguf (version GGUF V3 (latest))
  233. llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
  234. llama_model_loader: - kv 0: general.architecture str = glm4moe
  235. llama_model_loader: - kv 1: general.type str = model
  236. llama_model_loader: - kv 2: general.name str = GLM 4.5 Air
  237. llama_model_loader: - kv 3: general.size_label str = 128x9.4B
  238. llama_model_loader: - kv 4: general.license str = mit
  239. llama_model_loader: - kv 5: general.tags arr[str,1] = ["text-generation"]
  240. llama_model_loader: - kv 6: general.languages arr[str,2] = ["en", "zh"]
  241. llama_model_loader: - kv 7: glm4moe.block_count u32 = 47
  242. llama_model_loader: - kv 8: glm4moe.context_length u32 = 131072
  243. llama_model_loader: - kv 9: glm4moe.embedding_length u32 = 4096
  244. llama_model_loader: - kv 10: glm4moe.feed_forward_length u32 = 10944
  245. llama_model_loader: - kv 11: glm4moe.attention.head_count u32 = 96
  246. llama_model_loader: - kv 12: glm4moe.attention.head_count_kv u32 = 8
  247. llama_model_loader: - kv 13: glm4moe.rope.freq_base f32 = 1000000.000000
  248. llama_model_loader: - kv 14: glm4moe.attention.layer_norm_rms_epsilon f32 = 0.000010
  249. llama_model_loader: - kv 15: glm4moe.expert_used_count u32 = 8
  250. llama_model_loader: - kv 16: glm4moe.attention.key_length u32 = 128
  251. llama_model_loader: - kv 17: glm4moe.attention.value_length u32 = 128
  252. llama_model_loader: - kv 18: glm4moe.rope.dimension_count u32 = 64
  253. llama_model_loader: - kv 19: glm4moe.expert_count u32 = 128
  254. llama_model_loader: - kv 20: glm4moe.expert_feed_forward_length u32 = 1408
  255. llama_model_loader: - kv 21: glm4moe.expert_shared_count u32 = 1
  256. llama_model_loader: - kv 22: glm4moe.leading_dense_block_count u32 = 1
  257. llama_model_loader: - kv 23: glm4moe.expert_gating_func u32 = 2
  258. llama_model_loader: - kv 24: glm4moe.expert_weights_scale f32 = 1.000000
  259. llama_model_loader: - kv 25: glm4moe.expert_weights_norm bool = true
  260. llama_model_loader: - kv 26: glm4moe.nextn_predict_layers u32 = 1
  261. llama_model_loader: - kv 27: tokenizer.ggml.model str = gpt2
  262. llama_model_loader: - kv 28: tokenizer.ggml.pre str = glm4
  263. llama_model_loader: - kv 29: tokenizer.ggml.tokens arr[str,151552] = ["!", "\"", "#", "$", "%", "&", "'", ...
  264. llama_model_loader: - kv 30: tokenizer.ggml.token_type arr[i32,151552] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
  265. llama_model_loader: - kv 31: tokenizer.ggml.merges arr[str,318088] = ["Ġ Ġ", "Ġ ĠĠĠ", "ĠĠ ĠĠ", "...
  266. llama_model_loader: - kv 32: tokenizer.ggml.eos_token_id u32 = 151329
  267. llama_model_loader: - kv 33: tokenizer.ggml.padding_token_id u32 = 151329
  268. llama_model_loader: - kv 34: tokenizer.ggml.bos_token_id u32 = 151331
  269. llama_model_loader: - kv 35: tokenizer.ggml.eot_token_id u32 = 151336
  270. llama_model_loader: - kv 36: tokenizer.ggml.unknown_token_id u32 = 151329
  271. llama_model_loader: - kv 37: tokenizer.ggml.eom_token_id u32 = 151338
  272. llama_model_loader: - kv 38: tokenizer.chat_template str = [gMASK]<sop>\n{%- if tools -%}\n<|syste...
  273. llama_model_loader: - kv 39: general.quantization_version u32 = 2
  274. llama_model_loader: - kv 40: general.file_type u32 = 15
  275. llama_model_loader: - kv 41: quantize.imatrix.file str = /models_out/GLM-4.5-Air-GGUF/zai-org_...
  276. llama_model_loader: - kv 42: quantize.imatrix.dataset str = /training_dir/calibration_datav3.txt
  277. llama_model_loader: - kv 43: quantize.imatrix.entries_count u32 = 502
  278. llama_model_loader: - kv 44: quantize.imatrix.chunks_count u32 = 485
  279. llama_model_loader: - kv 45: split.no u16 = 0
  280. llama_model_loader: - kv 46: split.tensors.count i32 = 803
  281. llama_model_loader: - kv 47: split.count u16 = 2
  282. llama_model_loader: - type f32: 331 tensors
  283. llama_model_loader: - type q5_0: 24 tensors
  284. llama_model_loader: - type q8_0: 160 tensors
  285. llama_model_loader: - type q4_K: 169 tensors
  286. llama_model_loader: - type q5_K: 47 tensors
  287. llama_model_loader: - type q6_K: 72 tensors
  288. print_info: file format = GGUF V3 (latest)
  289. print_info: file type = Q4_K - Medium
  290. print_info: file size = 68.45 GiB (5.32 BPW)
  291. load: special_eot_id is not in special_eog_ids - the tokenizer config may be incorrect
  292. load: special_eom_id is not in special_eog_ids - the tokenizer config may be incorrect
  293. load: printing all EOG tokens:
  294. load: - 151329 ('<|endoftext|>')
  295. load: - 151336 ('<|user|>')
  296. load: - 151338 ('<|observation|>')
  297. load: special tokens cache size = 36
  298. load: token to piece cache size = 0.9713 MB
  299. print_info: arch = glm4moe
  300. print_info: vocab_only = 0
  301. print_info: n_ctx_train = 131072
  302. print_info: n_embd = 4096
  303. print_info: n_layer = 47
  304. print_info: n_head = 96
  305. print_info: n_head_kv = 8
  306. print_info: n_rot = 64
  307. print_info: n_swa = 0
  308. print_info: is_swa_any = 0
  309. print_info: n_embd_head_k = 128
  310. print_info: n_embd_head_v = 128
  311. print_info: n_gqa = 12
  312. print_info: n_embd_k_gqa = 1024
  313. print_info: n_embd_v_gqa = 1024
  314. print_info: f_norm_eps = 0.0e+00
  315. print_info: f_norm_rms_eps = 1.0e-05
  316. print_info: f_clamp_kqv = 0.0e+00
  317. print_info: f_max_alibi_bias = 0.0e+00
  318. print_info: f_logit_scale = 0.0e+00
  319. print_info: f_attn_scale = 0.0e+00
  320. print_info: n_ff = 10944
  321. print_info: n_expert = 128
  322. print_info: n_expert_used = 8
  323. print_info: causal attn = 1
  324. print_info: pooling type = 0
  325. print_info: rope type = 2
  326. print_info: rope scaling = linear
  327. print_info: freq_base_train = 1000000.0
  328. print_info: freq_scale_train = 1
  329. print_info: n_ctx_orig_yarn = 131072
  330. print_info: rope_finetuned = unknown
  331. print_info: model type = 106B.A12B
  332. print_info: model params = 110.47 B
  333. print_info: general.name = GLM 4.5 Air
  334. print_info: vocab type = BPE
  335. print_info: n_vocab = 151552
  336. print_info: n_merges = 318088
  337. print_info: BOS token = 151331 '[gMASK]'
  338. print_info: EOS token = 151329 '<|endoftext|>'
  339. print_info: EOT token = 151336 '<|user|>'
  340. print_info: EOM token = 151338 '<|observation|>'
  341. print_info: UNK token = 151329 '<|endoftext|>'
  342. print_info: PAD token = 151329 '<|endoftext|>'
  343. print_info: LF token = 198 'Ċ'
  344. print_info: FIM PRE token = 151347 '<|code_prefix|>'
  345. print_info: FIM SUF token = 151349 '<|code_suffix|>'
  346. print_info: FIM MID token = 151348 '<|code_middle|>'
  347. print_info: EOG token = 151329 '<|endoftext|>'
  348. print_info: EOG token = 151336 '<|user|>'
  349. print_info: EOG token = 151338 '<|observation|>'
  350. print_info: max token length = 1024
  351. load_tensors: loading model tensors, this can take a while... (mmap = true)
  352. model has unused tensor blk.46.attn_norm.weight (size = 16384 bytes) -- ignoring
  353. model has unused tensor blk.46.attn_q.weight (size = 28311552 bytes) -- ignoring
  354. model has unused tensor blk.46.attn_k.weight (size = 4456448 bytes) -- ignoring
  355. model has unused tensor blk.46.attn_v.weight (size = 3440640 bytes) -- ignoring
  356. model has unused tensor blk.46.attn_q.bias (size = 49152 bytes) -- ignoring
  357. model has unused tensor blk.46.attn_k.bias (size = 4096 bytes) -- ignoring
  358. model has unused tensor blk.46.attn_v.bias (size = 4096 bytes) -- ignoring
  359. model has unused tensor blk.46.attn_output.weight (size = 34603008 bytes) -- ignoring
  360. model has unused tensor blk.46.post_attention_norm.weight (size = 16384 bytes) -- ignoring
  361. model has unused tensor blk.46.ffn_gate_inp.weight (size = 2097152 bytes) -- ignoring
  362. model has unused tensor blk.46.exp_probs_b.bias (size = 512 bytes) -- ignoring
  363. model has unused tensor blk.46.ffn_gate_exps.weight (size = 415236096 bytes) -- ignoring
  364. model has unused tensor blk.46.ffn_down_exps.weight (size = 784334848 bytes) -- ignoring
  365. model has unused tensor blk.46.ffn_up_exps.weight (size = 415236096 bytes) -- ignoring
  366. model has unused tensor blk.46.ffn_gate_shexp.weight (size = 6127616 bytes) -- ignoring
  367. model has unused tensor blk.46.ffn_down_shexp.weight (size = 6127616 bytes) -- ignoring
  368. model has unused tensor blk.46.ffn_up_shexp.weight (size = 6127616 bytes) -- ignoring
  369. model has unused tensor blk.46.nextn.eh_proj.weight (size = 18874368 bytes) -- ignoring
  370. model has unused tensor blk.46.nextn.enorm.weight (size = 16384 bytes) -- ignoring
  371. model has unused tensor blk.46.nextn.hnorm.weight (size = 16384 bytes) -- ignoring
  372. model has unused tensor blk.46.nextn.embed_tokens.weight (size = 349175808 bytes) -- ignoring
  373. model has unused tensor blk.46.nextn.shared_head_head.weight (size = 349175808 bytes) -- ignoring
  374. model has unused tensor blk.46.nextn.shared_head_norm.weight (size = 16384 bytes) -- ignoring
  375. load_tensors: offloading 47 repeating layers to GPU
  376. load_tensors: offloading output layer to GPU
  377. load_tensors: offloaded 48/48 layers to GPU
  378. load_tensors: Metal_Mapped model buffer size = 37977.33 MiB
  379. load_tensors: Metal_Mapped model buffer size = 29799.45 MiB
  380. load_tensors: CPU_Mapped model buffer size = 333.00 MiB
  381. ..................................................................................................
  382. llama_init_from_model: model default pooling_type is [0], but [-1] was specified
  383. llama_context: constructing llama_context
  384. llama_context: n_seq_max = 32
  385. llama_context: n_ctx = 300000
  386. llama_context: n_ctx_per_seq = 9375
  387. llama_context: n_batch = 2048
  388. llama_context: n_ubatch = 2048
  389. llama_context: causal_attn = 1
  390. llama_context: flash_attn = enabled
  391. llama_context: kv_unified = false
  392. llama_context: freq_base = 1000000.0
  393. llama_context: freq_scale = 1
  394. llama_context: n_ctx_per_seq (9375) < n_ctx_train (131072) -- the full capacity of the model will not be utilized
  395. ggml_metal_init: allocating
  396. ggml_metal_init: picking default device: Apple M4 Max
  397. ggml_metal_init: use bfloat = true
  398. ggml_metal_init: use fusion = true
  399. ggml_metal_init: use concurrency = true
  400. ggml_metal_init: use graph optimize = true
  401. llama_context: CPU output buffer size = 18.50 MiB
  402. llama_kv_cache: Metal KV buffer size = 54464.00 MiB
  403. llama_kv_cache: size = 54464.00 MiB ( 9472 cells, 46 layers, 32/32 seqs), K (f16): 27232.00 MiB, V (f16): 27232.00 MiB
  404. llama_context: Metal compute buffer size = 12387.07 MiB
  405. llama_context: CPU compute buffer size = 106.05 MiB
  406. llama_context: graph nodes = 3193
  407. llama_context: graph splits = 2
  408.  
  409. main: n_kv_max = 303104, n_batch = 2048, n_ubatch = 2048, flash_attn = 1, is_pp_shared = 0, n_gpu_layers = 999, n_threads = 12, n_threads_batch = 12
  410.  
  411. | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s |
  412. |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
  413. | 4096 | 32 | 1 | 4128 | 40.042 | 102.29 | 437.052 | 0.07 | 477.094 | 8.65 |
  414. | 4096 | 32 | 2 | 8256 | 77.299 | 105.98 | 501.545 | 0.13 | 578.844 | 14.26 |
  415. | 4096 | 32 | 4 | 16512 | 161.838 | 101.24 | 499.020 | 0.26 | 660.858 | 24.99 |
  416. | 4096 | 32 | 8 | 33024 | 326.307 | 100.42 | 517.406 | 0.49 | 843.713 | 39.14 |
  417. | 4096 | 32 | 16 | 66048 | 658.933 | 99.46 | 538.631 | 0.95 | 1197.564 | 55.15 |
  418. | 4096 | 32 | 32 | 132096 | 1352.850 | 96.89 | 585.793 | 1.75 | 1938.643 | 68.14 |
  419. | 8192 | 32 | 1 | 8224 | 89.265 | 91.77 | 583.327 | 0.05 | 672.591 | 12.23 |
  420. | 8192 | 32 | 2 | 16448 | 178.881 | 91.59 | 583.590 | 0.11 | 762.471 | 21.57 |
  421. | 8192 | 32 | 4 | 32896 | 363.187 | 90.22 | 597.051 | 0.21 | 960.238 | 34.26 |
  422. | 8192 | 32 | 8 | 65792 | 740.917 | 88.45 | 586.614 | 0.44 | 1327.531 | 49.56 |
  423. | 8192 | 32 | 16 | 131584 | 1457.437 | 89.93 | 598.005 | 0.86 | 2055.442 | 64.02 |
  424. | 8192 | 32 | 32 | 263168 | 3294.007 | 79.58 | 1375.630 | 0.74 | 4669.637 | 56.36 |
  425.  
  426. llama_perf_context_print: load time = 16229.62 ms
  427. llama_perf_context_print: prompt eval time = 15132036.12 ms / 778128 tokens ( 19.45 ms per token, 51.42 tokens per second)
  428. llama_perf_context_print: eval time = 1020375.68 ms / 64 runs (15943.37 ms per token, 0.06 tokens per second)
  429. llama_perf_context_print: total time = 16160981.34 ms / 778192 tokens
  430. llama_perf_context_print: graphs reused = 372
  431. ggml_metal_free: deallocating
  432.  
  433.  
  434. === Parallel Benchmark Results ===
  435. Model: gpt-oss-120b-mxfp4-00001-of-00003.gguf
  436. Date: Sun Oct 19 17:59:54 EDT 2025
  437. macOS Version: 26.0.1
  438. Hardware: Mac16,6
  439. Chip: Apple M4 Max
  440. Total RAM: 128.00 GB
  441. CPU Cores: 16
  442. GPU Layers: 999
  443. Flash Attention: 1
  444.  
  445. === Benchmark Parameters ===
  446. Context Size: 300000
  447. Ubatch Size: 2048
  448. Prompt Tokens: 4096,8192
  449. Generation Tokens: 32
  450. Parallel Requests: 1,2,4,8,16,32
  451.  
  452. === Results ===
  453. ggml_metal_library_init: using embedded metal library
  454. ggml_metal_library_init: loaded in 0.006 sec
  455. ggml_metal_device_init: GPU name: Apple M4 Max
  456. ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
  457. ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
  458. ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
  459. ggml_metal_device_init: simdgroup reduction = true
  460. ggml_metal_device_init: simdgroup matrix mul. = true
  461. ggml_metal_device_init: has unified memory = true
  462. ggml_metal_device_init: has bfloat = true
  463. ggml_metal_device_init: use residency sets = true
  464. ggml_metal_device_init: use shared buffers = true
  465. ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
  466. build: 6789 (3d4e86bb) with Apple clang version 17.0.0 (clang-1700.3.19.1) for arm64-apple-darwin25.0.0
  467. llama_model_load_from_file_impl: using device Metal (Apple M4 Max) (unknown id) - 110100 MiB free
  468. llama_model_loader: additional 2 GGUFs metadata loaded.
  469. llama_model_loader: loaded meta data with 38 key-value pairs and 687 tensors from /Usersusermodels/dgx/gpt-oss-120b-mxfp4-00001-of-00003.gguf (version GGUF V3 (latest))
  470. llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
  471. llama_model_loader: - kv 0: general.architecture str = gpt-oss
  472. llama_model_loader: - kv 1: general.type str = model
  473. llama_model_loader: - kv 2: general.name str = Gpt Oss 120b
  474. llama_model_loader: - kv 3: general.basename str = gpt-oss
  475. llama_model_loader: - kv 4: general.size_label str = 120B
  476. llama_model_loader: - kv 5: general.license str = apache-2.0
  477. llama_model_loader: - kv 6: general.tags arr[str,2] = ["vllm", "text-generation"]
  478. llama_model_loader: - kv 7: gpt-oss.block_count u32 = 36
  479. llama_model_loader: - kv 8: gpt-oss.context_length u32 = 131072
  480. llama_model_loader: - kv 9: gpt-oss.embedding_length u32 = 2880
  481. llama_model_loader: - kv 10: gpt-oss.feed_forward_length u32 = 2880
  482. llama_model_loader: - kv 11: gpt-oss.attention.head_count u32 = 64
  483. llama_model_loader: - kv 12: gpt-oss.attention.head_count_kv u32 = 8
  484. llama_model_loader: - kv 13: gpt-oss.rope.freq_base f32 = 150000.000000
  485. llama_model_loader: - kv 14: gpt-oss.attention.layer_norm_rms_epsilon f32 = 0.000010
  486. llama_model_loader: - kv 15: gpt-oss.expert_count u32 = 128
  487. llama_model_loader: - kv 16: gpt-oss.expert_used_count u32 = 4
  488. llama_model_loader: - kv 17: gpt-oss.attention.key_length u32 = 64
  489. llama_model_loader: - kv 18: gpt-oss.attention.value_length u32 = 64
  490. llama_model_loader: - kv 19: gpt-oss.attention.sliding_window u32 = 128
  491. llama_model_loader: - kv 20: gpt-oss.expert_feed_forward_length u32 = 2880
  492. llama_model_loader: - kv 21: gpt-oss.rope.scaling.type str = yarn
  493. llama_model_loader: - kv 22: gpt-oss.rope.scaling.factor f32 = 32.000000
  494. llama_model_loader: - kv 23: gpt-oss.rope.scaling.original_context_length u32 = 4096
  495. llama_model_loader: - kv 24: tokenizer.ggml.model str = gpt2
  496. llama_model_loader: - kv 25: tokenizer.ggml.pre str = gpt-4o
  497. llama_model_loader: - kv 26: tokenizer.ggml.tokens arr[str,201088] = ["!", "\"", "#", "$", "%", "&", "'", ...
  498. llama_model_loader: - kv 27: tokenizer.ggml.token_type arr[i32,201088] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
  499. llama_model_loader: - kv 28: tokenizer.ggml.merges arr[str,446189] = ["Ġ Ġ", "Ġ ĠĠĠ", "ĠĠ ĠĠ", "...
  500. llama_model_loader: - kv 29: tokenizer.ggml.bos_token_id u32 = 199998
  501. llama_model_loader: - kv 30: tokenizer.ggml.eos_token_id u32 = 200002
  502. llama_model_loader: - kv 31: tokenizer.ggml.padding_token_id u32 = 199999
  503. llama_model_loader: - kv 32: tokenizer.chat_template str = {#-\n In addition to the normal input...
  504. llama_model_loader: - kv 33: general.quantization_version u32 = 2
  505. llama_model_loader: - kv 34: general.file_type u32 = 38
  506. llama_model_loader: - kv 35: split.no u16 = 0
  507. llama_model_loader: - kv 36: split.tensors.count i32 = 687
  508. llama_model_loader: - kv 37: split.count u16 = 3
  509. llama_model_loader: - type f32: 433 tensors
  510. llama_model_loader: - type q8_0: 146 tensors
  511. llama_model_loader: - type mxfp4: 108 tensors
  512. print_info: file format = GGUF V3 (latest)
  513. print_info: file type = MXFP4 MoE
  514. print_info: file size = 59.02 GiB (4.34 BPW)
  515. load: printing all EOG tokens:
  516. load: - 199999 ('<|endoftext|>')
  517. load: - 200002 ('<|return|>')
  518. load: - 200007 ('<|end|>')
  519. load: - 200012 ('<|call|>')
  520. load: special_eog_ids contains both '<|return|>' and '<|call|>' tokens, removing '<|end|>' token from EOG list
  521. load: special tokens cache size = 21
  522. load: token to piece cache size = 1.3332 MB
  523. print_info: arch = gpt-oss
  524. print_info: vocab_only = 0
  525. print_info: n_ctx_train = 131072
  526. print_info: n_embd = 2880
  527. print_info: n_layer = 36
  528. print_info: n_head = 64
  529. print_info: n_head_kv = 8
  530. print_info: n_rot = 64
  531. print_info: n_swa = 128
  532. print_info: is_swa_any = 1
  533. print_info: n_embd_head_k = 64
  534. print_info: n_embd_head_v = 64
  535. print_info: n_gqa = 8
  536. print_info: n_embd_k_gqa = 512
  537. print_info: n_embd_v_gqa = 512
  538. print_info: f_norm_eps = 0.0e+00
  539. print_info: f_norm_rms_eps = 1.0e-05
  540. print_info: f_clamp_kqv = 0.0e+00
  541. print_info: f_max_alibi_bias = 0.0e+00
  542. print_info: f_logit_scale = 0.0e+00
  543. print_info: f_attn_scale = 0.0e+00
  544. print_info: n_ff = 2880
  545. print_info: n_expert = 128
  546. print_info: n_expert_used = 4
  547. print_info: causal attn = 1
  548. print_info: pooling type = 0
  549. print_info: rope type = 2
  550. print_info: rope scaling = yarn
  551. print_info: freq_base_train = 150000.0
  552. print_info: freq_scale_train = 0.03125
  553. print_info: n_ctx_orig_yarn = 4096
  554. print_info: rope_finetuned = unknown
  555. print_info: model type = 120B
  556. print_info: model params = 116.83 B
  557. print_info: general.name = Gpt Oss 120b
  558. print_info: n_ff_exp = 2880
  559. print_info: vocab type = BPE
  560. print_info: n_vocab = 201088
  561. print_info: n_merges = 446189
  562. print_info: BOS token = 199998 '<|startoftext|>'
  563. print_info: EOS token = 200002 '<|return|>'
  564. print_info: EOT token = 200007 '<|end|>'
  565. print_info: PAD token = 199999 '<|endoftext|>'
  566. print_info: LF token = 198 'Ċ'
  567. print_info: EOG token = 199999 '<|endoftext|>'
  568. print_info: EOG token = 200002 '<|return|>'
  569. print_info: EOG token = 200012 '<|call|>'
  570. print_info: max token length = 256
  571. load_tensors: loading model tensors, this can take a while... (mmap = true)
  572. load_tensors: offloading 36 repeating layers to GPU
  573. load_tensors: offloading output layer to GPU
  574. load_tensors: offloaded 37/37 layers to GPU
  575. load_tensors: Metal_Mapped model buffer size = 30268.16 MiB
  576. load_tensors: Metal_Mapped model buffer size = 30170.30 MiB
  577. load_tensors: CPU_Mapped model buffer size = 586.82 MiB
  578. ....................................................................................................
  579. llama_init_from_model: model default pooling_type is [0], but [-1] was specified
  580. llama_context: constructing llama_context
  581. llama_context: n_seq_max = 32
  582. llama_context: n_ctx = 300000
  583. llama_context: n_ctx_per_seq = 9375
  584. llama_context: n_batch = 2048
  585. llama_context: n_ubatch = 2048
  586. llama_context: causal_attn = 1
  587. llama_context: flash_attn = enabled
  588. llama_context: kv_unified = false
  589. llama_context: freq_base = 150000.0
  590. llama_context: freq_scale = 0.03125
  591. llama_context: n_ctx_per_seq (9375) < n_ctx_train (131072) -- the full capacity of the model will not be utilized
  592. ggml_metal_init: allocating
  593. ggml_metal_init: picking default device: Apple M4 Max
  594. ggml_metal_init: use bfloat = true
  595. ggml_metal_init: use fusion = true
  596. ggml_metal_init: use concurrency = true
  597. ggml_metal_init: use graph optimize = true
  598. llama_context: CPU output buffer size = 24.55 MiB
  599. llama_kv_cache_iswa: creating non-SWA KV cache, size = 9472 cells
  600. llama_kv_cache: Metal KV buffer size = 10656.00 MiB
  601. llama_kv_cache: size = 10656.00 MiB ( 9472 cells, 18 layers, 32/32 seqs), K (f16): 5328.00 MiB, V (f16): 5328.00 MiB
  602. llama_kv_cache_iswa: creating SWA KV cache, size = 2304 cells
  603. llama_kv_cache: Metal KV buffer size = 2592.00 MiB
  604. llama_kv_cache: size = 2592.00 MiB ( 2304 cells, 18 layers, 32/32 seqs), K (f16): 1296.00 MiB, V (f16): 1296.00 MiB
  605. llama_context: Metal compute buffer size = 11211.14 MiB
  606. llama_context: CPU compute buffer size = 114.59 MiB
  607. llama_context: graph nodes = 2096
  608. llama_context: graph splits = 2
  609.  
  610. main: n_kv_max = 303104, n_batch = 2048, n_ubatch = 2048, flash_attn = 1, is_pp_shared = 0, n_gpu_layers = 999, n_threads = 12, n_threads_batch = 12
  611.  
  612. | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s |
  613. |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
  614. | 4096 | 32 | 1 | 4128 | 3.840 | 1066.55 | 0.414 | 77.31 | 4.254 | 970.30 |
  615. | 4096 | 32 | 2 | 8256 | 10.012 | 818.22 | 0.699 | 91.56 | 10.711 | 770.79 |
  616. | 4096 | 32 | 4 | 16512 | 19.070 | 859.15 | 1.223 | 104.64 | 20.293 | 813.67 |
  617. | 4096 | 32 | 8 | 33024 | 36.727 | 892.20 | 2.361 | 108.42 | 39.088 | 844.86 |
  618. | 4096 | 32 | 16 | 66048 | 73.762 | 888.48 | 4.075 | 125.65 | 77.837 | 848.54 |
  619. | 4096 | 32 | 32 | 132096 | 147.384 | 889.32 | 6.361 | 160.97 | 153.746 | 859.19 |
  620. | 8192 | 32 | 1 | 8224 | 10.347 | 791.72 | 0.460 | 69.49 | 10.808 | 760.95 |
  621. | 8192 | 32 | 2 | 16448 | 20.998 | 780.25 | 0.743 | 86.09 | 21.742 | 756.52 |
  622. | 8192 | 32 | 4 | 32896 | 43.610 | 751.39 | 1.393 | 91.90 | 45.003 | 730.98 |
  623. | 8192 | 32 | 8 | 65792 | 81.570 | 803.43 | 2.807 | 91.21 | 84.377 | 779.74 |
  624. | 8192 | 32 | 16 | 131584 | 159.933 | 819.54 | 4.209 | 121.63 | 164.143 | 801.64 |
  625. | 8192 | 32 | 32 | 263168 | 320.016 | 819.16 | 7.321 | 139.87 | 327.337 | 803.97 |
  626.  
  627. llama_perf_context_print: load time = 3230.50 ms
  628. llama_perf_context_print: prompt eval time = 958569.92 ms / 778128 tokens ( 1.23 ms per token, 811.76 tokens per second)
  629. llama_perf_context_print: eval time = 873.87 ms / 64 runs ( 13.65 ms per token, 73.24 tokens per second)
  630. llama_perf_context_print: total time = 962608.89 ms / 778192 tokens
  631. llama_perf_context_print: graphs reused = 372
  632. ggml_metal_free: deallocating
  633.  
  634.  
  635. === Parallel Benchmark Results ===
  636. Model: gpt-oss-20b-mxfp4.gguf
  637. Date: Sun Oct 19 17:32:49 EDT 2025
  638. macOS Version: 26.0.1
  639. Hardware: Mac16,6
  640. Chip: Apple M4 Max
  641. Total RAM: 128.00 GB
  642. CPU Cores: 16
  643. GPU Layers: 999
  644. Flash Attention: 1
  645.  
  646. === Benchmark Parameters ===
  647. Context Size: 300000
  648. Ubatch Size: 2048
  649. Prompt Tokens: 4096,8192
  650. Generation Tokens: 32
  651. Parallel Requests: 1,2,4,8,16,32
  652.  
  653. === Results ===
  654. ggml_metal_library_init: using embedded metal library
  655. ggml_metal_library_init: loaded in 0.006 sec
  656. ggml_metal_device_init: GPU name: Apple M4 Max
  657. ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
  658. ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
  659. ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
  660. ggml_metal_device_init: simdgroup reduction = true
  661. ggml_metal_device_init: simdgroup matrix mul. = true
  662. ggml_metal_device_init: has unified memory = true
  663. ggml_metal_device_init: has bfloat = true
  664. ggml_metal_device_init: use residency sets = true
  665. ggml_metal_device_init: use shared buffers = true
  666. ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
  667. build: 6789 (3d4e86bb) with Apple clang version 17.0.0 (clang-1700.3.19.1) for arm64-apple-darwin25.0.0
  668. llama_model_load_from_file_impl: using device Metal (Apple M4 Max) (unknown id) - 110100 MiB free
  669. llama_model_loader: loaded meta data with 35 key-value pairs and 459 tensors from /Usersusermodels/dgx/gpt-oss-20b-mxfp4.gguf (version GGUF V3 (latest))
  670. llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
  671. llama_model_loader: - kv 0: general.architecture str = gpt-oss
  672. llama_model_loader: - kv 1: general.type str = model
  673. llama_model_loader: - kv 2: general.name str = Gpt Oss 20b
  674. llama_model_loader: - kv 3: general.basename str = gpt-oss
  675. llama_model_loader: - kv 4: general.size_label str = 20B
  676. llama_model_loader: - kv 5: general.license str = apache-2.0
  677. llama_model_loader: - kv 6: general.tags arr[str,2] = ["vllm", "text-generation"]
  678. llama_model_loader: - kv 7: gpt-oss.block_count u32 = 24
  679. llama_model_loader: - kv 8: gpt-oss.context_length u32 = 131072
  680. llama_model_loader: - kv 9: gpt-oss.embedding_length u32 = 2880
  681. llama_model_loader: - kv 10: gpt-oss.feed_forward_length u32 = 2880
  682. llama_model_loader: - kv 11: gpt-oss.attention.head_count u32 = 64
  683. llama_model_loader: - kv 12: gpt-oss.attention.head_count_kv u32 = 8
  684. llama_model_loader: - kv 13: gpt-oss.rope.freq_base f32 = 150000.000000
  685. llama_model_loader: - kv 14: gpt-oss.attention.layer_norm_rms_epsilon f32 = 0.000010
  686. llama_model_loader: - kv 15: gpt-oss.expert_count u32 = 32
  687. llama_model_loader: - kv 16: gpt-oss.expert_used_count u32 = 4
  688. llama_model_loader: - kv 17: gpt-oss.attention.key_length u32 = 64
  689. llama_model_loader: - kv 18: gpt-oss.attention.value_length u32 = 64
  690. llama_model_loader: - kv 19: gpt-oss.attention.sliding_window u32 = 128
  691. llama_model_loader: - kv 20: gpt-oss.expert_feed_forward_length u32 = 2880
  692. llama_model_loader: - kv 21: gpt-oss.rope.scaling.type str = yarn
  693. llama_model_loader: - kv 22: gpt-oss.rope.scaling.factor f32 = 32.000000
  694. llama_model_loader: - kv 23: gpt-oss.rope.scaling.original_context_length u32 = 4096
  695. llama_model_loader: - kv 24: tokenizer.ggml.model str = gpt2
  696. llama_model_loader: - kv 25: tokenizer.ggml.pre str = gpt-4o
  697. llama_model_loader: - kv 26: tokenizer.ggml.tokens arr[str,201088] = ["!", "\"", "#", "$", "%", "&", "'", ...
  698. llama_model_loader: - kv 27: tokenizer.ggml.token_type arr[i32,201088] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
  699. llama_model_loader: - kv 28: tokenizer.ggml.merges arr[str,446189] = ["Ġ Ġ", "Ġ ĠĠĠ", "ĠĠ ĠĠ", "...
  700. llama_model_loader: - kv 29: tokenizer.ggml.bos_token_id u32 = 199998
  701. llama_model_loader: - kv 30: tokenizer.ggml.eos_token_id u32 = 200002
  702. llama_model_loader: - kv 31: tokenizer.ggml.padding_token_id u32 = 199999
  703. llama_model_loader: - kv 32: tokenizer.chat_template str = {#-\n In addition to the normal input...
  704. llama_model_loader: - kv 33: general.quantization_version u32 = 2
  705. llama_model_loader: - kv 34: general.file_type u32 = 38
  706. llama_model_loader: - type f32: 289 tensors
  707. llama_model_loader: - type q8_0: 98 tensors
  708. llama_model_loader: - type mxfp4: 72 tensors
  709. print_info: file format = GGUF V3 (latest)
  710. print_info: file type = MXFP4 MoE
  711. print_info: file size = 11.27 GiB (4.63 BPW)
  712. load: printing all EOG tokens:
  713. load: - 199999 ('<|endoftext|>')
  714. load: - 200002 ('<|return|>')
  715. load: - 200007 ('<|end|>')
  716. load: - 200012 ('<|call|>')
  717. load: special_eog_ids contains both '<|return|>' and '<|call|>' tokens, removing '<|end|>' token from EOG list
  718. load: special tokens cache size = 21
  719. load: token to piece cache size = 1.3332 MB
  720. print_info: arch = gpt-oss
  721. print_info: vocab_only = 0
  722. print_info: n_ctx_train = 131072
  723. print_info: n_embd = 2880
  724. print_info: n_layer = 24
  725. print_info: n_head = 64
  726. print_info: n_head_kv = 8
  727. print_info: n_rot = 64
  728. print_info: n_swa = 128
  729. print_info: is_swa_any = 1
  730. print_info: n_embd_head_k = 64
  731. print_info: n_embd_head_v = 64
  732. print_info: n_gqa = 8
  733. print_info: n_embd_k_gqa = 512
  734. print_info: n_embd_v_gqa = 512
  735. print_info: f_norm_eps = 0.0e+00
  736. print_info: f_norm_rms_eps = 1.0e-05
  737. print_info: f_clamp_kqv = 0.0e+00
  738. print_info: f_max_alibi_bias = 0.0e+00
  739. print_info: f_logit_scale = 0.0e+00
  740. print_info: f_attn_scale = 0.0e+00
  741. print_info: n_ff = 2880
  742. print_info: n_expert = 32
  743. print_info: n_expert_used = 4
  744. print_info: causal attn = 1
  745. print_info: pooling type = 0
  746. print_info: rope type = 2
  747. print_info: rope scaling = yarn
  748. print_info: freq_base_train = 150000.0
  749. print_info: freq_scale_train = 0.03125
  750. print_info: n_ctx_orig_yarn = 4096
  751. print_info: rope_finetuned = unknown
  752. print_info: model type = 20B
  753. print_info: model params = 20.91 B
  754. print_info: general.name = Gpt Oss 20b
  755. print_info: n_ff_exp = 2880
  756. print_info: vocab type = BPE
  757. print_info: n_vocab = 201088
  758. print_info: n_merges = 446189
  759. print_info: BOS token = 199998 '<|startoftext|>'
  760. print_info: EOS token = 200002 '<|return|>'
  761. print_info: EOT token = 200007 '<|end|>'
  762. print_info: PAD token = 199999 '<|endoftext|>'
  763. print_info: LF token = 198 'Ċ'
  764. print_info: EOG token = 199999 '<|endoftext|>'
  765. print_info: EOG token = 200002 '<|return|>'
  766. print_info: EOG token = 200012 '<|call|>'
  767. print_info: max token length = 256
  768. load_tensors: loading model tensors, this can take a while... (mmap = true)
  769. load_tensors: offloading 24 repeating layers to GPU
  770. load_tensors: offloading output layer to GPU
  771. load_tensors: offloaded 25/25 layers to GPU
  772. load_tensors: Metal_Mapped model buffer size = 11536.18 MiB
  773. load_tensors: CPU_Mapped model buffer size = 586.82 MiB
  774. ................................................................................
  775. llama_init_from_model: model default pooling_type is [0], but [-1] was specified
  776. llama_context: constructing llama_context
  777. llama_context: n_seq_max = 32
  778. llama_context: n_ctx = 300000
  779. llama_context: n_ctx_per_seq = 9375
  780. llama_context: n_batch = 2048
  781. llama_context: n_ubatch = 2048
  782. llama_context: causal_attn = 1
  783. llama_context: flash_attn = enabled
  784. llama_context: kv_unified = false
  785. llama_context: freq_base = 150000.0
  786. llama_context: freq_scale = 0.03125
  787. llama_context: n_ctx_per_seq (9375) < n_ctx_train (131072) -- the full capacity of the model will not be utilized
  788. ggml_metal_init: allocating
  789. ggml_metal_init: picking default device: Apple M4 Max
  790. ggml_metal_init: use bfloat = true
  791. ggml_metal_init: use fusion = true
  792. ggml_metal_init: use concurrency = true
  793. ggml_metal_init: use graph optimize = true
  794. llama_context: CPU output buffer size = 24.55 MiB
  795. llama_kv_cache_iswa: creating non-SWA KV cache, size = 9472 cells
  796. llama_kv_cache: Metal KV buffer size = 7104.00 MiB
  797. llama_kv_cache: size = 7104.00 MiB ( 9472 cells, 12 layers, 32/32 seqs), K (f16): 3552.00 MiB, V (f16): 3552.00 MiB
  798. llama_kv_cache_iswa: creating SWA KV cache, size = 2304 cells
  799. llama_kv_cache: Metal KV buffer size = 1728.00 MiB
  800. llama_kv_cache: size = 1728.00 MiB ( 2304 cells, 12 layers, 32/32 seqs), K (f16): 864.00 MiB, V (f16): 864.00 MiB
  801. llama_context: Metal compute buffer size = 7885.59 MiB
  802. llama_context: CPU compute buffer size = 114.59 MiB
  803. llama_context: graph nodes = 1400
  804. llama_context: graph splits = 2
  805.  
  806. main: n_kv_max = 303104, n_batch = 2048, n_ubatch = 2048, flash_attn = 1, is_pp_shared = 0, n_gpu_layers = 999, n_threads = 12, n_threads_batch = 12
  807.  
  808. | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s |
  809. |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
  810. | 4096 | 32 | 1 | 4128 | 2.655 | 1542.84 | 0.310 | 103.25 | 2.965 | 1392.35 |
  811. | 4096 | 32 | 2 | 8256 | 6.764 | 1211.13 | 0.566 | 113.12 | 7.330 | 1126.38 |
  812. | 4096 | 32 | 4 | 16512 | 12.218 | 1341.02 | 0.919 | 139.35 | 13.136 | 1257.00 |
  813. | 4096 | 32 | 8 | 33024 | 24.281 | 1349.52 | 1.751 | 146.24 | 26.032 | 1268.60 |
  814. | 4096 | 32 | 16 | 66048 | 47.249 | 1387.04 | 2.898 | 176.67 | 50.147 | 1317.09 |
  815. | 4096 | 32 | 32 | 132096 | 90.107 | 1454.62 | 3.180 | 322.00 | 93.287 | 1416.01 |
  816. | 8192 | 32 | 1 | 8224 | 5.998 | 1365.71 | 0.304 | 105.40 | 6.302 | 1304.99 |
  817. | 8192 | 32 | 2 | 16448 | 11.846 | 1383.03 | 0.492 | 129.97 | 12.339 | 1333.02 |
  818. | 8192 | 32 | 4 | 32896 | 23.743 | 1380.08 | 0.877 | 146.01 | 24.620 | 1336.14 |
  819. | 8192 | 32 | 8 | 65792 | 48.095 | 1362.65 | 1.690 | 151.48 | 49.785 | 1321.53 |
  820. | 8192 | 32 | 16 | 131584 | 96.484 | 1358.49 | 2.780 | 184.17 | 99.264 | 1325.60 |
  821. | 8192 | 32 | 32 | 263168 | 192.008 | 1365.28 | 3.552 | 288.28 | 195.560 | 1345.72 |
  822.  
  823. llama_perf_context_print: load time = 1143.70 ms
  824. llama_perf_context_print: prompt eval time = 580230.14 ms / 778128 tokens ( 0.75 ms per token, 1341.07 tokens per second)
  825. llama_perf_context_print: eval time = 613.09 ms / 64 runs ( 9.58 ms per token, 104.39 tokens per second)
  826. llama_perf_context_print: total time = 581948.28 ms / 778192 tokens
  827. llama_perf_context_print: graphs reused = 372
  828. ggml_metal_free: deallocating
  829.  
  830.  
  831. === Parallel Benchmark Results ===
  832. Model: qwen2.5-coder-7b-instruct-q8_0.gguf
  833. Date: Sun Oct 19 19:08:54 EDT 2025
  834. macOS Version: 26.0.1
  835. Hardware: Mac16,6
  836. Chip: Apple M4 Max
  837. Total RAM: 128.00 GB
  838. CPU Cores: 16
  839. GPU Layers: 999
  840. Flash Attention: 1
  841.  
  842. === Benchmark Parameters ===
  843. Context Size: 300000
  844. Ubatch Size: 2048
  845. Prompt Tokens: 4096,8192
  846. Generation Tokens: 32
  847. Parallel Requests: 1,2,4,8,16,32
  848.  
  849. === Results ===
  850. ggml_metal_library_init: using embedded metal library
  851. ggml_metal_library_init: loaded in 0.005 sec
  852. ggml_metal_device_init: GPU name: Apple M4 Max
  853. ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
  854. ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
  855. ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
  856. ggml_metal_device_init: simdgroup reduction = true
  857. ggml_metal_device_init: simdgroup matrix mul. = true
  858. ggml_metal_device_init: has unified memory = true
  859. ggml_metal_device_init: has bfloat = true
  860. ggml_metal_device_init: use residency sets = true
  861. ggml_metal_device_init: use shared buffers = true
  862. ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
  863. build: 6789 (3d4e86bb) with Apple clang version 17.0.0 (clang-1700.3.19.1) for arm64-apple-darwin25.0.0
  864. llama_model_load_from_file_impl: using device Metal (Apple M4 Max) (unknown id) - 110100 MiB free
  865. llama_model_loader: loaded meta data with 29 key-value pairs and 339 tensors from /Usersusermodels/dgx/qwen2.5-coder-7b-instruct-q8_0.gguf (version GGUF V3 (latest))
  866. llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
  867. llama_model_loader: - kv 0: general.architecture str = qwen2
  868. llama_model_loader: - kv 1: general.type str = model
  869. llama_model_loader: - kv 2: general.name str = Qwen2.5 Coder 7B Instruct GGUF
  870. llama_model_loader: - kv 3: general.finetune str = Instruct-GGUF
  871. llama_model_loader: - kv 4: general.basename str = Qwen2.5-Coder
  872. llama_model_loader: - kv 5: general.size_label str = 7B
  873. llama_model_loader: - kv 6: qwen2.block_count u32 = 28
  874. llama_model_loader: - kv 7: qwen2.context_length u32 = 131072
  875. llama_model_loader: - kv 8: qwen2.embedding_length u32 = 3584
  876. llama_model_loader: - kv 9: qwen2.feed_forward_length u32 = 18944
  877. llama_model_loader: - kv 10: qwen2.attention.head_count u32 = 28
  878. llama_model_loader: - kv 11: qwen2.attention.head_count_kv u32 = 4
  879. llama_model_loader: - kv 12: qwen2.rope.freq_base f32 = 1000000.000000
  880. llama_model_loader: - kv 13: qwen2.attention.layer_norm_rms_epsilon f32 = 0.000001
  881. llama_model_loader: - kv 14: general.file_type u32 = 7
  882. llama_model_loader: - kv 15: tokenizer.ggml.model str = gpt2
  883. llama_model_loader: - kv 16: tokenizer.ggml.pre str = qwen2
  884. llama_model_loader: - kv 17: tokenizer.ggml.tokens arr[str,152064] = ["!", "\"", "#", "$", "%", "&", "'", ...
  885. llama_model_loader: - kv 18: tokenizer.ggml.token_type arr[i32,152064] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
  886. llama_model_loader: - kv 19: tokenizer.ggml.merges arr[str,151387] = ["Ġ Ġ", "ĠĠ ĠĠ", "i n", "Ġ t",...
  887. llama_model_loader: - kv 20: tokenizer.ggml.eos_token_id u32 = 151645
  888. llama_model_loader: - kv 21: tokenizer.ggml.padding_token_id u32 = 151643
  889. llama_model_loader: - kv 22: tokenizer.ggml.bos_token_id u32 = 151643
  890. llama_model_loader: - kv 23: tokenizer.ggml.add_bos_token bool = false
  891. llama_model_loader: - kv 24: tokenizer.chat_template str = {%- if tools %}\n {{- '<|im_start|>...
  892. llama_model_loader: - kv 25: general.quantization_version u32 = 2
  893. llama_model_loader: - kv 26: split.no u16 = 0
  894. llama_model_loader: - kv 27: split.count u16 = 0
  895. llama_model_loader: - kv 28: split.tensors.count i32 = 339
  896. llama_model_loader: - type f32: 141 tensors
  897. llama_model_loader: - type q8_0: 198 tensors
  898. print_info: file format = GGUF V3 (latest)
  899. print_info: file type = Q8_0
  900. print_info: file size = 7.54 GiB (8.50 BPW)
  901. load: printing all EOG tokens:
  902. load: - 151643 ('<|endoftext|>')
  903. load: - 151645 ('<|im_end|>')
  904. load: - 151662 ('<|fim_pad|>')
  905. load: - 151663 ('<|repo_name|>')
  906. load: - 151664 ('<|file_sep|>')
  907. load: special tokens cache size = 22
  908. load: token to piece cache size = 0.9310 MB
  909. print_info: arch = qwen2
  910. print_info: vocab_only = 0
  911. print_info: n_ctx_train = 131072
  912. print_info: n_embd = 3584
  913. print_info: n_layer = 28
  914. print_info: n_head = 28
  915. print_info: n_head_kv = 4
  916. print_info: n_rot = 128
  917. print_info: n_swa = 0
  918. print_info: is_swa_any = 0
  919. print_info: n_embd_head_k = 128
  920. print_info: n_embd_head_v = 128
  921. print_info: n_gqa = 7
  922. print_info: n_embd_k_gqa = 512
  923. print_info: n_embd_v_gqa = 512
  924. print_info: f_norm_eps = 0.0e+00
  925. print_info: f_norm_rms_eps = 1.0e-06
  926. print_info: f_clamp_kqv = 0.0e+00
  927. print_info: f_max_alibi_bias = 0.0e+00
  928. print_info: f_logit_scale = 0.0e+00
  929. print_info: f_attn_scale = 0.0e+00
  930. print_info: n_ff = 18944
  931. print_info: n_expert = 0
  932. print_info: n_expert_used = 0
  933. print_info: causal attn = 1
  934. print_info: pooling type = -1
  935. print_info: rope type = 2
  936. print_info: rope scaling = linear
  937. print_info: freq_base_train = 1000000.0
  938. print_info: freq_scale_train = 1
  939. print_info: n_ctx_orig_yarn = 131072
  940. print_info: rope_finetuned = unknown
  941. print_info: model type = 7B
  942. print_info: model params = 7.62 B
  943. print_info: general.name = Qwen2.5 Coder 7B Instruct GGUF
  944. print_info: vocab type = BPE
  945. print_info: n_vocab = 152064
  946. print_info: n_merges = 151387
  947. print_info: BOS token = 151643 '<|endoftext|>'
  948. print_info: EOS token = 151645 '<|im_end|>'
  949. print_info: EOT token = 151645 '<|im_end|>'
  950. print_info: PAD token = 151643 '<|endoftext|>'
  951. print_info: LF token = 198 'Ċ'
  952. print_info: FIM PRE token = 151659 '<|fim_prefix|>'
  953. print_info: FIM SUF token = 151661 '<|fim_suffix|>'
  954. print_info: FIM MID token = 151660 '<|fim_middle|>'
  955. print_info: FIM PAD token = 151662 '<|fim_pad|>'
  956. print_info: FIM REP token = 151663 '<|repo_name|>'
  957. print_info: FIM SEP token = 151664 '<|file_sep|>'
  958. print_info: EOG token = 151643 '<|endoftext|>'
  959. print_info: EOG token = 151645 '<|im_end|>'
  960. print_info: EOG token = 151662 '<|fim_pad|>'
  961. print_info: EOG token = 151663 '<|repo_name|>'
  962. print_info: EOG token = 151664 '<|file_sep|>'
  963. print_info: max token length = 256
  964. load_tensors: loading model tensors, this can take a while... (mmap = true)
  965. load_tensors: offloading 28 repeating layers to GPU
  966. load_tensors: offloading output layer to GPU
  967. load_tensors: offloaded 29/29 layers to GPU
  968. load_tensors: Metal_Mapped model buffer size = 7165.44 MiB
  969. load_tensors: CPU_Mapped model buffer size = 552.23 MiB
  970. .......................................................................................
  971. llama_context: constructing llama_context
  972. llama_context: n_seq_max = 32
  973. llama_context: n_ctx = 300000
  974. llama_context: n_ctx_per_seq = 9375
  975. llama_context: n_batch = 2048
  976. llama_context: n_ubatch = 2048
  977. llama_context: causal_attn = 1
  978. llama_context: flash_attn = enabled
  979. llama_context: kv_unified = false
  980. llama_context: freq_base = 1000000.0
  981. llama_context: freq_scale = 1
  982. llama_context: n_ctx_per_seq (9375) < n_ctx_train (131072) -- the full capacity of the model will not be utilized
  983. ggml_metal_init: allocating
  984. ggml_metal_init: picking default device: Apple M4 Max
  985. ggml_metal_init: use bfloat = true
  986. ggml_metal_init: use fusion = true
  987. ggml_metal_init: use concurrency = true
  988. ggml_metal_init: use graph optimize = true
  989. llama_context: CPU output buffer size = 18.56 MiB
  990. llama_kv_cache: Metal KV buffer size = 16576.00 MiB
  991. llama_kv_cache: size = 16576.00 MiB ( 9472 cells, 28 layers, 32/32 seqs), K (f16): 8288.00 MiB, V (f16): 8288.00 MiB
  992. llama_context: Metal compute buffer size = 1216.00 MiB
  993. llama_context: CPU compute buffer size = 102.05 MiB
  994. llama_context: graph nodes = 1015
  995. llama_context: graph splits = 2
  996.  
  997. main: n_kv_max = 303104, n_batch = 2048, n_ubatch = 2048, flash_attn = 1, is_pp_shared = 0, n_gpu_layers = 999, n_threads = 12, n_threads_batch = 12
  998.  
  999. | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s |
  1000. |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
  1001. | 4096 | 32 | 1 | 4128 | 5.597 | 731.85 | 0.573 | 55.85 | 6.170 | 669.08 |
  1002. | 4096 | 32 | 2 | 8256 | 11.508 | 711.88 | 0.641 | 99.90 | 12.148 | 679.61 |
  1003. | 4096 | 32 | 4 | 16512 | 22.183 | 738.57 | 1.226 | 104.40 | 23.409 | 705.35 |
  1004. | 4096 | 32 | 8 | 33024 | 47.485 | 690.07 | 2.780 | 92.08 | 50.265 | 656.99 |
  1005. | 4096 | 32 | 16 | 66048 | 95.468 | 686.47 | 2.453 | 208.69 | 97.921 | 674.50 |
  1006. | 4096 | 32 | 32 | 132096 | 177.133 | 739.96 | 2.953 | 346.74 | 180.086 | 733.52 |
  1007. | 8192 | 32 | 1 | 8224 | 11.835 | 692.19 | 0.594 | 53.90 | 12.429 | 661.70 |
  1008. | 8192 | 32 | 2 | 16448 | 23.895 | 685.68 | 0.702 | 91.13 | 24.597 | 668.70 |
  1009. | 8192 | 32 | 4 | 32896 | 47.030 | 696.75 | 1.344 | 95.26 | 48.374 | 680.04 |
  1010. | 8192 | 32 | 8 | 65792 | 95.530 | 686.03 | 2.633 | 97.23 | 98.163 | 670.23 |
  1011. | 8192 | 32 | 16 | 131584 | 191.551 | 684.27 | 2.880 | 177.76 | 194.431 | 676.76 |
  1012. | 8192 | 32 | 32 | 263168 | 382.786 | 684.83 | 4.050 | 252.83 | 386.837 | 680.31 |
  1013.  
  1014. llama_perf_context_print: load time = 1139.37 ms
  1015. llama_perf_context_print: prompt eval time = 1133723.39 ms / 778128 tokens ( 1.46 ms per token, 686.35 tokens per second)
  1016. llama_perf_context_print: eval time = 1166.11 ms / 64 runs ( 18.22 ms per token, 54.88 tokens per second)
  1017. llama_perf_context_print: total time = 1135998.75 ms / 778192 tokens
  1018. llama_perf_context_print: graphs reused = 372
  1019. ggml_metal_free: deallocating
  1020.  
  1021.  
  1022. === Parallel Benchmark Results ===
  1023. Model: qwen3-30b-a3b-q8_0.gguf
  1024. Date: Sun Oct 19 18:35:23 EDT 2025
  1025. macOS Version: 26.0.1
  1026. Hardware: Mac16,6
  1027. Chip: Apple M4 Max
  1028. Total RAM: 128.00 GB
  1029. CPU Cores: 16
  1030. GPU Layers: 999
  1031. Flash Attention: 1
  1032.  
  1033. === Benchmark Parameters ===
  1034. Context Size: 300000
  1035. Ubatch Size: 2048
  1036. Prompt Tokens: 4096,8192
  1037. Generation Tokens: 32
  1038. Parallel Requests: 1,2,4,8,16,32
  1039.  
  1040. === Results ===
  1041. ggml_metal_library_init: using embedded metal library
  1042. ggml_metal_library_init: loaded in 0.006 sec
  1043. ggml_metal_device_init: GPU name: Apple M4 Max
  1044. ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
  1045. ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
  1046. ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
  1047. ggml_metal_device_init: simdgroup reduction = true
  1048. ggml_metal_device_init: simdgroup matrix mul. = true
  1049. ggml_metal_device_init: has unified memory = true
  1050. ggml_metal_device_init: has bfloat = true
  1051. ggml_metal_device_init: use residency sets = true
  1052. ggml_metal_device_init: use shared buffers = true
  1053. ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
  1054. build: 6789 (3d4e86bb) with Apple clang version 17.0.0 (clang-1700.3.19.1) for arm64-apple-darwin25.0.0
  1055. llama_model_load_from_file_impl: using device Metal (Apple M4 Max) (unknown id) - 110100 MiB free
  1056. llama_model_loader: loaded meta data with 31 key-value pairs and 579 tensors from /Usersusermodels/dgx/qwen3-30b-a3b-q8_0.gguf (version GGUF V3 (latest))
  1057. llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
  1058. llama_model_loader: - kv 0: general.architecture str = qwen3moe
  1059. llama_model_loader: - kv 1: general.type str = model
  1060. llama_model_loader: - kv 2: general.name str = Qwen3 30Ba3 Instruct
  1061. llama_model_loader: - kv 3: general.finetune str = 30Ba3-Instruct
  1062. llama_model_loader: - kv 4: general.basename str = Qwen3
  1063. llama_model_loader: - kv 5: general.size_label str = 128x1.8B
  1064. llama_model_loader: - kv 6: qwen3moe.block_count u32 = 48
  1065. llama_model_loader: - kv 7: qwen3moe.context_length u32 = 40960
  1066. llama_model_loader: - kv 8: qwen3moe.embedding_length u32 = 2048
  1067. llama_model_loader: - kv 9: qwen3moe.feed_forward_length u32 = 6144
  1068. llama_model_loader: - kv 10: qwen3moe.attention.head_count u32 = 32
  1069. llama_model_loader: - kv 11: qwen3moe.attention.head_count_kv u32 = 4
  1070. llama_model_loader: - kv 12: qwen3moe.rope.freq_base f32 = 1000000.000000
  1071. llama_model_loader: - kv 13: qwen3moe.attention.layer_norm_rms_epsilon f32 = 0.000001
  1072. llama_model_loader: - kv 14: qwen3moe.expert_used_count u32 = 8
  1073. llama_model_loader: - kv 15: qwen3moe.attention.key_length u32 = 128
  1074. llama_model_loader: - kv 16: qwen3moe.attention.value_length u32 = 128
  1075. llama_model_loader: - kv 17: qwen3moe.expert_count u32 = 128
  1076. llama_model_loader: - kv 18: qwen3moe.expert_feed_forward_length u32 = 768
  1077. llama_model_loader: - kv 19: tokenizer.ggml.model str = gpt2
  1078. llama_model_loader: - kv 20: tokenizer.ggml.pre str = qwen2
  1079. llama_model_loader: - kv 21: tokenizer.ggml.tokens arr[str,151936] = ["!", "\"", "#", "$", "%", "&", "'", ...
  1080. llama_model_loader: - kv 22: tokenizer.ggml.token_type arr[i32,151936] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
  1081. llama_model_loader: - kv 23: tokenizer.ggml.merges arr[str,151387] = ["Ġ Ġ", "ĠĠ ĠĠ", "i n", "Ġ t",...
  1082. llama_model_loader: - kv 24: tokenizer.ggml.eos_token_id u32 = 151645
  1083. llama_model_loader: - kv 25: tokenizer.ggml.padding_token_id u32 = 151643
  1084. llama_model_loader: - kv 26: tokenizer.ggml.bos_token_id u32 = 151643
  1085. llama_model_loader: - kv 27: tokenizer.ggml.add_bos_token bool = false
  1086. llama_model_loader: - kv 28: tokenizer.chat_template str = {%- if tools %}\n {{- '<|im_start|>...
  1087. llama_model_loader: - kv 29: general.quantization_version u32 = 2
  1088. llama_model_loader: - kv 30: general.file_type u32 = 7
  1089. llama_model_loader: - type f32: 241 tensors
  1090. llama_model_loader: - type q8_0: 338 tensors
  1091. print_info: file format = GGUF V3 (latest)
  1092. print_info: file type = Q8_0
  1093. print_info: file size = 30.25 GiB (8.51 BPW)
  1094. load: printing all EOG tokens:
  1095. load: - 151643 ('<|endoftext|>')
  1096. load: - 151645 ('<|im_end|>')
  1097. load: - 151662 ('<|fim_pad|>')
  1098. load: - 151663 ('<|repo_name|>')
  1099. load: - 151664 ('<|file_sep|>')
  1100. load: special tokens cache size = 26
  1101. load: token to piece cache size = 0.9311 MB
  1102. print_info: arch = qwen3moe
  1103. print_info: vocab_only = 0
  1104. print_info: n_ctx_train = 40960
  1105. print_info: n_embd = 2048
  1106. print_info: n_layer = 48
  1107. print_info: n_head = 32
  1108. print_info: n_head_kv = 4
  1109. print_info: n_rot = 128
  1110. print_info: n_swa = 0
  1111. print_info: is_swa_any = 0
  1112. print_info: n_embd_head_k = 128
  1113. print_info: n_embd_head_v = 128
  1114. print_info: n_gqa = 8
  1115. print_info: n_embd_k_gqa = 512
  1116. print_info: n_embd_v_gqa = 512
  1117. print_info: f_norm_eps = 0.0e+00
  1118. print_info: f_norm_rms_eps = 1.0e-06
  1119. print_info: f_clamp_kqv = 0.0e+00
  1120. print_info: f_max_alibi_bias = 0.0e+00
  1121. print_info: f_logit_scale = 0.0e+00
  1122. print_info: f_attn_scale = 0.0e+00
  1123. print_info: n_ff = 6144
  1124. print_info: n_expert = 128
  1125. print_info: n_expert_used = 8
  1126. print_info: causal attn = 1
  1127. print_info: pooling type = 0
  1128. print_info: rope type = 2
  1129. print_info: rope scaling = linear
  1130. print_info: freq_base_train = 1000000.0
  1131. print_info: freq_scale_train = 1
  1132. print_info: n_ctx_orig_yarn = 40960
  1133. print_info: rope_finetuned = unknown
  1134. print_info: model type = 30B.A3B
  1135. print_info: model params = 30.53 B
  1136. print_info: general.name = Qwen3 30Ba3 Instruct
  1137. print_info: n_ff_exp = 768
  1138. print_info: vocab type = BPE
  1139. print_info: n_vocab = 151936
  1140. print_info: n_merges = 151387
  1141. print_info: BOS token = 151643 '<|endoftext|>'
  1142. print_info: EOS token = 151645 '<|im_end|>'
  1143. print_info: EOT token = 151645 '<|im_end|>'
  1144. print_info: PAD token = 151643 '<|endoftext|>'
  1145. print_info: LF token = 198 'Ċ'
  1146. print_info: FIM PRE token = 151659 '<|fim_prefix|>'
  1147. print_info: FIM SUF token = 151661 '<|fim_suffix|>'
  1148. print_info: FIM MID token = 151660 '<|fim_middle|>'
  1149. print_info: FIM PAD token = 151662 '<|fim_pad|>'
  1150. print_info: FIM REP token = 151663 '<|repo_name|>'
  1151. print_info: FIM SEP token = 151664 '<|file_sep|>'
  1152. print_info: EOG token = 151643 '<|endoftext|>'
  1153. print_info: EOG token = 151645 '<|im_end|>'
  1154. print_info: EOG token = 151662 '<|fim_pad|>'
  1155. print_info: EOG token = 151663 '<|repo_name|>'
  1156. print_info: EOG token = 151664 '<|file_sep|>'
  1157. print_info: max token length = 256
  1158. load_tensors: loading model tensors, this can take a while... (mmap = true)
  1159. load_tensors: offloading 48 repeating layers to GPU
  1160. load_tensors: offloading output layer to GPU
  1161. load_tensors: offloaded 49/49 layers to GPU
  1162. load_tensors: Metal_Mapped model buffer size = 30973.40 MiB
  1163. load_tensors: CPU_Mapped model buffer size = 315.30 MiB
  1164. ...................................................................................................
  1165. llama_init_from_model: model default pooling_type is [0], but [-1] was specified
  1166. llama_context: constructing llama_context
  1167. llama_context: n_seq_max = 32
  1168. llama_context: n_ctx = 300000
  1169. llama_context: n_ctx_per_seq = 9375
  1170. llama_context: n_batch = 2048
  1171. llama_context: n_ubatch = 2048
  1172. llama_context: causal_attn = 1
  1173. llama_context: flash_attn = enabled
  1174. llama_context: kv_unified = false
  1175. llama_context: freq_base = 1000000.0
  1176. llama_context: freq_scale = 1
  1177. llama_context: n_ctx_per_seq (9375) < n_ctx_train (40960) -- the full capacity of the model will not be utilized
  1178. ggml_metal_init: allocating
  1179. ggml_metal_init: picking default device: Apple M4 Max
  1180. ggml_metal_init: use bfloat = true
  1181. ggml_metal_init: use fusion = true
  1182. ggml_metal_init: use concurrency = true
  1183. ggml_metal_init: use graph optimize = true
  1184. llama_context: CPU output buffer size = 18.55 MiB
  1185. llama_kv_cache: Metal KV buffer size = 28416.00 MiB
  1186. llama_kv_cache: size = 28416.00 MiB ( 9472 cells, 48 layers, 32/32 seqs), K (f16): 14208.00 MiB, V (f16): 14208.00 MiB
  1187. llama_context: Metal compute buffer size = 7265.07 MiB
  1188. llama_context: CPU compute buffer size = 90.05 MiB
  1189. llama_context: graph nodes = 3079
  1190. llama_context: graph splits = 2
  1191.  
  1192. main: n_kv_max = 303104, n_batch = 2048, n_ubatch = 2048, flash_attn = 1, is_pp_shared = 0, n_gpu_layers = 999, n_threads = 12, n_threads_batch = 12
  1193.  
  1194. | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s |
  1195. |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
  1196. | 4096 | 32 | 1 | 4128 | 2.635 | 1554.20 | 0.433 | 73.87 | 3.069 | 1345.22 |
  1197. | 4096 | 32 | 2 | 8256 | 6.712 | 1220.48 | 0.645 | 99.18 | 7.357 | 1122.13 |
  1198. | 4096 | 32 | 4 | 16512 | 13.668 | 1198.74 | 1.109 | 115.43 | 14.777 | 1117.44 |
  1199. | 4096 | 32 | 8 | 33024 | 26.887 | 1218.74 | 2.115 | 121.02 | 29.002 | 1138.67 |
  1200. | 4096 | 32 | 16 | 66048 | 53.227 | 1231.25 | 4.077 | 125.58 | 57.304 | 1152.58 |
  1201. | 4096 | 32 | 32 | 132096 | 114.297 | 1146.77 | 5.387 | 190.10 | 119.683 | 1103.71 |
  1202. | 8192 | 32 | 1 | 8224 | 8.294 | 987.73 | 0.484 | 66.11 | 8.778 | 936.91 |
  1203. | 8192 | 32 | 2 | 16448 | 16.465 | 995.09 | 0.751 | 85.24 | 17.216 | 955.41 |
  1204. | 8192 | 32 | 4 | 32896 | 32.029 | 1023.07 | 1.373 | 93.22 | 33.402 | 984.85 |
  1205. | 8192 | 32 | 8 | 65792 | 64.867 | 1010.32 | 2.433 | 105.22 | 67.300 | 977.60 |
  1206. | 8192 | 32 | 16 | 131584 | 129.668 | 1010.83 | 4.935 | 103.76 | 134.603 | 977.57 |
  1207. | 8192 | 32 | 32 | 263168 | 260.440 | 1006.54 | 6.647 | 154.05 | 267.087 | 985.33 |
  1208.  
  1209. llama_perf_context_print: load time = 2699.68 ms
  1210. llama_perf_context_print: prompt eval time = 758745.76 ms / 778128 tokens ( 0.98 ms per token, 1025.55 tokens per second)
  1211. llama_perf_context_print: eval time = 916.61 ms / 64 runs ( 14.32 ms per token, 69.82 tokens per second)
  1212. llama_perf_context_print: total time = 762309.19 ms / 778192 tokens
  1213. llama_perf_context_print: graphs reused = 372
  1214. ggml_metal_free: deallocating
  1215.  
  1216.  
  1217. === Sequential Benchmark Results ===
  1218. Model: gemma-3-4b-it-q4_0.gguf
  1219. Date: Sun Oct 19 19:27:50 EDT 2025
  1220. macOS Version: 26.0.1
  1221. Hardware: Mac16,6
  1222. Chip: Apple M4 Max
  1223. Total RAM: 128.00 GB
  1224. CPU Cores: 16
  1225. GPU Layers: 999
  1226. Flash Attention: 1
  1227.  
  1228. === Benchmark Parameters ===
  1229. Prompt Sizes: 2048
  1230. Generation Sizes: 32
  1231. Batch Size:
  1232.  
  1233. === Results ===
  1234. ggml_metal_library_init: using embedded metal library
  1235. ggml_metal_library_init: loaded in 0.007 sec
  1236. ggml_metal_device_init: GPU name: Apple M4 Max
  1237. ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
  1238. ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
  1239. ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
  1240. ggml_metal_device_init: simdgroup reduction = true
  1241. ggml_metal_device_init: simdgroup matrix mul. = true
  1242. ggml_metal_device_init: has unified memory = true
  1243. ggml_metal_device_init: has bfloat = true
  1244. ggml_metal_device_init: use residency sets = true
  1245. ggml_metal_device_init: use shared buffers = true
  1246. ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
  1247. | model | size | params | backend | threads | n_ubatch | fa | test | t/s |
  1248. | ------------------------------ | ---------: | ---------: | ---------- | ------: | -------: | -: | --------------: | -------------------: |
  1249. | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 | 1639.55 ± 49.17 |
  1250. | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | tg32 | 132.69 ± 1.35 |
  1251. | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d4096 | 1536.50 ± 94.10 |
  1252. | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d4096 | 121.45 ± 1.36 |
  1253. | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d8192 | 1498.66 ± 19.88 |
  1254. | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d8192 | 118.23 ± 1.59 |
  1255. | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d16384 | 1385.92 ± 23.74 |
  1256. | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d16384 | 111.51 ± 0.26 |
  1257. | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d32768 | 1166.95 ± 16.65 |
  1258. | gemma3 4B Q4_0 | 2.20 GiB | 3.88 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d32768 | 99.56 ± 0.54 |
  1259.  
  1260. build: 3d4e86bb (6789)
  1261. === Sequential Benchmark Results ===
  1262. Model: glm-4.5-air-q4_k_m-00001-of-00002.gguf
  1263. Date: Sun Oct 19 19:44:11 EDT 2025
  1264. macOS Version: 26.0.1
  1265. Hardware: Mac16,6
  1266. Chip: Apple M4 Max
  1267. Total RAM: 128.00 GB
  1268. CPU Cores: 16
  1269. GPU Layers: 999
  1270. Flash Attention: 1
  1271.  
  1272. === Benchmark Parameters ===
  1273. Prompt Sizes: 2048
  1274. Generation Sizes: 32
  1275. Batch Size:
  1276.  
  1277. === Results ===
  1278. ggml_metal_library_init: using embedded metal library
  1279. ggml_metal_library_init: loaded in 0.005 sec
  1280. ggml_metal_device_init: GPU name: Apple M4 Max
  1281. ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
  1282. ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
  1283. ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
  1284. ggml_metal_device_init: simdgroup reduction = true
  1285. ggml_metal_device_init: simdgroup matrix mul. = true
  1286. ggml_metal_device_init: has unified memory = true
  1287. ggml_metal_device_init: has bfloat = true
  1288. ggml_metal_device_init: use residency sets = true
  1289. ggml_metal_device_init: use shared buffers = true
  1290. ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
  1291. | model | size | params | backend | threads | n_ubatch | fa | test | t/s |
  1292. | ------------------------------ | ---------: | ---------: | ---------- | ------: | -------: | -: | --------------: | -------------------: |
  1293. | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 | 312.02 ± 12.39 |
  1294. | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | tg32 | 32.06 ± 2.94 |
  1295. | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d4096 | 235.64 ± 2.57 |
  1296. | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d4096 | 29.56 ± 0.88 |
  1297. | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d8192 | 184.17 ± 6.39 |
  1298. | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d8192 | 28.33 ± 0.21 |
  1299. | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d16384 | 139.00 ± 1.75 |
  1300. | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d16384 | 23.06 ± 0.40 |
  1301. | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d32768 | 85.37 ± 0.75 |
  1302. | glm4moe 106B.A12B Q4_K - Medium | 68.45 GiB | 110.47 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d32768 | 16.23 ± 0.73 |
  1303.  
  1304. build: 3d4e86bb (6789)
  1305. === Sequential Benchmark Results ===
  1306. Model: gpt-oss-120b-mxfp4-00001-of-00003.gguf
  1307. Date: Sun Oct 19 17:42:31 EDT 2025
  1308. macOS Version: 26.0.1
  1309. Hardware: Mac16,6
  1310. Chip: Apple M4 Max
  1311. Total RAM: 128.00 GB
  1312. CPU Cores: 16
  1313. GPU Layers: 999
  1314. Flash Attention: 1
  1315.  
  1316. === Benchmark Parameters ===
  1317. Prompt Sizes: 2048
  1318. Generation Sizes: 32
  1319. Batch Size:
  1320.  
  1321. === Results ===
  1322. ggml_metal_library_init: using embedded metal library
  1323. ggml_metal_library_init: loaded in 0.006 sec
  1324. ggml_metal_device_init: GPU name: Apple M4 Max
  1325. ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
  1326. ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
  1327. ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
  1328. ggml_metal_device_init: simdgroup reduction = true
  1329. ggml_metal_device_init: simdgroup matrix mul. = true
  1330. ggml_metal_device_init: has unified memory = true
  1331. ggml_metal_device_init: has bfloat = true
  1332. ggml_metal_device_init: use residency sets = true
  1333. ggml_metal_device_init: use shared buffers = true
  1334. ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
  1335. | model | size | params | backend | threads | n_ubatch | fa | test | t/s |
  1336. | ------------------------------ | ---------: | ---------: | ---------- | ------: | -------: | -: | --------------: | -------------------: |
  1337. | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 | 949.11 ± 87.73 |
  1338. | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | tg32 | 81.68 ± 1.36 |
  1339. | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d4096 | 812.53 ± 13.49 |
  1340. | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d4096 | 74.77 ± 0.96 |
  1341. | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d8192 | 716.74 ± 4.71 |
  1342. | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d8192 | 73.03 ± 0.54 |
  1343. | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d16384 | 580.13 ± 6.30 |
  1344. | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d16384 | 66.01 ± 2.89 |
  1345. | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d32768 | 411.04 ± 3.79 |
  1346. | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d32768 | 57.12 ± 0.99 |
  1347.  
  1348. build: 3d4e86bb (6789)
  1349. === Sequential Benchmark Results ===
  1350. Model: gpt-oss-20b-mxfp4.gguf
  1351. Date: Sun Oct 19 17:22:09 EDT 2025
  1352. macOS Version: 26.0.1
  1353. Hardware: Mac16,6
  1354. Chip: Apple M4 Max
  1355. Total RAM: 128.00 GB
  1356. CPU Cores: 16
  1357. GPU Layers: 999
  1358. Flash Attention: 1
  1359.  
  1360. === Benchmark Parameters ===
  1361. Prompt Sizes: 2048
  1362. Generation Sizes: 32
  1363. Batch Size:
  1364.  
  1365. === Results ===
  1366. ggml_metal_library_init: using embedded metal library
  1367. ggml_metal_library_init: loaded in 0.004 sec
  1368. ggml_metal_device_init: GPU name: Apple M4 Max
  1369. ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
  1370. ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
  1371. ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
  1372. ggml_metal_device_init: simdgroup reduction = true
  1373. ggml_metal_device_init: simdgroup matrix mul. = true
  1374. ggml_metal_device_init: has unified memory = true
  1375. ggml_metal_device_init: has bfloat = true
  1376. ggml_metal_device_init: use residency sets = true
  1377. ggml_metal_device_init: use shared buffers = true
  1378. ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
  1379. | model | size | params | backend | threads | n_ubatch | fa | test | t/s |
  1380. | ------------------------------ | ---------: | ---------: | ---------- | ------: | -------: | -: | --------------: | -------------------: |
  1381. | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 | 1832.03 ± 22.74 |
  1382. | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | tg32 | 122.27 ± 3.11 |
  1383. | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d4096 | 1370.87 ± 44.31 |
  1384. | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d4096 | 106.12 ± 1.90 |
  1385. | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d8192 | 1174.47 ± 19.20 |
  1386. | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d8192 | 100.58 ± 3.19 |
  1387. | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d16384 | 918.14 ± 13.86 |
  1388. | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d16384 | 92.41 ± 1.14 |
  1389. | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d32768 | 634.84 ± 8.48 |
  1390. | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d32768 | 76.55 ± 4.77 |
  1391.  
  1392. build: 3d4e86bb (6789)
  1393. === Sequential Benchmark Results ===
  1394. Model: qwen2.5-coder-7b-instruct-q8_0.gguf
  1395. Date: Sun Oct 19 18:48:06 EDT 2025
  1396. macOS Version: 26.0.1
  1397. Hardware: Mac16,6
  1398. Chip: Apple M4 Max
  1399. Total RAM: 128.00 GB
  1400. CPU Cores: 16
  1401. GPU Layers: 999
  1402. Flash Attention: 1
  1403.  
  1404. === Benchmark Parameters ===
  1405. Prompt Sizes: 2048
  1406. Generation Sizes: 32
  1407. Batch Size:
  1408.  
  1409. === Results ===
  1410. ggml_metal_library_init: using embedded metal library
  1411. ggml_metal_library_init: loaded in 0.006 sec
  1412. ggml_metal_device_init: GPU name: Apple M4 Max
  1413. ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
  1414. ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
  1415. ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
  1416. ggml_metal_device_init: simdgroup reduction = true
  1417. ggml_metal_device_init: simdgroup matrix mul. = true
  1418. ggml_metal_device_init: has unified memory = true
  1419. ggml_metal_device_init: has bfloat = true
  1420. ggml_metal_device_init: use residency sets = true
  1421. ggml_metal_device_init: use shared buffers = true
  1422. ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
  1423. | model | size | params | backend | threads | n_ubatch | fa | test | t/s |
  1424. | ------------------------------ | ---------: | ---------: | ---------- | ------: | -------: | -: | --------------: | -------------------: |
  1425. | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 | 752.70 ± 62.73 |
  1426. | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | tg32 | 50.23 ± 7.60 |
  1427. | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d4096 | 671.11 ± 14.22 |
  1428. | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d4096 | 56.64 ± 0.28 |
  1429. | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d8192 | 581.00 ± 20.79 |
  1430. | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d8192 | 53.73 ± 0.79 |
  1431. | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d16384 | 477.27 ± 12.32 |
  1432. | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d16384 | 49.62 ± 0.23 |
  1433. | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d32768 | 340.93 ± 1.76 |
  1434. | qwen2 7B Q8_0 | 7.54 GiB | 7.62 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d32768 | 42.93 ± 0.69 |
  1435.  
  1436. build: 3d4e86bb (6789)
  1437. === Sequential Benchmark Results ===
  1438. Model: qwen3-30b-a3b-q8_0.gguf
  1439. Date: Sun Oct 19 18:15:57 EDT 2025
  1440. macOS Version: 26.0.1
  1441. Hardware: Mac16,6
  1442. Chip: Apple M4 Max
  1443. Total RAM: 128.00 GB
  1444. CPU Cores: 16
  1445. GPU Layers: 999
  1446. Flash Attention: 1
  1447.  
  1448. === Benchmark Parameters ===
  1449. Prompt Sizes: 2048
  1450. Generation Sizes: 32
  1451. Batch Size:
  1452.  
  1453. === Results ===
  1454. ggml_metal_library_init: using embedded metal library
  1455. ggml_metal_library_init: loaded in 0.006 sec
  1456. ggml_metal_device_init: GPU name: Apple M4 Max
  1457. ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
  1458. ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
  1459. ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
  1460. ggml_metal_device_init: simdgroup reduction = true
  1461. ggml_metal_device_init: simdgroup matrix mul. = true
  1462. ggml_metal_device_init: has unified memory = true
  1463. ggml_metal_device_init: has bfloat = true
  1464. ggml_metal_device_init: use residency sets = true
  1465. ggml_metal_device_init: use shared buffers = true
  1466. ggml_metal_device_init: recommendedMaxWorkingSetSize = 115448.73 MB
  1467. | model | size | params | backend | threads | n_ubatch | fa | test | t/s |
  1468. | ------------------------------ | ---------: | ---------: | ---------- | ------: | -------: | -: | --------------: | -------------------: |
  1469. | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 | 1745.62 ± 29.06 |
  1470. | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | tg32 | 84.96 ± 1.19 |
  1471. | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d4096 | 918.50 ± 73.61 |
  1472. | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d4096 | 70.22 ± 1.13 |
  1473. | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d8192 | 698.42 ± 21.62 |
  1474. | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d8192 | 64.64 ± 0.63 |
  1475. | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d16384 | 449.50 ± 16.10 |
  1476. | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d16384 | 53.08 ± 1.34 |
  1477. | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | pp2048 @ d32768 | 253.05 ± 9.51 |
  1478. | qwen3moe 30B.A3B Q8_0 | 30.25 GiB | 30.53 B | Metal,BLAS | 12 | 2048 | 1 | tg32 @ d32768 | 40.53 ± 2.98 |
  1479.  
  1480. build: 3d4e86bb (6789)
  1481.  
Advertisement
Add Comment
Please, Sign In to add comment