Guest User

Qwen 3.8 chat template analysis

a guest
Aug 14th, 2026
230
0
Never
Not a member of Pastebin yet? Sign Up, it unlocks many cool features!
text 29.45 KB | Software | 0 0
  1. # 1. Reasoning support
  2.  
  3. ## Reasoning effort comparison
  4.  
  5. | Behavior | Original | Modified | Assessment |
  6. |---|---|---|---|
  7. | Default effort | `xhigh` | `xhigh` | Same |
  8. | `xhigh` | Accepted | Accepted | Same |
  9. | `high` | Raises error | Aliased to `xhigh` | Improvement |
  10. | `medium` | Thinking enabled, no extra instruction | Same | Same |
  11. | `low` | Low-effort instruction | Same | Same |
  12. | Invalid value | Raises error | Silently becomes `xhigh` | Regression |
  13. | Actual token budget control | No | No | Both only prompt the model |
  14.  
  15. The modified normalization:
  16.  
  17. ```jinja
  18. {%- else %}
  19. {%- set _reasoning_effort = 'xhigh' %}
  20. {%- endif %}
  21. ```
  22.  
  23. means values such as:
  24.  
  25. - `none`
  26. - `minimal`
  27. - `very_high`
  28. - `HIGH`
  29. - a typo such as `medum`
  30.  
  31. all silently select maximum reasoning.
  32.  
  33. That is especially undesirable because an attempt to disable or reduce reasoning can accidentally enable the most expensive mode.
  34.  
  35. ### Recommended behavior
  36.  
  37. Keep the useful `high` alias, but reject unknown values:
  38.  
  39. ```jinja
  40. {%- if _effort_raw in ('high', 'xhigh') %}
  41. {%- set _reasoning_effort = 'xhigh' %}
  42. {%- elif _effort_raw in ('medium', 'low') %}
  43. {%- set _reasoning_effort = _effort_raw %}
  44. {%- else %}
  45. {{- raise_exception(
  46. 'Unexpected reasoning effort ' ~ _effort_raw ~
  47. '. Supported values are high/xhigh, medium, and low.'
  48. ) }}
  49. {%- endif %}
  50. ```
  51.  
  52. If `none` is intended to disable thinking, it should be implemented explicitly rather than falling back to `xhigh`.
  53.  
  54. ## Reasoning effort is only a prompt instruction
  55.  
  56. In both templates, `reasoning_effort` only controls system text such as:
  57.  
  58. > Reasoning effort is set to xhigh...
  59.  
  60. It does not directly control:
  61.  
  62. - maximum reasoning tokens,
  63. - sampling parameters,
  64. - a server-side reasoning budget,
  65. - inference-engine reasoning settings.
  66.  
  67. Thus `medium` simply means “thinking mode is on, but no special effort instruction is added.” Whether that behaves as medium depends entirely on the model.
  68.  
  69. ---
  70.  
  71. ## `enable_thinking` handling
  72.  
  73. The modified template is more internally consistent because it derives everything from:
  74.  
  75. ```jinja
  76. ns_state.thinking
  77. ```
  78.  
  79. The original uses exact Boolean tests in some places:
  80.  
  81. ```jinja
  82. enable_thinking is true
  83. enable_thinking is false
  84. ```
  85.  
  86. This can behave inconsistently for non-Boolean values such as `1`, `None`, or `"false"`.
  87.  
  88. The modified version uses truthiness instead. That is generally better, although callers should still pass actual Booleans.
  89.  
  90. ---
  91.  
  92. ## Think-control token bugs
  93.  
  94. The modified template adds:
  95.  
  96. ```text
  97. <|think_off|>
  98. <|think_on|>
  99. ```
  100.  
  101. This is useful, but the implementation has several correctness problems.
  102.  
  103. ### 1. `off` always wins within one string
  104.  
  105. The scan uses:
  106.  
  107. ```jinja
  108. {%- if '<|think_off|>' in msg.content %}
  109. ...
  110. {%- elif '<|think_on|>' in msg.content %}
  111. ```
  112.  
  113. Therefore this content:
  114.  
  115. ```text
  116. <|think_off|> ... <|think_on|>
  117. ```
  118.  
  119. leaves thinking disabled, even though the last directive is `think_on`.
  120.  
  121. It does not process directives in occurrence order.
  122.  
  123. ### 2. One token may remain in the rendered prompt
  124.  
  125. Removal also uses `if`/`elif`:
  126.  
  127. ```jinja
  128. {%- if '<|think_off|>' in content %}
  129. remove think_off
  130. {%- elif '<|think_on|>' in content %}
  131. remove think_on
  132. ```
  133.  
  134. If both are present, only one type is removed. For example:
  135.  
  136. ```text
  137. <|think_off|>foo<|think_on|>
  138. ```
  139.  
  140. can leave `<|think_on|>` visible to the model.
  141.  
  142. Both markers should always be removed independently:
  143.  
  144. ```jinja
  145. {%- set content = content
  146. | replace('<|think_off|>', '')
  147. | replace('<|think_on|>', '')
  148. | trim
  149. %}
  150. ```
  151.  
  152. State selection still needs separate occurrence-order logic.
  153.  
  154. ### 3. Users can override system/developer thinking policy
  155.  
  156. The scan includes:
  157.  
  158. ```jinja
  159. msg.role == 'system' or
  160. msg.role == 'developer' or
  161. msg.role == 'user'
  162. ```
  163.  
  164. Consequently, a later user message can override a system-level `<|think_off|>` with `<|think_on|>`.
  165.  
  166. That may be intentional, but if these are privileged control tokens, it is a hierarchy violation. Consider scanning only system/developer messages, or explicitly document that users are allowed to control thinking.
  167.  
  168. ### 4. Literal quoted tokens also change behavior
  169.  
  170. Any historical user message containing the literal text—such as a discussion explaining `<|think_off|>`—changes generation behavior. There is no escaping mechanism.
  171.  
  172. ### 5. List-of-string content is inconsistent
  173.  
  174. The new renderer permits non-mapping list items by stringifying them:
  175.  
  176. ```jinja
  177. {{- item | string }}
  178. ```
  179.  
  180. But the preliminary thinking scan only checks mapping items with a `text` field. A list item equal to `"<|think_off|>"` can be removed later without ever changing `ns_state.thinking`.
  181.  
  182. ---
  183.  
  184. ## Historical thinking reconstruction regression
  185.  
  186. The original always emits a thinking block when preservation applies, even when reasoning is empty:
  187.  
  188. ```text
  189. <think>
  190.  
  191. </think>
  192.  
  193. <tool_call>...
  194. ```
  195.  
  196. The modified version requires nonempty reasoning:
  197.  
  198. ```jinja
  199. {%- if (_preserve_thinking or loop.index0 > ns.last_query_index)
  200. and reasoning_content %}
  201. ```
  202.  
  203. Thus an assistant tool call with no stored reasoning becomes:
  204.  
  205. ```text
  206. <tool_call>...
  207. ```
  208.  
  209. This is inconsistent with the generation prompt when thinking is disabled, which emits:
  210.  
  211. ```text
  212. <think>
  213.  
  214. </think>
  215.  
  216. ```
  217.  
  218. This matters during multi-step tool use:
  219.  
  220. 1. The generation prompt starts with an empty thinking block.
  221. 2. The model emits a tool call.
  222. 3. The application stores structured `tool_calls`, but no `reasoning_content`.
  223. 4. On the next render, the modified template reconstructs the assistant turn without the empty thinking block.
  224.  
  225. The original template reconstructs it consistently.
  226.  
  227. Whether the empty block is strictly required depends on the checkpoint, but this is a concrete format change and should be tested. To preserve original behavior, remove `and reasoning_content`.
  228.  
  229. ---
  230.  
  231. ## Reasoning extraction improvements
  232.  
  233. The modified template does improve compatibility by accepting:
  234.  
  235. - `message.reasoning_content`,
  236. - `message.thinking`,
  237. - non-string reasoning values,
  238. - some inline `<think>...</think>` content.
  239.  
  240. That is a useful fix.
  241.  
  242. However, the fallback parser is heuristic:
  243.  
  244. - It misses forms such as `<think>x</think>` without a newline before the closing tag.
  245. - With multiple closing tags, `split(...)[-1]` may discard intermediate content.
  246. - A literal `</think>` appearing in an answer can be mistaken for a reasoning delimiter.
  247. - If `reasoning_content` exists but is empty, inline reasoning is not parsed.
  248.  
  249. ---
  250.  
  251. # 2. Tool-call support
  252.  
  253. ## XML arguments: original limitation and modified partial fix
  254.  
  255. ### Original behavior
  256.  
  257. The original assumes `tool_call.arguments` is a mapping:
  258.  
  259. ```jinja
  260. {%- for args_name, args_value in tool_call.arguments|items %}
  261. ```
  262.  
  263. This works for:
  264.  
  265. ```json
  266. {
  267. "name": "weather",
  268. "arguments": {
  269. "city": "Paris"
  270. }
  271. }
  272. ```
  273.  
  274. But many OpenAI-compatible APIs store arguments as a JSON string:
  275.  
  276. ```json
  277. {
  278. "name": "weather",
  279. "arguments": "{\"city\":\"Paris\"}"
  280. }
  281. ```
  282.  
  283. Applying `|items` to that string normally raises a template error.
  284.  
  285. This is a real problem in the original.
  286.  
  287. ### Modified XML behavior
  288.  
  289. The modified template avoids the exception, but does not actually convert the JSON string into XML parameters:
  290.  
  291. ```jinja
  292. {%- elif tc.arguments is string and tc.arguments %}
  293. {{- tc.arguments }}
  294. ```
  295.  
  296. It produces:
  297.  
  298. ```xml
  299. <tool_call>
  300. <function=weather>
  301. {"city":"Paris"}</function>
  302. </tool_call>
  303. ```
  304.  
  305. That does not follow the instructed XML dialect:
  306.  
  307. ```xml
  308. <parameter=city>
  309. Paris
  310. </parameter>
  311. ```
  312.  
  313. So the new XML branch changes a hard failure into malformed or parser-incompatible output.
  314.  
  315. ### Recommended solution
  316.  
  317. For XML mode, either:
  318.  
  319. 1. Require arguments to be normalized to a mapping by the host application, or
  320. 2. Parse JSON-string arguments before applying the template, or
  321. 3. Raise a clear exception instructing the caller to use JSON mode.
  322.  
  323. Do not silently place raw JSON inside `<function>`.
  324.  
  325. ---
  326.  
  327. ## JSON tool mode
  328.  
  329. The new JSON mode can correctly handle both mapping and JSON-string arguments:
  330.  
  331. ```xml
  332. <tool_call>
  333. {"name":"weather","arguments":{"city":"Paris"}}
  334. </tool_call>
  335. ```
  336.  
  337. That is a meaningful improvement.
  338.  
  339. However, this only works end-to-end if all components agree on the format:
  340.  
  341. - the model/checkpoint,
  342. - the chat template,
  343. - the inference server,
  344. - the tool-call parser,
  345. - the API response adapter.
  346.  
  347. Changing the prompt format does not automatically teach a server’s Qwen tool parser to recognize it. If the server expects:
  348.  
  349. ```xml
  350. <function=...>
  351. <parameter=...>
  352. ```
  353.  
  354. then JSON mode may generate apparently correct text but no structured `tool_calls` in the API response.
  355.  
  356. Likewise, a checkpoint primarily trained on the XML dialect may be less reliable in JSON mode, even though the prompt contains an example.
  357.  
  358. `tool_call_format` should also be validated. Currently any value other than `"json"` silently selects XML.
  359.  
  360. ---
  361.  
  362. ## Scalar XML argument regression
  363.  
  364. The original serializes every non-string value using JSON:
  365.  
  366. ```jinja
  367. args_value | tojson
  368. ```
  369.  
  370. Therefore:
  371.  
  372. - `true` remains `true`
  373. - `false` remains `false`
  374. - `null` remains `null`
  375.  
  376. The modified version only uses JSON for mappings and sequences:
  377.  
  378. ```jinja
  379. {%- if args_value is mapping or
  380. (args_value is sequence and args_value is not string) %}
  381. {{ args_value | tojson }}
  382. {%- else %}
  383. {{ args_value | string }}
  384. {%- endif %}
  385. ```
  386.  
  387. This can render:
  388.  
  389. ```text
  390. True
  391. False
  392. None
  393. ```
  394.  
  395. instead of:
  396.  
  397. ```text
  398. true
  399. false
  400. null
  401. ```
  402.  
  403. That is a regression, particularly if the tool-call parser interprets parameter values as JSON-like data.
  404.  
  405. A better rule is the original one:
  406.  
  407. ```jinja
  408. {%- if args_value is string %}
  409. {%- set _av = args_value %}
  410. {%- else %}
  411. {%- set _av = args_value | tojson %}
  412. {%- endif %}
  413. ```
  414.  
  415. This also handles nested values correctly.
  416.  
  417. ---
  418.  
  419. ## Tool instructions
  420.  
  421. The modified instructions improve parseability by telling the model to put pre-call reasoning inside `<think>` rather than as arbitrary text before the tool call. This is better than the original statement:
  422.  
  423. > You may provide optional reasoning ... in natural language BEFORE the function call
  424.  
  425. Free text outside a reasoning block can interfere with strict tool parsers.
  426.  
  427. However, the new wording has two issues:
  428.  
  429. 1. The same `<think>` instructions are included when thinking is disabled.
  430. 2. “ALL explanation and reasoning MUST be placed strictly inside `<think>`” is broad enough that the model could interpret it as applying to final answers as well.
  431.  
  432. It would be better to scope the instruction:
  433.  
  434. > When making a tool call, any planning before the call must remain inside the thinking block. Do not emit conversational text between the thinking block and the tool call.
  435.  
  436. Also, the serializer still permits nonempty `message.content` before structured tool calls, even though the new prompt discourages it.
  437.  
  438. ---
  439.  
  440. ## Multiple tool calls and responses
  441.  
  442. Both templates support:
  443.  
  444. - multiple tool calls in one assistant message,
  445. - multiple consecutive tool results,
  446. - grouping tool results inside one synthetic user turn.
  447.  
  448. The modified formatting adds slightly more separation between tool-call blocks, but remains structurally valid.
  449.  
  450. Neither template includes tool-call IDs in the serialized tool responses. Results are effectively associated by position. That is inherited behavior and may be sufficient for the intended Qwen format, but it is weaker for parallel calls where IDs matter.
  451.  
  452. ---
  453.  
  454. ## Tool-response error detection is unsafe
  455.  
  456. The modified template automatically classifies short tool results as failures and injects warnings.
  457.  
  458. For example, this valid result:
  459.  
  460. ```json
  461. {"error": null, "data": [1, 2, 3]}
  462. ```
  463.  
  464. matches:
  465.  
  466. ```jinja
  467. '"error":' in _content_head
  468. ```
  469.  
  470. and receives:
  471.  
  472. ```text
  473. ⚠️ SYSTEM WARNING: The previous tool call returned an error...
  474. ```
  475.  
  476. Other false positives include:
  477.  
  478. ```text
  479. error: 0
  480. no error: success
  481. ```
  482.  
  483. Meanwhile, real errors can be missed when:
  484.  
  485. - they are over 500 characters,
  486. - the error marker occurs after the first 80 characters,
  487. - output contains `$ `,
  488. - output contains `took `.
  489.  
  490. The injected warning is also not a real system message. It is text inside a tool response wrapped as a user turn, despite saying `SYSTEM WARNING`.
  491.  
  492. This behavior should be:
  493.  
  494. - removed,
  495. - made opt-in, or
  496. - based on structured metadata such as `message.is_error`.
  497.  
  498. It should not infer success/failure from arbitrary response text by default.
  499.  
  500. ---
  501.  
  502. ## Truncation support
  503.  
  504. Optional truncation is useful and defaults to disabled, so it does not affect baseline behavior.
  505.  
  506. There are inconsistencies:
  507.  
  508. - `max_tool_arg_chars` only applies to mapped XML arguments.
  509. - It does not apply to raw string arguments.
  510. - It does not apply in JSON mode.
  511. - Truncating a nested JSON value can make the value invalid JSON.
  512. - Tool-response truncation can remove essential result data.
  513.  
  514. These are acceptable only if documented as lossy context controls.
  515.  
  516. ---
  517.  
  518. # 3. Query and reasoning-preservation logic
  519.  
  520. The original raises when it cannot find a real user query:
  521.  
  522. ```jinja
  523. {{- raise_exception('No user query found in messages.') }}
  524. ```
  525.  
  526. The modified version instead uses:
  527.  
  528. ```jinja
  529. {%- if _last_idx > 50 %}
  530. {%- set ns.last_query_index = _last_idx %}
  531. {%- else %}
  532. {%- set ns.last_query_index = 0 %}
  533. {%- endif %}
  534. ```
  535.  
  536. This is difficult to justify:
  537.  
  538. - For a 50-message history, one preservation behavior is selected.
  539. - For a 52-message history, the opposite behavior may be selected.
  540. - A conversation with no user query is silently accepted.
  541. - The result of `preserve_thinking=false` becomes dependent on an arbitrary history length.
  542.  
  543. This should be replaced with a deterministic policy:
  544.  
  545. - retain the original exception, or
  546. - explicitly treat the entire history as a continuation, e.g. use `-1`, or
  547. - explicitly treat all history as prior context.
  548.  
  549. The `> 50` heuristic is a correctness bug.
  550.  
  551. ---
  552.  
  553. # 4. System and developer messages
  554.  
  555. ## Improvements
  556.  
  557. The original only supports `system` as the first message. A `developer` role eventually causes:
  558.  
  559. ```text
  560. Unexpected message role.
  561. ```
  562.  
  563. The modified template supports `developer` and maps it to Qwen’s `system` role. That is useful for OpenAI-style message inputs.
  564.  
  565. ## Regressions
  566.  
  567. The original enforces:
  568.  
  569. ```text
  570. System message must be at the beginning.
  571. ```
  572.  
  573. The modified template allows later system/developer messages and emits them wherever they occur.
  574.  
  575. It also converts any unknown role into:
  576.  
  577. ```text
  578. [user]
  579. [role]: content
  580. ```
  581.  
  582. instead of raising.
  583.  
  584. This is more permissive, but less safe:
  585.  
  586. - malformed message histories are hidden rather than detected,
  587. - a later system message can appear after user turns,
  588. - unknown observation/function roles may become user instructions,
  589. - model-role semantics are changed silently.
  590.  
  591. A better policy is:
  592.  
  593. 1. Merge all leading `system`/`developer` messages into the initial system prompt.
  594. 2. Reject system/developer messages after the first non-system message.
  595. 3. Reject unknown roles unless explicit compatibility conversion is enabled.
  596.  
  597. ---
  598.  
  599. # 5. Content and vision handling
  600.  
  601. The modified renderer is safer for normal dictionary content items because it checks:
  602.  
  603. ```jinja
  604. item is mapping
  605. ```
  606.  
  607. before testing keys. The original can fail strangely on non-mapping list items.
  608.  
  609. The modified version also explicitly defaults `add_vision_id`, which is cleaner.
  610.  
  611. However, it now stringifies unsupported list items instead of raising:
  612.  
  613. ```jinja
  614. {{- item | string }}
  615. ```
  616.  
  617. That may hide malformed multimodal inputs. It also stops supporting object-like items with attributes such as `.type` unless they are mappings.
  618.  
  619. Both templates correctly reject image/video content in system messages; the modified version extends that restriction to developer messages.
  620.  
  621. ---
  622.  
  623. # 6. Practical compatibility matrix
  624.  
  625. | Input / feature | Original | Modified XML | Modified JSON |
  626. |---|---:|---:|---:|
  627. | Basic text chat | Works | Works | Works |
  628. | `reasoning_effort=xhigh` | Works | Works | Works |
  629. | `reasoning_effort=high` | Error | Works as xhigh | Works as xhigh |
  630. | Invalid effort | Clear error | Silently xhigh | Silently xhigh |
  631. | Thinking disabled | Works | Works, but tool prompt still mentions thinking | Same |
  632. | `reasoning_content` string | Works | Works | Works |
  633. | `thinking` alias | Ignored | Works | Works |
  634. | Inline `<think>` extraction | No | Partial support | Partial support |
  635. | Mapping tool arguments | Works | Mostly works | Works |
  636. | Boolean/null XML args | Correct JSON spelling | Python spelling regression | Correct JSON |
  637. | JSON-string arguments | Template error | Malformed XML dialect | Works if valid JSON |
  638. | Multiple tool calls | Works | Works | Works |
  639. | Developer role | Error | Supported as system | Supported as system |
  640. | Noninitial system role | Error | Silently accepted | Silently accepted |
  641. | No real user query | Clear error | Arbitrary fallback | Arbitrary fallback |
  642. | External parser compatibility | Native format only | Usually closest to base | Must be verified |
  643. | Valid `{"error": null}` result | Unmodified | Falsely warned | Falsely warned |
  644.  
  645. ---
  646.  
  647. # 7. Recommended changes before deployment
  648.  
  649. At minimum:
  650.  
  651. 1. **Validate reasoning effort** while retaining `high -> xhigh`.
  652. 2. **Validate `tool_call_format`** instead of silently defaulting unknown values to XML.
  653. 3. **Require mapping arguments in XML mode**, or parse JSON strings before templating.
  654. 4. **Serialize every non-string XML argument with `tojson`**.
  655. 5. **Restore consistent empty thinking-block reconstruction**, or deliberately test that the target checkpoint does not need it.
  656. 6. **Remove the `> 50` no-query fallback**.
  657. 7. **Remove both think-control markers unconditionally** and define proper precedence.
  658. 8. **Decide whether users are allowed to override system thinking policy**.
  659. 9. **Remove or opt in to heuristic tool-error warnings**.
  660. 10. **Merge only leading system/developer messages and reject later ones**.
  661. 11. **Verify JSON tool mode against the actual inference server’s tool parser**.
  662.  
  663. ## Overall assessment
  664.  
  665. - The **original** is narrower and stricter. It has weaker compatibility, especially for developer roles and JSON-string tool arguments, but its failure modes are usually obvious.
  666. - The **modified** version adds valuable functionality, particularly JSON tool formatting and broader reasoning extraction, but it replaces several explicit errors with silent or malformed behavior.
  667. - For production use, the modified version should be treated as a feature branch requiring validation—not as a drop-in corrected version of the original.
  668.  
  669. ---
  670.  
  671. ## When the original works correctly
  672.  
  673. The original expects reasoning to be stored separately:
  674.  
  675. ```json
  676. {
  677. "role": "assistant",
  678. "reasoning_content": "I should inspect the weather.",
  679. "content": "The weather is sunny."
  680. }
  681. ```
  682.  
  683. It renders that correctly:
  684.  
  685. ```text
  686. <think>
  687. I should inspect the weather.
  688. </think>
  689.  
  690. The weather is sunny.
  691. ```
  692.  
  693. There is no blank block before the actual reasoning.
  694.  
  695. ## When the original breaks
  696.  
  697. Some clients store the raw generated completion entirely in `content`:
  698.  
  699. ```json
  700. {
  701. "role": "assistant",
  702. "content": "I should inspect the weather.\n</think>\n\nThe weather is sunny."
  703. }
  704. ```
  705.  
  706. This is plausible because the generation prompt already supplied the opening:
  707.  
  708. ```text
  709. <think>
  710. ```
  711.  
  712. and the generated text contains the reasoning followed by `</think>`.
  713.  
  714. The original does not parse reasoning out of `content`. Since `reasoning_content` is absent, it emits an empty block and then appends the raw content:
  715.  
  716. ```text
  717. <think>
  718.  
  719. </think>
  720.  
  721. I should inspect the weather.
  722. </think>
  723.  
  724. The weather is sunny.
  725. ```
  726.  
  727. That is malformed history: it has an empty thinking block followed by reasoning and an unmatched closing tag.
  728.  
  729. ## What the modified template fixes
  730.  
  731. The modified template tries to detect a closing `</think>` inside `content`, split out the preceding reasoning, and reconstruct:
  732.  
  733. ```text
  734. <think>
  735. I should inspect the weather.
  736. </think>
  737.  
  738. The weather is sunny.
  739. ```
  740.  
  741. That is a real compatibility improvement for clients that store raw reasoning in `content`.
  742.  
  743. ## Important qualifications
  744.  
  745. Calling this “the official template poisons history” is too broad:
  746.  
  747. - The original works correctly when `reasoning_content` is populated as expected.
  748. - Empty `<think></think>` blocks are not inherently poisonous. They can be the correct representation of a no-thinking assistant turn.
  749. - The modified fallback parser is heuristic and does not recognize every possible formatting variation.
  750. - If `reasoning_content` exists but is an empty string, the modified template will not fall back to parsing reasoning from `content`.
  751. - The modified template introduces the opposite inconsistency: it suppresses historical empty thinking blocks because of:
  752.  
  753. ```jinja
  754. and reasoning_content
  755. ```
  756.  
  757. This can make a reconstructed no-thinking turn differ from the turn that was originally generated.
  758.  
  759. For example, with thinking disabled, the generation prompt includes:
  760.  
  761. ```text
  762. <think>
  763.  
  764. </think>
  765.  
  766. ```
  767.  
  768. But if the stored assistant message only contains the answer, the modified template reconstructs it without that empty block. The original reconstructs it with the block.
  769.  
  770. ### Better documentation wording
  771.  
  772. A more accurate claim would be:
  773.  
  774. > Recovers reasoning from assistant `content` when clients do not store it separately in `reasoning_content`, preventing duplicate empty thinking blocks and unmatched closing tags in that representation.
  775.  
  776. ---
  777.  
  778. # 2. “Tool calling crashes”
  779.  
  780. > If your client passes arguments as JSON strings—the standard OpenAI API format—the official template crashes.
  781.  
  782. This is substantially correct.
  783.  
  784. ## Original behavior
  785.  
  786. The original assumes `arguments` is a mapping:
  787.  
  788. ```jinja
  789. {%- for args_name, args_value in tool_call.arguments|items %}
  790. ```
  791.  
  792. It expects:
  793.  
  794. ```json
  795. {
  796. "name": "get_weather",
  797. "arguments": {
  798. "city": "Paris"
  799. }
  800. }
  801. ```
  802.  
  803. But OpenAI-style API tool calls commonly use a JSON-encoded string:
  804.  
  805. ```json
  806. {
  807. "name": "get_weather",
  808. "arguments": "{\"city\":\"Paris\"}"
  809. }
  810. ```
  811.  
  812. On modern Jinja, applying `|items` to a string normally raises an error such as:
  813.  
  814. ```text
  815. TypeError: Can only get item pairs from a mapping.
  816. ```
  817.  
  818. Therefore, yes: the original can crash if the adapter passes the wire-format JSON string directly to the template.
  819.  
  820. There is a nuance: many inference servers or client adapters parse the JSON string into a dictionary before invoking the template. In those environments, the original works. The original template assumes that normalization has already occurred.
  821.  
  822. ## Does the modified template fix it?
  823.  
  824. ### In JSON tool-call mode: yes, mostly
  825.  
  826. With:
  827.  
  828. ```text
  829. tool_call_format = "json"
  830. ```
  831.  
  832. the modified template accepts either:
  833.  
  834. - a mapping, serialized with `tojson`, or
  835. - an existing JSON string, inserted as the `arguments` value.
  836.  
  837. It can produce:
  838.  
  839. ```xml
  840. <tool_call>
  841. {"name": "get_weather", "arguments": {"city":"Paris"}}
  842. </tool_call>
  843. ```
  844.  
  845. That is a valid fix, assuming:
  846.  
  847. 1. The argument string is valid JSON.
  848. 2. The model supports this tool-call dialect.
  849. 3. The inference server’s tool parser recognizes JSON inside `<tool_call>`.
  850.  
  851. The template does not validate the raw JSON string.
  852.  
  853. ### In default XML mode: it avoids the crash but does not serialize correctly
  854.  
  855. The modified template defaults to:
  856.  
  857. ```jinja
  858. _tool_format = 'xml'
  859. ```
  860.  
  861. For string arguments, it emits the string directly:
  862.  
  863. ```jinja
  864. {%- elif tc.arguments is string and tc.arguments %}
  865. {{- tc.arguments }}
  866. ```
  867.  
  868. The result is:
  869.  
  870. ```xml
  871. <tool_call>
  872. <function=get_weather>
  873. {"city":"Paris"}</function>
  874. </tool_call>
  875. ```
  876.  
  877. But the documented XML format requires:
  878.  
  879. ```xml
  880. <tool_call>
  881. <function=get_weather>
  882. <parameter=city>
  883. Paris
  884. </parameter>
  885. </function>
  886. </tool_call>
  887. ```
  888.  
  889. So in default XML mode, the modified template changes the failure from:
  890.  
  891. - template crash
  892.  
  893. to:
  894.  
  895. - nonconforming tool-call history that the model or parser may not understand.
  896.  
  897. That is not a complete fix.
  898.  
  899. ## Additional XML regression
  900.  
  901. The original JSON-serializes all non-string values. The modified XML branch stringifies scalar values. This can change:
  902.  
  903. ```json
  904. true
  905. false
  906. null
  907. ```
  908.  
  909. into Python-style:
  910.  
  911. ```text
  912. True
  913. False
  914. None
  915. ```
  916.  
  917. That may be incorrectly parsed by downstream tool-call parsers.
  918.  
  919. The modified XML branch should use:
  920.  
  921. ```jinja
  922. {%- if args_value is string %}
  923. {%- set _av = args_value %}
  924. {%- else %}
  925. {%- set _av = args_value | tojson %}
  926. {%- endif %}
  927. ```
  928.  
  929. ## Better documentation wording
  930.  
  931. > Adds JSON-string argument support in JSON tool-call mode. The original template expects arguments to be pre-parsed into a mapping and may raise a Jinja error otherwise.
  932.  
  933. It should not claim complete support for JSON-string arguments in the default XML mode.
  934.  
  935. ---
  936.  
  937. # 3. “100% KV Cache hits”
  938.  
  939. > Keeps past thoughts intact by default so your prefix cache stays warm across turns.
  940.  
  941. This is the weakest claim.
  942.  
  943. ## The original already preserves thoughts by default
  944.  
  945. The original condition is:
  946.  
  947. ```jinja
  948. preserve_thinking is undefined
  949. or preserve_thinking is true
  950. or loop.index0 > ns.last_query_index
  951. ```
  952.  
  953. If `preserve_thinking` is not specified, historical reasoning is preserved.
  954.  
  955. The modified template also defaults preservation to true:
  956.  
  957. ```jinja
  958. {%- set _preserve_thinking = true %}
  959. ```
  960.  
  961. Therefore, “keeps past thoughts intact by default” is not a new behavior. The original already does it.
  962.  
  963. The modified version adds compatibility with a `preserve_reasoning` option and can recover reasoning embedded in `content`, but preservation itself was already present.
  964.  
  965. ## Preserving reasoning does not guarantee cache hits
  966.  
  967. Prefix caching requires the newly rendered prompt to have an exact token prefix matching a previously processed sequence.
  968.  
  969. That depends on much more than retaining reasoning:
  970.  
  971. - identical system prompt,
  972. - identical tool definitions and ordering,
  973. - identical whitespace and delimiters,
  974. - identical reasoning serialization,
  975. - identical thinking settings,
  976. - no changed truncation,
  977. - no changed template version,
  978. - support for prefix caching in the inference server,
  979. - the cached entry still being resident.
  980.  
  981. A template cannot promise “100% KV cache hits.”
  982.  
  983. At best, it can improve the chance that the previously generated prefix is reconstructed exactly.
  984.  
  985. ## Cases where the modified template may improve caching
  986.  
  987. If raw reasoning is stored in `content`, the original reconstructs a malformed and different prefix:
  988.  
  989. ```text
  990. <think>
  991.  
  992. </think>
  993.  
  994. actual reasoning
  995. </think>
  996. ```
  997.  
  998. The modified parser may reconstruct the original sequence more accurately. That can improve cache reuse.
  999.  
  1000. ## Cases where the modified template can reduce cache reuse
  1001.  
  1002. ### 1. It removes empty historical thinking blocks
  1003.  
  1004. As noted above, the modified template only emits historical thinking when `reasoning_content` is nonempty:
  1005.  
  1006. ```jinja
  1007. and reasoning_content
  1008. ```
  1009.  
  1010. If the previous generation used:
  1011.  
  1012. ```text
  1013. <think>
  1014.  
  1015. </think>
  1016.  
  1017. Answer
  1018. ```
  1019.  
  1020. but the next render reconstructs:
  1021.  
  1022. ```text
  1023. Answer
  1024. ```
  1025.  
  1026. the prefix differs at the beginning of that assistant turn. The cache cannot be reused past that point.
  1027.  
  1028. ### 2. A new user think-control token can change the initial system prompt
  1029.  
  1030. The modified template scans all messages for:
  1031.  
  1032. ```text
  1033. <|think_off|>
  1034. <|think_on|>
  1035. ```
  1036.  
  1037. A new user message can change `ns_state.thinking`, which changes whether the reasoning-effort instruction appears in the initial system message.
  1038.  
  1039. For example, adding `<|think_off|>` to the latest user message can remove the initial xhigh instruction on the next render. That changes the prompt near its beginning and can invalidate essentially the whole prefix cache.
  1040.  
  1041. ### 3. Reformatting can change tokens
  1042.  
  1043. The modified template trims and reconstructs reasoning/content. Even semantically equivalent changes in newlines or spaces can prevent exact token-prefix matching.
  1044.  
  1045. ## What “100%” could reasonably mean
  1046.  
  1047. It might mean:
  1048.  
  1049. > All unchanged tokens from the preceding conversation are eligible for prefix-cache reuse.
  1050.  
  1051. Even then, it is only true if the template reproduces those tokens exactly. Newly added user and generation-prompt tokens obviously are not already cached, so the entire new request cannot generally be a 100% cache hit.
  1052.  
  1053. ## Better documentation wording
  1054.  
  1055. > Preserves historical reasoning by default and can reconstruct reasoning embedded in assistant content, helping maintain a stable prefix for KV-cache reuse. Actual cache-hit rates depend on exact token stability and inference-server support.
  1056.  
  1057. ---
  1058.  
  1059. # Overall assessment of the claims
  1060.  
  1061. ### Claim 1: blank thinking blocks
  1062.  
  1063. **Based on a real issue, but missing the condition.**
  1064. The original fails when raw reasoning remains in `content`; it works when `reasoning_content` is correctly populated. The modified template improves compatibility but introduces empty-think reconstruction inconsistencies.
  1065.  
  1066. ### Claim 2: JSON-string tool arguments
  1067.  
  1068. **Correct about the original crash.**
  1069. However, the modified version properly supports such arguments only in JSON tool mode. Its default XML mode avoids the exception but emits the arguments in the wrong dialect.
  1070.  
  1071. ### Claim 3: 100% KV cache hits
  1072.  
  1073. **Mostly marketing.**
  1074. The original already preserves reasoning by default, and no chat template can guarantee a cache-hit percentage. The modified template can improve prefix stability in one message-storage scenario, but can also reduce it in others.
Advertisement
Add Comment
Please, Sign In to add comment