Not a member of Pastebin yet?
Sign Up,
it unlocks many cool features!
- # 1. Reasoning support
- ## Reasoning effort comparison
- | Behavior | Original | Modified | Assessment |
- |---|---|---|---|
- | Default effort | `xhigh` | `xhigh` | Same |
- | `xhigh` | Accepted | Accepted | Same |
- | `high` | Raises error | Aliased to `xhigh` | Improvement |
- | `medium` | Thinking enabled, no extra instruction | Same | Same |
- | `low` | Low-effort instruction | Same | Same |
- | Invalid value | Raises error | Silently becomes `xhigh` | Regression |
- | Actual token budget control | No | No | Both only prompt the model |
- The modified normalization:
- ```jinja
- {%- else %}
- {%- set _reasoning_effort = 'xhigh' %}
- {%- endif %}
- ```
- means values such as:
- - `none`
- - `minimal`
- - `very_high`
- - `HIGH`
- - a typo such as `medum`
- all silently select maximum reasoning.
- That is especially undesirable because an attempt to disable or reduce reasoning can accidentally enable the most expensive mode.
- ### Recommended behavior
- Keep the useful `high` alias, but reject unknown values:
- ```jinja
- {%- if _effort_raw in ('high', 'xhigh') %}
- {%- set _reasoning_effort = 'xhigh' %}
- {%- elif _effort_raw in ('medium', 'low') %}
- {%- set _reasoning_effort = _effort_raw %}
- {%- else %}
- {{- raise_exception(
- 'Unexpected reasoning effort ' ~ _effort_raw ~
- '. Supported values are high/xhigh, medium, and low.'
- ) }}
- {%- endif %}
- ```
- If `none` is intended to disable thinking, it should be implemented explicitly rather than falling back to `xhigh`.
- ## Reasoning effort is only a prompt instruction
- In both templates, `reasoning_effort` only controls system text such as:
- > Reasoning effort is set to xhigh...
- It does not directly control:
- - maximum reasoning tokens,
- - sampling parameters,
- - a server-side reasoning budget,
- - inference-engine reasoning settings.
- Thus `medium` simply means “thinking mode is on, but no special effort instruction is added.” Whether that behaves as medium depends entirely on the model.
- ---
- ## `enable_thinking` handling
- The modified template is more internally consistent because it derives everything from:
- ```jinja
- ns_state.thinking
- ```
- The original uses exact Boolean tests in some places:
- ```jinja
- enable_thinking is true
- enable_thinking is false
- ```
- This can behave inconsistently for non-Boolean values such as `1`, `None`, or `"false"`.
- The modified version uses truthiness instead. That is generally better, although callers should still pass actual Booleans.
- ---
- ## Think-control token bugs
- The modified template adds:
- ```text
- <|think_off|>
- <|think_on|>
- ```
- This is useful, but the implementation has several correctness problems.
- ### 1. `off` always wins within one string
- The scan uses:
- ```jinja
- {%- if '<|think_off|>' in msg.content %}
- ...
- {%- elif '<|think_on|>' in msg.content %}
- ```
- Therefore this content:
- ```text
- <|think_off|> ... <|think_on|>
- ```
- leaves thinking disabled, even though the last directive is `think_on`.
- It does not process directives in occurrence order.
- ### 2. One token may remain in the rendered prompt
- Removal also uses `if`/`elif`:
- ```jinja
- {%- if '<|think_off|>' in content %}
- remove think_off
- {%- elif '<|think_on|>' in content %}
- remove think_on
- ```
- If both are present, only one type is removed. For example:
- ```text
- <|think_off|>foo<|think_on|>
- ```
- can leave `<|think_on|>` visible to the model.
- Both markers should always be removed independently:
- ```jinja
- {%- set content = content
- | replace('<|think_off|>', '')
- | replace('<|think_on|>', '')
- | trim
- %}
- ```
- State selection still needs separate occurrence-order logic.
- ### 3. Users can override system/developer thinking policy
- The scan includes:
- ```jinja
- msg.role == 'system' or
- msg.role == 'developer' or
- msg.role == 'user'
- ```
- Consequently, a later user message can override a system-level `<|think_off|>` with `<|think_on|>`.
- That may be intentional, but if these are privileged control tokens, it is a hierarchy violation. Consider scanning only system/developer messages, or explicitly document that users are allowed to control thinking.
- ### 4. Literal quoted tokens also change behavior
- Any historical user message containing the literal text—such as a discussion explaining `<|think_off|>`—changes generation behavior. There is no escaping mechanism.
- ### 5. List-of-string content is inconsistent
- The new renderer permits non-mapping list items by stringifying them:
- ```jinja
- {{- item | string }}
- ```
- But the preliminary thinking scan only checks mapping items with a `text` field. A list item equal to `"<|think_off|>"` can be removed later without ever changing `ns_state.thinking`.
- ---
- ## Historical thinking reconstruction regression
- The original always emits a thinking block when preservation applies, even when reasoning is empty:
- ```text
- <think>
- </think>
- <tool_call>...
- ```
- The modified version requires nonempty reasoning:
- ```jinja
- {%- if (_preserve_thinking or loop.index0 > ns.last_query_index)
- and reasoning_content %}
- ```
- Thus an assistant tool call with no stored reasoning becomes:
- ```text
- <tool_call>...
- ```
- This is inconsistent with the generation prompt when thinking is disabled, which emits:
- ```text
- <think>
- </think>
- ```
- This matters during multi-step tool use:
- 1. The generation prompt starts with an empty thinking block.
- 2. The model emits a tool call.
- 3. The application stores structured `tool_calls`, but no `reasoning_content`.
- 4. On the next render, the modified template reconstructs the assistant turn without the empty thinking block.
- The original template reconstructs it consistently.
- Whether the empty block is strictly required depends on the checkpoint, but this is a concrete format change and should be tested. To preserve original behavior, remove `and reasoning_content`.
- ---
- ## Reasoning extraction improvements
- The modified template does improve compatibility by accepting:
- - `message.reasoning_content`,
- - `message.thinking`,
- - non-string reasoning values,
- - some inline `<think>...</think>` content.
- That is a useful fix.
- However, the fallback parser is heuristic:
- - It misses forms such as `<think>x</think>` without a newline before the closing tag.
- - With multiple closing tags, `split(...)[-1]` may discard intermediate content.
- - A literal `</think>` appearing in an answer can be mistaken for a reasoning delimiter.
- - If `reasoning_content` exists but is empty, inline reasoning is not parsed.
- ---
- # 2. Tool-call support
- ## XML arguments: original limitation and modified partial fix
- ### Original behavior
- The original assumes `tool_call.arguments` is a mapping:
- ```jinja
- {%- for args_name, args_value in tool_call.arguments|items %}
- ```
- This works for:
- ```json
- {
- "name": "weather",
- "arguments": {
- "city": "Paris"
- }
- }
- ```
- But many OpenAI-compatible APIs store arguments as a JSON string:
- ```json
- {
- "name": "weather",
- "arguments": "{\"city\":\"Paris\"}"
- }
- ```
- Applying `|items` to that string normally raises a template error.
- This is a real problem in the original.
- ### Modified XML behavior
- The modified template avoids the exception, but does not actually convert the JSON string into XML parameters:
- ```jinja
- {%- elif tc.arguments is string and tc.arguments %}
- {{- tc.arguments }}
- ```
- It produces:
- ```xml
- <tool_call>
- <function=weather>
- {"city":"Paris"}</function>
- </tool_call>
- ```
- That does not follow the instructed XML dialect:
- ```xml
- <parameter=city>
- Paris
- </parameter>
- ```
- So the new XML branch changes a hard failure into malformed or parser-incompatible output.
- ### Recommended solution
- For XML mode, either:
- 1. Require arguments to be normalized to a mapping by the host application, or
- 2. Parse JSON-string arguments before applying the template, or
- 3. Raise a clear exception instructing the caller to use JSON mode.
- Do not silently place raw JSON inside `<function>`.
- ---
- ## JSON tool mode
- The new JSON mode can correctly handle both mapping and JSON-string arguments:
- ```xml
- <tool_call>
- {"name":"weather","arguments":{"city":"Paris"}}
- </tool_call>
- ```
- That is a meaningful improvement.
- However, this only works end-to-end if all components agree on the format:
- - the model/checkpoint,
- - the chat template,
- - the inference server,
- - the tool-call parser,
- - the API response adapter.
- Changing the prompt format does not automatically teach a server’s Qwen tool parser to recognize it. If the server expects:
- ```xml
- <function=...>
- <parameter=...>
- ```
- then JSON mode may generate apparently correct text but no structured `tool_calls` in the API response.
- Likewise, a checkpoint primarily trained on the XML dialect may be less reliable in JSON mode, even though the prompt contains an example.
- `tool_call_format` should also be validated. Currently any value other than `"json"` silently selects XML.
- ---
- ## Scalar XML argument regression
- The original serializes every non-string value using JSON:
- ```jinja
- args_value | tojson
- ```
- Therefore:
- - `true` remains `true`
- - `false` remains `false`
- - `null` remains `null`
- The modified version only uses JSON for mappings and sequences:
- ```jinja
- {%- if args_value is mapping or
- (args_value is sequence and args_value is not string) %}
- {{ args_value | tojson }}
- {%- else %}
- {{ args_value | string }}
- {%- endif %}
- ```
- This can render:
- ```text
- True
- False
- None
- ```
- instead of:
- ```text
- true
- false
- null
- ```
- That is a regression, particularly if the tool-call parser interprets parameter values as JSON-like data.
- A better rule is the original one:
- ```jinja
- {%- if args_value is string %}
- {%- set _av = args_value %}
- {%- else %}
- {%- set _av = args_value | tojson %}
- {%- endif %}
- ```
- This also handles nested values correctly.
- ---
- ## Tool instructions
- The modified instructions improve parseability by telling the model to put pre-call reasoning inside `<think>` rather than as arbitrary text before the tool call. This is better than the original statement:
- > You may provide optional reasoning ... in natural language BEFORE the function call
- Free text outside a reasoning block can interfere with strict tool parsers.
- However, the new wording has two issues:
- 1. The same `<think>` instructions are included when thinking is disabled.
- 2. “ALL explanation and reasoning MUST be placed strictly inside `<think>`” is broad enough that the model could interpret it as applying to final answers as well.
- It would be better to scope the instruction:
- > When making a tool call, any planning before the call must remain inside the thinking block. Do not emit conversational text between the thinking block and the tool call.
- Also, the serializer still permits nonempty `message.content` before structured tool calls, even though the new prompt discourages it.
- ---
- ## Multiple tool calls and responses
- Both templates support:
- - multiple tool calls in one assistant message,
- - multiple consecutive tool results,
- - grouping tool results inside one synthetic user turn.
- The modified formatting adds slightly more separation between tool-call blocks, but remains structurally valid.
- Neither template includes tool-call IDs in the serialized tool responses. Results are effectively associated by position. That is inherited behavior and may be sufficient for the intended Qwen format, but it is weaker for parallel calls where IDs matter.
- ---
- ## Tool-response error detection is unsafe
- The modified template automatically classifies short tool results as failures and injects warnings.
- For example, this valid result:
- ```json
- {"error": null, "data": [1, 2, 3]}
- ```
- matches:
- ```jinja
- '"error":' in _content_head
- ```
- and receives:
- ```text
- ⚠️ SYSTEM WARNING: The previous tool call returned an error...
- ```
- Other false positives include:
- ```text
- error: 0
- no error: success
- ```
- Meanwhile, real errors can be missed when:
- - they are over 500 characters,
- - the error marker occurs after the first 80 characters,
- - output contains `$ `,
- - output contains `took `.
- The injected warning is also not a real system message. It is text inside a tool response wrapped as a user turn, despite saying `SYSTEM WARNING`.
- This behavior should be:
- - removed,
- - made opt-in, or
- - based on structured metadata such as `message.is_error`.
- It should not infer success/failure from arbitrary response text by default.
- ---
- ## Truncation support
- Optional truncation is useful and defaults to disabled, so it does not affect baseline behavior.
- There are inconsistencies:
- - `max_tool_arg_chars` only applies to mapped XML arguments.
- - It does not apply to raw string arguments.
- - It does not apply in JSON mode.
- - Truncating a nested JSON value can make the value invalid JSON.
- - Tool-response truncation can remove essential result data.
- These are acceptable only if documented as lossy context controls.
- ---
- # 3. Query and reasoning-preservation logic
- The original raises when it cannot find a real user query:
- ```jinja
- {{- raise_exception('No user query found in messages.') }}
- ```
- The modified version instead uses:
- ```jinja
- {%- if _last_idx > 50 %}
- {%- set ns.last_query_index = _last_idx %}
- {%- else %}
- {%- set ns.last_query_index = 0 %}
- {%- endif %}
- ```
- This is difficult to justify:
- - For a 50-message history, one preservation behavior is selected.
- - For a 52-message history, the opposite behavior may be selected.
- - A conversation with no user query is silently accepted.
- - The result of `preserve_thinking=false` becomes dependent on an arbitrary history length.
- This should be replaced with a deterministic policy:
- - retain the original exception, or
- - explicitly treat the entire history as a continuation, e.g. use `-1`, or
- - explicitly treat all history as prior context.
- The `> 50` heuristic is a correctness bug.
- ---
- # 4. System and developer messages
- ## Improvements
- The original only supports `system` as the first message. A `developer` role eventually causes:
- ```text
- Unexpected message role.
- ```
- The modified template supports `developer` and maps it to Qwen’s `system` role. That is useful for OpenAI-style message inputs.
- ## Regressions
- The original enforces:
- ```text
- System message must be at the beginning.
- ```
- The modified template allows later system/developer messages and emits them wherever they occur.
- It also converts any unknown role into:
- ```text
- [user]
- [role]: content
- ```
- instead of raising.
- This is more permissive, but less safe:
- - malformed message histories are hidden rather than detected,
- - a later system message can appear after user turns,
- - unknown observation/function roles may become user instructions,
- - model-role semantics are changed silently.
- A better policy is:
- 1. Merge all leading `system`/`developer` messages into the initial system prompt.
- 2. Reject system/developer messages after the first non-system message.
- 3. Reject unknown roles unless explicit compatibility conversion is enabled.
- ---
- # 5. Content and vision handling
- The modified renderer is safer for normal dictionary content items because it checks:
- ```jinja
- item is mapping
- ```
- before testing keys. The original can fail strangely on non-mapping list items.
- The modified version also explicitly defaults `add_vision_id`, which is cleaner.
- However, it now stringifies unsupported list items instead of raising:
- ```jinja
- {{- item | string }}
- ```
- That may hide malformed multimodal inputs. It also stops supporting object-like items with attributes such as `.type` unless they are mappings.
- Both templates correctly reject image/video content in system messages; the modified version extends that restriction to developer messages.
- ---
- # 6. Practical compatibility matrix
- | Input / feature | Original | Modified XML | Modified JSON |
- |---|---:|---:|---:|
- | Basic text chat | Works | Works | Works |
- | `reasoning_effort=xhigh` | Works | Works | Works |
- | `reasoning_effort=high` | Error | Works as xhigh | Works as xhigh |
- | Invalid effort | Clear error | Silently xhigh | Silently xhigh |
- | Thinking disabled | Works | Works, but tool prompt still mentions thinking | Same |
- | `reasoning_content` string | Works | Works | Works |
- | `thinking` alias | Ignored | Works | Works |
- | Inline `<think>` extraction | No | Partial support | Partial support |
- | Mapping tool arguments | Works | Mostly works | Works |
- | Boolean/null XML args | Correct JSON spelling | Python spelling regression | Correct JSON |
- | JSON-string arguments | Template error | Malformed XML dialect | Works if valid JSON |
- | Multiple tool calls | Works | Works | Works |
- | Developer role | Error | Supported as system | Supported as system |
- | Noninitial system role | Error | Silently accepted | Silently accepted |
- | No real user query | Clear error | Arbitrary fallback | Arbitrary fallback |
- | External parser compatibility | Native format only | Usually closest to base | Must be verified |
- | Valid `{"error": null}` result | Unmodified | Falsely warned | Falsely warned |
- ---
- # 7. Recommended changes before deployment
- At minimum:
- 1. **Validate reasoning effort** while retaining `high -> xhigh`.
- 2. **Validate `tool_call_format`** instead of silently defaulting unknown values to XML.
- 3. **Require mapping arguments in XML mode**, or parse JSON strings before templating.
- 4. **Serialize every non-string XML argument with `tojson`**.
- 5. **Restore consistent empty thinking-block reconstruction**, or deliberately test that the target checkpoint does not need it.
- 6. **Remove the `> 50` no-query fallback**.
- 7. **Remove both think-control markers unconditionally** and define proper precedence.
- 8. **Decide whether users are allowed to override system thinking policy**.
- 9. **Remove or opt in to heuristic tool-error warnings**.
- 10. **Merge only leading system/developer messages and reject later ones**.
- 11. **Verify JSON tool mode against the actual inference server’s tool parser**.
- ## Overall assessment
- - The **original** is narrower and stricter. It has weaker compatibility, especially for developer roles and JSON-string tool arguments, but its failure modes are usually obvious.
- - The **modified** version adds valuable functionality, particularly JSON tool formatting and broader reasoning extraction, but it replaces several explicit errors with silent or malformed behavior.
- - For production use, the modified version should be treated as a feature branch requiring validation—not as a drop-in corrected version of the original.
- ---
- ## When the original works correctly
- The original expects reasoning to be stored separately:
- ```json
- {
- "role": "assistant",
- "reasoning_content": "I should inspect the weather.",
- "content": "The weather is sunny."
- }
- ```
- It renders that correctly:
- ```text
- <think>
- I should inspect the weather.
- </think>
- The weather is sunny.
- ```
- There is no blank block before the actual reasoning.
- ## When the original breaks
- Some clients store the raw generated completion entirely in `content`:
- ```json
- {
- "role": "assistant",
- "content": "I should inspect the weather.\n</think>\n\nThe weather is sunny."
- }
- ```
- This is plausible because the generation prompt already supplied the opening:
- ```text
- <think>
- ```
- and the generated text contains the reasoning followed by `</think>`.
- The original does not parse reasoning out of `content`. Since `reasoning_content` is absent, it emits an empty block and then appends the raw content:
- ```text
- <think>
- </think>
- I should inspect the weather.
- </think>
- The weather is sunny.
- ```
- That is malformed history: it has an empty thinking block followed by reasoning and an unmatched closing tag.
- ## What the modified template fixes
- The modified template tries to detect a closing `</think>` inside `content`, split out the preceding reasoning, and reconstruct:
- ```text
- <think>
- I should inspect the weather.
- </think>
- The weather is sunny.
- ```
- That is a real compatibility improvement for clients that store raw reasoning in `content`.
- ## Important qualifications
- Calling this “the official template poisons history” is too broad:
- - The original works correctly when `reasoning_content` is populated as expected.
- - Empty `<think></think>` blocks are not inherently poisonous. They can be the correct representation of a no-thinking assistant turn.
- - The modified fallback parser is heuristic and does not recognize every possible formatting variation.
- - If `reasoning_content` exists but is an empty string, the modified template will not fall back to parsing reasoning from `content`.
- - The modified template introduces the opposite inconsistency: it suppresses historical empty thinking blocks because of:
- ```jinja
- and reasoning_content
- ```
- This can make a reconstructed no-thinking turn differ from the turn that was originally generated.
- For example, with thinking disabled, the generation prompt includes:
- ```text
- <think>
- </think>
- ```
- But if the stored assistant message only contains the answer, the modified template reconstructs it without that empty block. The original reconstructs it with the block.
- ### Better documentation wording
- A more accurate claim would be:
- > Recovers reasoning from assistant `content` when clients do not store it separately in `reasoning_content`, preventing duplicate empty thinking blocks and unmatched closing tags in that representation.
- ---
- # 2. “Tool calling crashes”
- > If your client passes arguments as JSON strings—the standard OpenAI API format—the official template crashes.
- This is substantially correct.
- ## Original behavior
- The original assumes `arguments` is a mapping:
- ```jinja
- {%- for args_name, args_value in tool_call.arguments|items %}
- ```
- It expects:
- ```json
- {
- "name": "get_weather",
- "arguments": {
- "city": "Paris"
- }
- }
- ```
- But OpenAI-style API tool calls commonly use a JSON-encoded string:
- ```json
- {
- "name": "get_weather",
- "arguments": "{\"city\":\"Paris\"}"
- }
- ```
- On modern Jinja, applying `|items` to a string normally raises an error such as:
- ```text
- TypeError: Can only get item pairs from a mapping.
- ```
- Therefore, yes: the original can crash if the adapter passes the wire-format JSON string directly to the template.
- There is a nuance: many inference servers or client adapters parse the JSON string into a dictionary before invoking the template. In those environments, the original works. The original template assumes that normalization has already occurred.
- ## Does the modified template fix it?
- ### In JSON tool-call mode: yes, mostly
- With:
- ```text
- tool_call_format = "json"
- ```
- the modified template accepts either:
- - a mapping, serialized with `tojson`, or
- - an existing JSON string, inserted as the `arguments` value.
- It can produce:
- ```xml
- <tool_call>
- {"name": "get_weather", "arguments": {"city":"Paris"}}
- </tool_call>
- ```
- That is a valid fix, assuming:
- 1. The argument string is valid JSON.
- 2. The model supports this tool-call dialect.
- 3. The inference server’s tool parser recognizes JSON inside `<tool_call>`.
- The template does not validate the raw JSON string.
- ### In default XML mode: it avoids the crash but does not serialize correctly
- The modified template defaults to:
- ```jinja
- _tool_format = 'xml'
- ```
- For string arguments, it emits the string directly:
- ```jinja
- {%- elif tc.arguments is string and tc.arguments %}
- {{- tc.arguments }}
- ```
- The result is:
- ```xml
- <tool_call>
- <function=get_weather>
- {"city":"Paris"}</function>
- </tool_call>
- ```
- But the documented XML format requires:
- ```xml
- <tool_call>
- <function=get_weather>
- <parameter=city>
- Paris
- </parameter>
- </function>
- </tool_call>
- ```
- So in default XML mode, the modified template changes the failure from:
- - template crash
- to:
- - nonconforming tool-call history that the model or parser may not understand.
- That is not a complete fix.
- ## Additional XML regression
- The original JSON-serializes all non-string values. The modified XML branch stringifies scalar values. This can change:
- ```json
- true
- false
- null
- ```
- into Python-style:
- ```text
- True
- False
- None
- ```
- That may be incorrectly parsed by downstream tool-call parsers.
- The modified XML branch should use:
- ```jinja
- {%- if args_value is string %}
- {%- set _av = args_value %}
- {%- else %}
- {%- set _av = args_value | tojson %}
- {%- endif %}
- ```
- ## Better documentation wording
- > Adds JSON-string argument support in JSON tool-call mode. The original template expects arguments to be pre-parsed into a mapping and may raise a Jinja error otherwise.
- It should not claim complete support for JSON-string arguments in the default XML mode.
- ---
- # 3. “100% KV Cache hits”
- > Keeps past thoughts intact by default so your prefix cache stays warm across turns.
- This is the weakest claim.
- ## The original already preserves thoughts by default
- The original condition is:
- ```jinja
- preserve_thinking is undefined
- or preserve_thinking is true
- or loop.index0 > ns.last_query_index
- ```
- If `preserve_thinking` is not specified, historical reasoning is preserved.
- The modified template also defaults preservation to true:
- ```jinja
- {%- set _preserve_thinking = true %}
- ```
- Therefore, “keeps past thoughts intact by default” is not a new behavior. The original already does it.
- The modified version adds compatibility with a `preserve_reasoning` option and can recover reasoning embedded in `content`, but preservation itself was already present.
- ## Preserving reasoning does not guarantee cache hits
- Prefix caching requires the newly rendered prompt to have an exact token prefix matching a previously processed sequence.
- That depends on much more than retaining reasoning:
- - identical system prompt,
- - identical tool definitions and ordering,
- - identical whitespace and delimiters,
- - identical reasoning serialization,
- - identical thinking settings,
- - no changed truncation,
- - no changed template version,
- - support for prefix caching in the inference server,
- - the cached entry still being resident.
- A template cannot promise “100% KV cache hits.”
- At best, it can improve the chance that the previously generated prefix is reconstructed exactly.
- ## Cases where the modified template may improve caching
- If raw reasoning is stored in `content`, the original reconstructs a malformed and different prefix:
- ```text
- <think>
- </think>
- actual reasoning
- </think>
- ```
- The modified parser may reconstruct the original sequence more accurately. That can improve cache reuse.
- ## Cases where the modified template can reduce cache reuse
- ### 1. It removes empty historical thinking blocks
- As noted above, the modified template only emits historical thinking when `reasoning_content` is nonempty:
- ```jinja
- and reasoning_content
- ```
- If the previous generation used:
- ```text
- <think>
- </think>
- Answer
- ```
- but the next render reconstructs:
- ```text
- Answer
- ```
- the prefix differs at the beginning of that assistant turn. The cache cannot be reused past that point.
- ### 2. A new user think-control token can change the initial system prompt
- The modified template scans all messages for:
- ```text
- <|think_off|>
- <|think_on|>
- ```
- A new user message can change `ns_state.thinking`, which changes whether the reasoning-effort instruction appears in the initial system message.
- For example, adding `<|think_off|>` to the latest user message can remove the initial xhigh instruction on the next render. That changes the prompt near its beginning and can invalidate essentially the whole prefix cache.
- ### 3. Reformatting can change tokens
- The modified template trims and reconstructs reasoning/content. Even semantically equivalent changes in newlines or spaces can prevent exact token-prefix matching.
- ## What “100%” could reasonably mean
- It might mean:
- > All unchanged tokens from the preceding conversation are eligible for prefix-cache reuse.
- Even then, it is only true if the template reproduces those tokens exactly. Newly added user and generation-prompt tokens obviously are not already cached, so the entire new request cannot generally be a 100% cache hit.
- ## Better documentation wording
- > Preserves historical reasoning by default and can reconstruct reasoning embedded in assistant content, helping maintain a stable prefix for KV-cache reuse. Actual cache-hit rates depend on exact token stability and inference-server support.
- ---
- # Overall assessment of the claims
- ### Claim 1: blank thinking blocks
- **Based on a real issue, but missing the condition.**
- The original fails when raw reasoning remains in `content`; it works when `reasoning_content` is correctly populated. The modified template improves compatibility but introduces empty-think reconstruction inconsistencies.
- ### Claim 2: JSON-string tool arguments
- **Correct about the original crash.**
- However, the modified version properly supports such arguments only in JSON tool mode. Its default XML mode avoids the exception but emits the arguments in the wrong dialect.
- ### Claim 3: 100% KV cache hits
- **Mostly marketing.**
- The original already preserves reasoning by default, and no chat template can guarantee a cache-hit percentage. The modified template can improve prefix stability in one message-storage scenario, but can also reduce it in others.
Advertisement
Add Comment
Please, Sign In to add comment