core.tokenizers.text.parsers.nemotron_v3_reasoning_parser#
Module Contents#
Classes#
Parser for NVIDIA Nemotron 3 (Super, Ultra) reasoning output. |
API#
- class core.tokenizers.text.parsers.nemotron_v3_reasoning_parser.NemotronV3ReasoningParser#
Bases:
megatron.core.tokenizers.text.parsers.deepseek_r1_reasoning_parser.DeepSeekR1ReasoningParserParser for NVIDIA Nemotron 3 (Super, Ultra) reasoning output.
Behaves like
DeepSeekR1ReasoningParser, except when reasoning is disabled viaenable_thinking=False, or the caller passesforce_nonempty_content=True: in that case, if no content would otherwise be returned (either because</think>never closes, e.g. reasoning exceeded the max length, or because it closes with nothing following it), the reasoning text is returned as content instead of being discarded, so callers always get a non-empty response.- static _should_force_content(chat_template_kwargs: dict | None) bool#
Whether would-be-empty content should be backfilled from reasoning.
Mirrors vLLM’s
SuperV3ReasoningParser._should_force_content: force content when reasoning is disabled (enable_thinking is False) or the caller explicitly requests it (force_nonempty_content is True). Both flags are supplied by the client insidechat_template_kwargs.
- static parse(text: str, **kwargs) tuple[str, dict[str, str]]#
Extract reasoning content delimited by
<think>...</think>tags.Delegates the
<think>/</think>split toDeepSeekR1ReasoningParser, then surfaces the reasoning as content (instead of discarding it) when reasoning was disabled or the caller forced non-empty content.- Parameters:
text (str) – The text to parse.
chat_template_kwargs (dict, optional) – The request’s
chat_template_kwargs. When it setsenable_thinking=Falseorforce_nonempty_content=True, reasoning is surfaced as content rather than discarded if there would otherwise be no content.
- Returns:
A tuple containing the unprocessed text and a dictionary with the extracted reasoning content.
- Return type:
tuple[str, dict[str, str]]