> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/datadesigner/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/datadesigner/_mcp/server.

# Use Case Recipes

> Self-contained code examples that demonstrate how to leverage Data Designer for specific use cases.

Recipes are a collection of code examples that demonstrate how to leverage Data Designer in specific use cases. Each recipe is a self-contained example that can be run independently.

#### New to Data Designer?

Recipes provide working code for specific use cases without detailed explanations. If you're learning Data Designer for the first time, start with our [tutorial notebooks](/tutorials/overview), which offer step-by-step guidance and explain core concepts. Once you're familiar with the basics, return here for practical, ready-to-use implementations.

#### Prerequisite

These recipes use the OpenAI model provider by default. Ensure your OpenAI provider is set up via the Data Designer CLI before running a recipe.

## Code Generation

#### [Text to Python](/recipes/code-generation/text-to-python)

Natural-language instructions paired with Python implementations across complexity levels and industries.

*Python code generation · validation · LLM-as-judge*

#### [Text to SQL](/recipes/code-generation/text-to-sql)

Natural-language instructions paired with SQL implementations across complexity levels and industries.

*SQL code generation · validation · LLM-as-judge*

#### [Nemotron Super Text to SQL](/recipes/code-generation/nemotron-super-text-to-sql)

Enterprise-grade text-to-SQL training data — dialect-specific SQL, distractor injection, dirty data, 5 LLM judges with 15 scoring dimensions.

*Multi-dialect SQL · `SubcategorySamplerParams` · 5 judges · 15 score columns*

## QA and Chat

#### [Product Info QA](/recipes/qa-and-chat/product-info-qa)

Product information paired with question/answer pairs.

*Structured outputs · expression columns · LLM-as-judge*

#### [Multi-Turn Chat](/recipes/qa-and-chat/multi-turn-chat)

Multi-turn chat conversations between a user and an AI assistant.

*Structured outputs · expression columns · LLM-as-judge*

## Trace Ingestion

#### [Agent Rollout Trace Distillation](/recipes/trace-ingestion/agent-rollout-trace-distillation)

Read agent rollout traces from disk and turn each one into a structured workflow record inside a Data Designer pipeline. See the [ingestion guide](/concepts/agent-rollout-ingestion) for the trace format.

*`AgentRolloutSeedSource` · ATIF, Claude Code, Codex, Hermes formats · trace-aware prompts*

## MCP and Tool Use

#### [Basic MCP Tool Use](/recipes/mcp-and-tool-use/basic-mcp-tool-use)

Minimal example of MCP tool calling — defines a simple MCP server and generates data that requires tool calls to complete.

*`LocalStdioMCPProvider` · simple tool server · tool-augmented text*

#### [PDF Document QA](/recipes/mcp-and-tool-use/pdf-document-qa)

Grounded Q\&A pairs from PDF documents using MCP tool calls and BM25 search.

*`LocalStdioMCPProvider` · BM25 retrieval · per-column trace capture*

#### [Nemotron Super Search Agent](/recipes/mcp-and-tool-use/nemotron-super-search-agent)

Multi-turn search agent trajectories — Tavily web search via MCP, Wikidata KG seeding, BrowseComp-style question generation.

*Tavily MCP · Wikidata seeding · two-stage question generation · trajectory capture*

## Plugin Development

#### [Markdown Section Seed Reader](/recipes/plugin-development/markdown-section-seed-reader-plugin)

Define a custom `FileSystemSeedReader` inline and turn Markdown files into one seed row per heading section.

*Single-file custom reader · `hydrate_row()` fanout · `DirectorySeedSource` customization*

## VLM Long-Document Understanding

A 9-recipe pipeline for generating visual QA training data from long PDF documents: OCR, page classification, single-page / multi-page / whole-document QA, and frontier-model quality filtering.

#### [Seed Dataset Preparation](/recipes/vlm-long-document-understanding/seed-dataset-preparation)

Download PDFs, render page images, and prepare seed datasets for the downstream VLM recipes.

#### [Nemotron Parse OCR](/recipes/vlm-long-document-understanding/nemotron-parse-ocr)

Run Nemotron Parse over document pages and save OCR transcripts for text-based QA generation.

#### [Text QA from OCR Transcripts](/recipes/vlm-long-document-understanding/text-qa-from-ocr-transcripts)

Generate text-grounded question-answer pairs from OCR transcripts.

#### [Page Classification](/recipes/vlm-long-document-understanding/page-classification)

Classify pages by visual reasoning potential before running more expensive QA generation.

#### [Visual QA](/recipes/vlm-long-document-understanding/visual-qa)

Generate visual question-answer pairs from classified page images.

#### [Single-Page QA](/recipes/vlm-long-document-understanding/single-page-qa)

Generate single-page VLM QA examples from page-level image seeds.

#### [Multi-Page Windowed QA](/recipes/vlm-long-document-understanding/multi-page-windowed-qa)

Generate cross-page QA examples over fixed-size page windows.

#### [Whole-Document QA](/recipes/vlm-long-document-understanding/whole-document-qa)

Generate document-level QA examples over grouped page images.

#### [Frontier Judge QA Filter](/recipes/vlm-long-document-understanding/frontier-judge-qa-filter)

Score and filter generated QA pairs with a stronger independent judge.