> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/sdgm/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/sdgm/_mcp/server.

# Buy-It-Again Recommendation

## Solution Background and Business Value

Buy-it-again recommendations make it easy for customers to repurchase products they already know and trust.
These recommendations:

* Increase repeat purchases by surfacing past buys at the right moment.

* Improve customer retention by keeping users engaged with relevant products.

* Personalize marketing campaigns, including push notifications, in-app recommendations, and emails.

## Data Requirements and Schema

To build a buy-it-again recommendation model, you need three core tables: Users, Items, and Transactions.
Kumo AI can incorporate additional signals beyond this minimum dataset to improve model quality.

**Core Tables**

1. **Users Table**

   Stores user details.

   * `user_id`: Unique identifier (Primary Key).
   * `join_timestamp`: When the user joined.
   * `age`, `location`, `other_features`: Optional user attributes.

2. **Items Table**

   Stores product details.

   * `item_id`: Unique identifier (Primary Key).
   * `item_name`, `category`: Product metadata.
   * `start_timestamp` / `end_timestamp`: Item availability window.
   * `price`, `color`, `other_features`: Additional item features.

3. **Transactions Table**

   Stores user purchase history.

   * `transaction_id`: Unique identifier (Primary Key).
   * `user_id`: Foreign Key linking to Users.
   * `item_id`: Foreign Key linking to Items.
   * `timestamp`: Purchase date.
   * `total_amount`, `payment_method`, `other_features`: Transaction metadata.

**Entity Relationship Diagram (ERD)**

```mermaid
erDiagram
    USERS {
        INT user_id PK
        TIMESTAMP join_timestamp
        INT age
        STRING location
        STRING other_features
    }
    
    ITEMS {
        INT item_id PK
        STRING item_name
        STRING category
        TIMESTAMP start_timestamp
        TIMESTAMP end_timestamp
        FLOAT price
        STRING color
        STRING other_features
    }
    
    TRANSACTIONS {
        INT transaction_id PK
        INT user_id FK
        INT item_id FK
        TIMESTAMP timestamp
        FLOAT total_amount
        STRING payment_method
        STRING other_features
    }

    USERS ||--o{ TRANSACTIONS : "has"
    ITEMS ||--o{ TRANSACTIONS : "includes"
```

## Predictive Queries

One challenge in buy-it-again recommendations is differentiating repeat purchases from one-time buys.
A model trained only on past repeat purchases misses important behavioral signals from non-repeat interactions.

Kumo trains a general item-to-user recommendation model and applies filters at prediction time, ensuring:

* The model learns overall user-item affinity.
* Users receive only buy-it-again recommendations.

```pql
PREDICT LIST_DISTINCT(transactions.item_id, 0, X, days) RANK TOP 50
FOR EACH users.user_id
WHERE COUNT(transactions.*, -D, 0, days) >= N
```

This Predictive Query Language (PQL) query:

* Predicts the top 50 distinct items a user is likely to buy again within a future X-day window.
* Limits predictions to active users who have made at least N purchases in the last D days, which avoids empty recommendation sets after filtering.

**Filtering Out Newly Introduced Items**

To exclude newly launched items that users have not had time to re-purchase, apply post-processing in SQL:

```sql
SELECT *
FROM (
    PREDICTIONS 
    JOIN (
        SELECT entity_id, item_id
        FROM <ORDERS>
        WHERE timestamp <= PREDICTION_ANCHOR_TIME
    ) AS CANDIDATES 
    ON PREDICTIONS.entity_id = CANDIDATES.entity_id 
       AND PREDICTIONS.item_id = CANDIDATES.item_id
);
```

## Building models in Kumo Fine-Tune SDK

Kumo AI simplifies ML modeling on relational data, making it well suited for this problem.

**1. Initialize the Kumo Fine-Tune SDK**

```python
import kumoai as kumo

kumo.init(url="https://<customer_id>.kumoai.cloud/api", api_key=API_KEY)
```

**2. Create a Connector for Data Storage**

```python
connector = kumo.S3Connector("s3://your-dataset-location/")
```

**3. Select tables**

```python
users = kumo.Table.from_source_table(
    source_table=connector.table('users'),
    primary_key='user_id',
).infer_metadata()

items = kumo.Table.from_source_table(
    source_table=connector.table('items'),
    primary_key='item_id',
).infer_metadata()

transactions = kumo.Table.from_source_table(
    source_table=connector.table('transactions'),
    time_column='timestamp',
).infer_metadata()
```

**4. Create graph schema**

```python
graph = kumo.Graph(
    tables={
        'users': users,
        'items': items,
        'transactions': transactions,
    },
    edges=[
        dict(src_table='transactions', fkey='user_id', dst_table='users'),
        dict(src_table='transactions', fkey='item_id', dst_table='items'),
    ],
)

graph.validate(verbose=True)
```

**5. Train the model**

```python
pquery = kumo.PredictiveQuery(
    graph=graph,
    query=(
        "PREDICT LIST_DISTINCT(transactions.item_id, 0, X, days) RANK TOP 50\n"
        "FOR EACH users.user_id"
    ),
)
pquery.validate(verbose=True)

model_plan = pquery.suggest_model_plan()
trainer = kumo.Trainer(model_plan)
training_job = trainer.fit(
    graph=graph,
    train_table=pquery.generate_training_table(non_blocking=True),
    non_blocking=False,
)
print(f"Training metrics: {training_job.metrics()}")
```

**6. Run the model**

```python
prediction_job = trainer.predict(
    graph=graph,
    prediction_table=pquery.generate_prediction_table(non_blocking=True),
    output_types={'predictions', 'embeddings'},
    output_connector=connector,
    output_table_name='buy_it_again_predictions',
    training_job_id=training_job.job_id,
    non_blocking=False,
)
print(f'Batch prediction job summary: {prediction_job.summary()}')
```