Loading Overrides
In the BaseModel Python API, when you call load_from_foundation_model or load_from_checkpoint, you can override several settings that were configured during training. This lets you adjust data loading, split strategy, and prediction scope without retraining.
Checkpoint Contents
Both foundation model and scenario model training write a checkpoint directory. Understanding what's inside helps when debugging or moving checkpoints between environments.
checkpoint_dir/
├── best_model.ckpt # Best weights (selected by validation metric)
├── data.yaml # Data source configuration snapshot
├── dataloader.yaml # Data loader configuration
├── task.yaml # Task name and parameters
├── target_fn.bin # Serialized target function (scenario models only)
└── lightning_checkpoints/ # Epoch-by-epoch checkpoints (PyTorch Lightning)
best_model.ckpt is automatically selected from lightning_checkpoints/ based on the monitored validation metric. Each task type has a default:
| Task | Default metric | Direction |
|---|---|---|
| Binary classification | val_auroc_0 | Maximize |
| Multiclass classification | val_precision_0 | Maximize |
| Multilabel classification | val_auroc_0 | Maximize |
| Regression | val_loss | Minimize |
| Recommendation | val_HR@10_0 | Maximize |
Override with metric_to_monitor and metric_monitoring_mode in TrainingParams. Custom metrics replace the task's default metrics, so the default key may no longer be logged. If the monitored key is not logged during validation and checkpoint_dir or early_stopping is set, fit() raises InvalidConfigurationError before training starts.
Foundation Model Path Resolution
pretrain(output_path=X) writes the foundation-model checkpoint to X/fm. Pass a directory, not the best_model.ckpt file. An explicit X/fm path is the most portable choice across versions.
From release 1.12, the loader can also resolve X to its ready X/fm subdirectory. If the supplied directory is itself a ready checkpoint, it takes precedence over a nested fm/. For earlier versions, use the explicit checkpoint directory rather than relying on automatic resolution.
| Operation | Loader | Directory |
|---|---|---|
| Create a downstream model from an FM | load_from_foundation_model | X/fm, or X with supported automatic resolution |
| Reload a trained scenario for evaluation or prediction | load_from_checkpoint | The scenario's checkpoint directory |
A normal completed checkpoint has an _FINISHED readiness marker and its required artifacts. If loading fails, check the run's exit status, the final logs, the directory visible inside the container and the completeness of the checkpoint. A missing or incomplete checkpoint cannot be fixed just by appending fm/ to the path. Do not create or change readiness markers manually; they are written when the checkpoint is saved successfully. Quick-check artifacts are diagnostic outputs; use minimal training when you need a loadable FM.
What you can override
Beyond the three required arguments (checkpoint_path, downstream_task, target_fn), load_from_foundation_model accepts optional overrides that let you reuse a foundation model head, filter predictions, change the split strategy, or adjust data loading — all without retraining.
For the full parameter list and types, see the load_from_foundation_model() reference. The sections below explain when and why to use each override.
Foundation Model Head
Setting with_head=True loads the foundation model's final prediction layer. This can improve recommendation quality when the foundation model was trained on similar item interactions.
from monad.ui.module import load_from_foundation_model, RecommendationTask
trainer = load_from_foundation_model(
checkpoint_path="./foundation_model",
downstream_task=RecommendationTask(),
target_fn=reco_target,
with_head=True,
)
Prefix head
For recommendation tasks, downstream_head="prefix" uses the block of the foundation model head that predicts the next basket of the target column directly as the scenario model's output, instead of adding a new dense head (the default, downstream_head="linear"). It adds no new parameters, so fine-tuning starts from the pretrained predictions. It implies with_head=True.
trainer = load_from_foundation_model(
checkpoint_path="./foundation_model",
downstream_task=RecommendationTask(),
target_fn=reco_target,
downstream_head="prefix",
)
The foundation model must have been trained with use_last_basket_sketches=True (the default); otherwise loading raises a ValueError. Other task types do not support "prefix" and raise a ValueError.
Prediction Filtering
Narrow or exclude items from recommendation predictions on a per-entity basis:
# Exclude items the entity has already interacted with
trainer = load_from_foundation_model(
checkpoint_path="./foundation_model",
downstream_task=RecommendationTask(),
target_fn=reco_target,
predictions_to_exclude_fn=lambda events, attrs, ctx: (
list(events["transactions"]["product_id"].events)
),
)
# Include only items from a specific category
ELECTRONICS_IDS = [...]
trainer = load_from_foundation_model(
checkpoint_path="./foundation_model",
downstream_task=RecommendationTask(),
target_fn=reco_target,
predictions_to_include_fn=lambda events, attrs, ctx: ELECTRONICS_IDS,
)
Split Override
Override the train/validation/test split strategy defined during foundation model training:
from monad.core.checkpoint import TimeSplitOverride
trainer = load_from_foundation_model(
checkpoint_path="./foundation_model",
downstream_task=BinaryClassificationTask(),
target_fn=my_target_fn,
split=TimeSplitOverride(...),
)
EntitySplitOverride declares these fields:
| Field | Type | Meaning |
|---|---|---|
training | int | Percentage of entities assigned to training, for example 90. |
validation | int | Percentage assigned to validation, for example 10. |
training_validation_end | datetime | Last timestamp available to both groups. |
In Monad 1.14, passing that field-style mapping to a checkpoint loader raises ValueError: Unknown DataMode: training when the foundation model was trained with a time split. Until that loader path is fixed, configure the foundation YAML entity split with type: entity, integer training and validation percentages, training_validation_end, and the nested test range, then train the scenario from that foundation checkpoint.
Data Parameters Override
Override individual data-loading settings — extra columns, date boundaries, sampling, or split-point count — by passing them as keyword arguments directly to load_from_foundation_model. Any keyword that matches a DataParams field is applied. Any other keyword raises InvalidConfigurationError: Unknown keyword argument(s): [...] before the checkpoint is read; the message lists the fields that can be overridden. Training settings belong in fit(training_params=...). In 1.13 and earlier, unknown keywords were silently ignored.
trainer = load_from_foundation_model(
checkpoint_path="./foundation_model",
downstream_task=BinaryClassificationTask(),
target_fn=my_target_fn,
extra_columns=[
{"data_source_name": "transactions", "columns": ["order_id", "unit_price"]}
],
)
Extra columns in verify_target()
verify_target() reuses data parameters stored in the foundation model checkpoint, including extra_columns. Pass data_params_overrides only when validation must change those settings or a legacy checkpoint lacks them.
Sampling strategy
Scenario models default to random for classification and regression tasks, and to valid for recommendation tasks. Three strategies are available:
| Strategy | Supported tasks | Behavior |
|---|---|---|
random | Classification, Regression | Samples up to √(event count) split points per entity, capped by maximum_splitpoints_per_entity. Default for classification and regression. |
valid | All tasks | Places split points between events, up to maximum_splitpoints_per_entity. Default for recommendation; a recommendation task given random falls back to valid with a warning. |
existing | All tasks | Uses event timestamps as split points, up to maximum_splitpoints_per_entity. Useful for same-basket prediction and event-triggered scoring, and required for one-future / one-future-all-variants. |
When to use which
random— the default for classification and regression scenario training. Split points are drawn at random moments across each entity's history, independent of when events occurred, so the model learns to predict at any point in time. Foundation-model training instead accepts onlyvalidand coerces another strategy tovalid.valid— the default strategy for recommendation; a recommendation task configured withrandomfalls back tovalidand logs a warning. Each split point falls between two consecutive events, leaving enough room before the next event to build a valid target. Choose it for recommendation, where every training example needs a real "next" event to predict. Recommendation also supportsexisting, described below.existing— split points land exactly on event timestamps, so the model is trained to predict at the moment an event happens rather than on a calendar date. Choose it for event-triggered scoring (for example, predicting at the instant of a transaction) and for same-basket or next-basket prediction. It is required whensplit_point_inclusion_overridesusesone-futureorone-future-all-variants.
All three are bounded by maximum_splitpoints_per_entity — see Additional Parameters.
Split-point behavior
Control how events at the exact split timestamp are assigned and how much history the model sees. These settings matter most for recommendation and next-basket scenarios.
trainer = load_from_foundation_model(
...,
history_future_split_overrides={
"entity_history_limit": {"transactions": 2592000}, # 30 days in seconds
"max_data_splits_per_split_point": 200,
"split_point_inclusion_overrides": {"transactions": "future"},
},
)
When to use these:
- Simulating limited retention — set
entity_history_limit(in seconds, per source) when your production environment only has access to recent events. For example,2592000limits history to the last 30 days. - Next-basket recommendations — set
split_point_inclusion_overridestofuturefor your transaction source. This places same-timestamp events into the future window, so the model learns to predict the full basket. For "given these cart items, suggest another" scenarios, useone-future(one item goes to future, rest stay in history) orone-future-all-variants(generates multiple training examples per basket). Both options requiretarget_sampling_strategy: existing— the configuration validator rejects any other strategy. - Contextual recommendations — increase
max_data_splits_per_split_pointto generate more examples per split point, giving the model more context-item combinations to learn from.