Run Training
In the BaseModel Python App, foundation model training can be run in two ways: as a single pipeline or as two separate stages. Both support Python and CLI execution.
Quick-check output path
A quick check on a new output path does not need overwrite. Rerunning it on the same path raises FileExistsError unless you pass overwrite or resume, so use the overwrite-enabled examples in Quick Check and keep that path disposable.
Joint Pipeline
The simplest option. The pretrain function validates your config, fits the behavioral representation, and trains the foundation model in one go.
Python
from monad.ui import pretrain
from pathlib import Path
pretrain(
config_path=Path("path/to/config.yaml"),
output_path=Path("path/to/store/pretrain/artifacts"),
)
CLI
python -m monad.run \
--pretrain \
--config-path "path/to/config.yaml" \
--output-path "path/to/store/pretrain/artifacts"
Modular Pipeline
Split training into two stages when you need different environments for each — for example, a CPU-heavy machine for fitting and a GPU machine for training.
Stage 1: Fit behavioral representation
Analyzes your data and builds the feature representation.
Python:
from monad.ui import fit_behavioral_representation
from pathlib import Path
fit_behavioral_representation(
config_path=Path("path/to/config.yaml"),
output_path=Path("path/to/store/pretrain/artifacts"),
)
CLI:
python -m monad.run \
--fit \
--config-path "path/to/config.yaml" \
--output-path "path/to/store/pretrain/artifacts"
Stage 2: Train foundation model
Trains the neural network using the representations from Stage 1. No config_path is needed — BaseModel reads the config stored during fitting.
Python:
from monad.ui import train_foundation_model
from pathlib import Path
train_foundation_model(
output_path=Path("path/to/store/pretrain/artifacts"),
)
CLI:
Offline Prediction
Once a model is trained, you can score new data from the command line with the --predict stage — no custom Python code required. It loads a trained checkpoint together with a TestingParams YAML, then writes predictions to the configured save location.
CLI:
python -m monad.run \
--predict \
--checkpoint-path "path/to/checkpoint" \
--testing-params-path "path/to/testing_params.yaml"
Optional flags: --storage-config-path for an external storage config, and --seed to fix the ordering of results.
Single GPU
--predict runs on a single GPU. Local output is written as TSV; set remote_save_location in the TestingParams YAML to write to Snowflake or Databricks instead.
See Inference for the full prediction workflow.
Quick Check (Dry Run)
Use a quick check to find configuration and data-access problems before a longer training run. It runs the selected stages with reduced data and batch counts. With pretrain, this includes fitting the behavioral representation and a short training loop. The duration also depends on startup, feature preparation and model size.
Complete Setup first. Run the examples below inside the Python App container. Replace the example config and output paths with your own.
Python
from pathlib import Path
from monad.ui import pretrain
pretrain(
config_path=Path("/workspace/configs/fm_config.yaml"),
output_path=Path("/workspace/project/quick-check-001"),
quick_check=True,
overwrite=True,
)
The flag is also available on fit_behavioral_representation() and train_foundation_model(); each checks its selected stage and still requires that stage's inputs.
CLI
python -m monad.run \
--pretrain \
--config-path /workspace/configs/fm_config.yaml \
--output-path /workspace/project/quick-check-001 \
--quick-check \
--overwrite
A quick check on a new output path runs without overwrite. The examples pass it so that you can rerun the check on the same path, which otherwise raises FileExistsError. Without an interactive terminal, --overwrite proceeds without a confirmation prompt (see Overwrite confirmation). Keep this output path new, disposable, and separate from a real model directory because overwrite deletes its contents. There is no --validate CLI flag; use the documented quick-check command for this diagnostic workflow.
What quick check does
For version 1.14, the default quick-check limits are:
- Fit stage — entity sampling limit of 1,000, history limit of 50 and approximately 5% entity filtering in SQL. Actual sample sizes depend on the data.
- Train stage — one epoch, five training batches and five validation batches.
These checks cover only the selected stages and sampled data. They do not test downstream training, prediction, or every row in the source. Quick-check outputs are marked _QUICK_CHECK, and the run logs that they are not suitable for production use. The loader does not refuse them, but they come from sampled data and a shortened run, so do not build downstream work on them. For a checkpoint to build on, use the minimal-training example below.
Choosing a Smoke Test
A smoke test is a short run that checks whether the parts of the workflow you need work together. Choose the test according to what you want to verify:
| Test | Purpose | What a pass does not establish |
|---|---|---|
| Quick check | Diagnose configuration, connectivity and selected stages on reduced data | A reusable FM checkpoint or a complete downstream pipeline |
| Minimal training | Execute real optimizer steps, save a normal checkpoint and verify loading | Production model capacity or useful predictive quality |
| Test with your intended model configuration | Keep the intended model size and feature types while limiting batches/data | Full-data throughput or final quality |
| End-to-end smoke | Exercise FM, downstream training and the required evaluation/prediction steps | Production-scale reliability or quality |
Minimal Training with a Loadable Checkpoint
Save a copy of your valid config, for example as /workspace/configs/fm_smoke.yaml. Keep its required data-source and split settings, and replace or update the following blocks (do not append duplicate YAML keys):
training_params:
devices: [0] # Choose an available GPU
epochs: 1
limit_train_batches: 2
limit_val_batches: 1
data_loader_params:
batch_size: 32
num_workers: 0
memory_constraining_params:
hidden_dim: 64
num_layers: 1
calibration_params:
enabled: false
Batch limits shorten the training loop; they do not limit the earlier feature-fitting work. To reduce that work, use a smaller input sample while keeping enough entity history and examples for both training and validation. Keep related tables consistent so joins still work. You can also reduce expensive feature types for a basic functional test. Any tables, features or joins left out of the test remain untested. If you want to check whether your intended model fits on the available hardware, keep its original dimensions and feature types.
Run without --quick-check, using a fresh output directory:
python -m monad.run \
--pretrain \
--config-path /workspace/configs/fm_smoke.yaml \
--output-path /workspace/project/smoke-001
A successful minimal-training check requires:
- A successful process exit and logs showing actual training steps and validation.
- A completed checkpoint under
/workspace/project/smoke-001/fm, including weights and the_FINISHEDreadiness marker. - Successful loading through
load_from_foundation_modelwith your real downstream task and target function; see Loading Overrides.
Successful loading confirms that the checkpoint is usable. A model trained on only a few batches is useful for checking the workflow, not for judging prediction quality. If this short run is slow, compare the time spent on startup, feature fitting and training to identify which stage needs attention.
Key Flags
All three functions accept the following flags:
resume— reuse partial results from a previous run and continue training from the last training checkpoint. If there is no training checkpoint, training starts from scratch;train_foundation_modelstill requires completed fit-stage outputs.overwrite— discard previous results before starting: withpretrainandfit_behavioral_representation, everything atoutput_path; withtrain_foundation_model, only itsfmsubdirectory, so the fitted features are kept. Destructive — the directory contents are deleted. The directory itself is kept and must be writable by the current user; this is checked before anything is deleted.seed— set a seed for reproducibility. As of 1.7, this also makes modality sketch-dropping reproducible, so seeded runs are reproducible end to end.
Cannot resume and overwrite simultaneously
resume and overwrite cannot both be True — this raises an error.
Function parameters override YAML config
Parameters passed to the function override those in the YAML config.
Adding --overwrite or --resume in CLI
python -m monad.run \
--pretrain \
--config-path "path/to/config.yaml" \
--output-path "path/to/store/pretrain/artifacts" \
--overwrite
Overwrite confirmation
When the output directory already holds completed steps, overwrite asks for confirmation on an interactive terminal. Any answer other than y or yes, or end of input without an answer, aborts the run with To reuse existing outputs, run with --resume flag. Aborting....
When stdin is not a terminal (for example docker exec without -t, or a CI job), the run logs a warning and proceeds without asking. It proceeds even when quick-check outputs are about to replace a production model, so check --output-path before you run a noninteractive overwrite.
1.13 and earlier
In 1.13 and earlier, a noninteractive run with no answer on stdin ended with EOFError. Pipe the answer (echo y | docker exec -i …) when you run those versions.
Resource Estimation
During the fit stage, BaseModel automatically estimates memory requirements for each feature computation task and logs a resource estimation report. The report shows:
- Available RAM and safety margin (default 20 %)
- Per-task memory estimates with component breakdowns
- Predicted peak memory usage
- Recommended vs. configured
num_concurrent_features
If the configured concurrency exceeds the recommendation, BaseModel prompts for confirmation in interactive mode or logs a warning in non-interactive mode.
To cap the RAM budget, set max_ram_gb in the query_optimization section of your YAML config:
DataLoader Calibration
When calibration_params.enabled is set to true in your YAML config, BaseModel automatically benchmarks DataLoader settings before foundation model training begins. It sweeps through candidate num_workers and prefetch_factor values, measures throughput, and applies the most efficient configuration.
This is especially useful when deploying to new hardware or when you are unsure which DataLoader settings work best. The calibration result is saved alongside model artifacts, so resumed runs skip re-calibration.
If calibration fails for any reason, training proceeds with the default data_loader_params from your config — no manual intervention required.
For details on all calibration parameters and tuning guidance, see Scaling & Memory → Automatic DataLoader Calibration.
Verifying Training Completed
How to confirm training completed
Training is complete when:
-
Console output confirms model checkpoints have been saved:
-
An empty
_FINISHEDmarker file appears in thefmsubdirectory of youroutput_path, next to the best model checkpoint (best_model.ckpt).
Training progress
Set show_entity_progress: true in training_params to add a Rich progress bar showing entity-level progress with percentage complete, elapsed time, and estimated time remaining:
The entity progress bar is off by default. Before each stage without a batch limit, it runs a distinct-entity count on global rank 0 to get the total; with the default, no counting queries run. In 1.13, the entity bar and its counts were on by default. The entity bar supplements the standard PyTorch Lightning training output. Set the ENABLE_PROGRESS_BAR environment variable to False to disable all progress bars, including the entity bar.
Columns Analysis Report
During the fit stage, BaseModel logs a columns analysis report for each data source. The report lists:
- Column types — every column grouped by its inferred type (decimal, categorical, categoricalCompressed, time_series, text, image)
- Skipped columns — columns excluded from training, with reasons (too many NaNs, text-like, low cardinality, date, mixed-object) and hints on how to include them
- Action recommendations — suggestions such as potential time-series or text columns that may benefit from a
column_type_overridesentry, and redundant column pairs to consider removing
Columns analysis report
====================================
Table: transactions
====================================
Column type Columns
------------------------------------
decimal price, quantity, total_amount
categorical channel, region
categoricalCompressed article_id, store_id
------------------------------------
Skipped columns
------------------------------------
Reason Columns
------------------------------------
Text column product_description
Hint: Text columns are skipped by default. To include it,
set its type via column_type_overrides.
------------------------------------
Action recommendations
------------------------------------
Redundant categorical columns product_code, product_id
Hint: These columns form a bijection (one-to-one mapping).
Keeping both adds no new information. Remove the redundant
column via disallowed_columns.
====================================
The report also logs the total fit duration: Fitting took X.XX seconds.
Review this report after every fit run:
- Verify column types — confirm that columns are classified as you expect. If a column landed in the wrong type, add a
column_type_overridesentry. - Check skipped columns — some skips are expected and harmless:
- Date columns — raw dates are not embedded (this is by design). If the column is your event timestamp, configure it as
date_column. If it should be a feature, transform it with ansql_lambda. - Text columns — skipped by default. Add a
column_type_overrides: textentry if the column should contribute semantic features. - Low cardinality — columns with fewer than 2 unique values carry no signal and are safe to ignore.
- Too many NaNs — over 90 % missing or empty values. Clean the source data or fill/impute if you need this column.
- Mixed-object columns — inconsistent types within a column (e.g., strings mixed with numbers). Normalize the data or override the type explicitly.
- Unexpected skips — if a column you need is skipped, check whether the data source itself is wrong (e.g., pointing at the wrong table or missing a filter).
- Date columns — raw dates are not embedded (this is by design). If the column is your event timestamp, configure it as
- Act on recommendations — remove redundant column pairs via
disallowed_columnsto reduce computation without losing signal. See also Troubleshooting → Redundancy Report.
Suggested Config
After the fit stage completes, BaseModel generates a suggested_config.yaml file in the output directory. This file is a copy of your original config with the column report findings already applied:
| What is applied | How |
|---|---|
| Time-series candidates | An sql_lambda is added with resolve_fn() and a matching column_type_overrides: time_series entry |
| Bijection columns | Redundant columns from each bijection group are added to disallowed_columns (keeping one representative column per group — local columns are preferred over joined ones) |
Auto-generated entries are annotated with inline comments so you can tell them apart from your original config:
data_sources:
- name: transactions
disallowed_columns:
- product_code # auto: bijection with product_id
column_type_overrides:
price_ts: time_series # auto: time series candidate
sql_lambdas:
- alias: price_ts # auto: time series candidate
expression: "{{ resolve_fn('price') }}"
How to use
Review the generated file, verify the suggestions make sense for your use case, then rename it to replace your original config. You can also cherry-pick individual suggestions and apply them manually.
Note
suggested_config.yaml is not generated in quick check mode — it requires a full fit run.
What Happens During Training
The pipeline runs in two phases regardless of whether you use the joint or modular approach:
-
Phase 1
Fit Behavioral Representation
- Validate config
- Sample data
- Analyze columns
- Compute representations
-
Phase 2
Train Foundation Model
- Load representations
- Initialize trainer
- Calibrate DataLoader (if enabled)
- Training loop
- Save best model
The fitting phase uses distributed computation — it runs locally on your server and no data leaves the environment. Ensure the server is secured against unauthorized access.
For details on log signatures at each step and how to diagnose failures, see Troubleshooting.