Data Connectors
BaseModel reads data from multiple database types. This page documents the connection_params block for each supported database within the YAML configuration.
Overview
| Database | database_type | Required fields |
|---|---|---|
| Parquet | parquet | path |
| Snowflake | snowflake | user, password, account, warehouse, database, db_schema |
| BigQuery | bigquery | filename |
| Databricks | databricks | host, warehouse_id, (token or client_id + client_secret) |
| Hive | hive | hive_params |
| ClickHouse | clickhouse | host |
| Synapse | synapse | server_name, database_name, user |
Parquet
Read Parquet files from a local or mounted filesystem, or directly from S3. Supports single files, directories, and glob patterns.
| Parameter | Type | Default | Description |
|---|---|---|---|
path | Path \| str | required | Path to the Parquet file or directory. Glob patterns supported (e.g., *.parquet). From 1.14, an s3:// URI is accepted too — see Reading Parquet from S3. |
cache_path | Path \| None | None | Database cache path. Can be reused between trainings for faster restarts. |
config_overrides | dict | {} | DuckDB connection overrides (e.g., {"max_memory": "100GB"}). See DuckDB configuration reference and DuckDB temporary files. |
data_location:
database_type: parquet
connection_params:
path: "/data/transactions.parquet"
cache_path: "/basemodel/db_cache/"
table_name: transactions
Glob pattern example:
Reading Parquet from S3
From 1.14, path accepts an s3:// URI. BaseModel passes the URI to DuckDB unchanged, and DuckDB reads the objects through its httpfs extension.
data_location:
database_type: parquet
connection_params:
path: "s3://my-bucket/transactions/*.parquet"
table_name: transactions
- URI form. Write the URI as a plain string with two slashes after the scheme:
s3://bucket/key. A value in which the double slash has collapsed to one (s3:/bucket/key, which is what wrapping the URI in a Pythonpathlib.Pathproduces) is rejected with a validation error.gs://andaz://URIs are kept as they are and passed to DuckDB in the same way; this section covers S3 only. - Credentials. BaseModel does not pass S3 credentials to DuckDB. The
httpfsextension reads them from the environment variablesAWS_ACCESS_KEY_ID,AWS_SECRET_ACCESS_KEY,AWS_SESSION_TOKEN, andAWS_DEFAULT_REGION.config_overridesis passed to DuckDB unchanged, so DuckDB S3 settings such ass3_endpoint(for S3-compatible storage) can be set there. The storage configuration file does not apply to Parquetpath. - Caching. Without
cache_path, queries read straight from the bucket. Withcache_path, ans3://source is loaded into the local cache once and reused on later runs; BaseModel does not check the bucket for changes, so new or modified objects are ignored. Delete thecache_pathdirectory to force a reload.
Reading s3:// paths requires outbound internet access
The BaseModel image does not include the DuckDB httpfs extension. DuckDB downloads it from the DuckDB extension repository (extensions.duckdb.org) the first time it needs it. Without outbound access to that repository, reading an s3:// path fails because the extension cannot be installed; setting a DuckDB S3 option in config_overrides fails the same way. In such environments, mount the bucket as a filesystem and point path at the mounted files.
DuckDB temporary files
When a query needs more memory than its limit allows, DuckDB writes temporary (spill) files to disk. From 1.14, each DuckDB instance that BaseModel opens gets its own subdirectory for these files, so concurrent instances do not overwrite each other's files. The subdirectory is created under the first of:
temp_directoryinconfig_overrides, if set;- the
MONAD_DUCKDB_TEMP_DIRECTORYenvironment variable, if set; monad_duckdbin the system temporary directory ($TMPDIR, usually/tmp).
By default, spill files therefore go to the system temporary directory. If the container's /tmp is small, point temp_directory or MONAD_DUCKDB_TEMP_DIRECTORY at a volume with enough free space. Set temp_directory: "" in config_overrides to disable spilling.
Snowflake
| Parameter | Type | Default | Description |
|---|---|---|---|
user | str | required | Login name of the user. Supports environment variables (e.g., ${SNOWFLAKE_USER}). |
password | str | required | Password for the user. Supports environment variables. |
account | str | required | Snowflake account identifier. |
warehouse | str | required | Virtual warehouse to use. |
database | str | required | Default database. |
db_schema | str | required | Default schema (alias: schema). |
role | str | "PUBLIC" | Snowflake role to use for the session. |
data_location:
database_type: snowflake
connection_params:
user: "${SNOWFLAKE_USER}"
password: "${SNOWFLAKE_PASSWORD}"
account: "${SNOWFLAKE_ACCOUNT}"
warehouse: "${SNOWFLAKE_WAREHOUSE}"
database: "${SNOWFLAKE_DATABASE}"
schema: "${SNOWFLAKE_SCHEMA}"
role: "${SNOWFLAKE_ROLE}"
table_name: transactions
Set role explicitly
role defaults to PUBLIC, which in most Snowflake accounts holds no grants on the warehouse, database, or schema you need — the session then fails with permission errors. Set it to a role that already has the required privileges (for example via ${SNOWFLAKE_ROLE}).
OAuth Token Authentication (Snowflake Container Services)
For Snowflake Container Services, use the OAuth token-based configuration. This variant reads the token from /snowflake/session/token and uses SNOWFLAKE_ACCOUNT and SNOWFLAKE_HOST environment variables. Fields: warehouse, database, db_schema, authenticator (set to "oauth").
BigQuery
| Parameter | Type | Default | Description |
|---|---|---|---|
filename | Path | required | Path to the service account JSON file. |
project_id | str \| None | None | BigQuery project ID. Only needed if different from the one in the service account. |
data_location:
database_type: bigquery
connection_params:
filename: "/secrets/bigquery-user.json"
project_id: "my-bigquery-project"
table_name: transactions
schema_name: my_dataset
Databricks
| Parameter | Type | Default | Description |
|---|---|---|---|
host | str | required | Server hostname for your Databricks cluster or SQL warehouse. |
warehouse_id | str | required | SQL warehouse ID for your Databricks SQL warehouse. |
token | str \| None | None | Personal access token. Mutually exclusive with client_id/client_secret. |
client_id | str \| None | None | Service principal client ID. Requires client_secret. |
client_secret | str \| None | None | Service principal client secret. Requires client_id. |
catalog | str \| None | None | Unity Catalog name. |
db_schema | str | "default" | Schema name (alias: schema). |
extra_connect_params | dict | {} | Additional kwargs forwarded verbatim to databricks.sql.connect. Use for options such as _socket_timeout, use_cloud_fetch, session_configuration, or http_headers. Keys that collide with the explicit connection kwargs (server_hostname, http_path, access_token, catalog, schema, credentials_provider) raise a TypeError at connection time. |
Warning
token and client_id/client_secret are mutually exclusive. Use either personal access token or service principal authentication, not both.
Connection timeout
Set a socket-level timeout (in seconds) for the Databricks connection via extra_connect_params._socket_timeout:
# Personal access token authentication
data_location:
database_type: databricks
connection_params:
host: "${DATABRICKS_SERVER_HOSTNAME}"
warehouse_id: "${DATABRICKS_WAREHOUSE_ID}"
token: "${DATABRICKS_TOKEN}"
table_name: transactions
# Service principal authentication
data_location:
database_type: databricks
connection_params:
host: "${DATABRICKS_SERVER_HOSTNAME}"
warehouse_id: "${DATABRICKS_WAREHOUSE_ID}"
client_id: "${DATABRICKS_CLIENT_ID}"
client_secret: "${DATABRICKS_CLIENT_SECRET}"
table_name: transactions
Hive
| Parameter | Type | Default | Description |
|---|---|---|---|
hive_params | HiveParamsConfig | required | Connection parameters. Configure via DSN or Driver + Port + HiveServerType. |
ini_file | str \| None | $ODBCINI env var | Path to ODBC .ini file. Required when using DSN. |
kerberos_params | KerberosParamsConfig \| None | None | Kerberos authentication parameters. Required for Kerberos-secured clusters. |
HiveParamsConfig
| Parameter | Type | Default | Description |
|---|---|---|---|
DSN | str \| None | None | ODBC Data Source Name. |
Driver | str \| None | None | ODBC driver path. Required if DSN is not set. |
Port | int \| None | None | Hive server port. Required if DSN is not set. |
HiveServerType | int \| None | None | Hive server type (e.g., 2). Required if DSN is not set. |
KerberosParamsConfig
| Parameter | Type | Default | Description |
|---|---|---|---|
user | str | required | Kerberos principal name. |
kinit_realm | str | required | Kerberos realm. |
kerberos_host | str | required | Kerberos service host IP. |
kerberos_service_name | str | required | Kerberos service name (e.g., "hive"). |
kerberos_fqdn | str | required | Fully qualified domain name. |
keytab_path | str \| None | None | Path to the keytab file. Must be readable by BaseModel; a keytab that group or other can access logs a warning. |
krb5_config_path | str | "/etc/krb5.conf" | Path to krb5.conf. |
password | str \| None | None | Password in plain text. A value that names an existing file is rejected; use password_file instead. |
password_file | str \| None | None | Path to a file that holds the password. Works only with the Heimdal Kerberos client. The file must be readable by BaseModel and not accessible to group or other (chmod 600; on Kubernetes, defaultMode: 0400 on the secret volume). |
kerberos_renewal_interval_minutes | int | 540 | Ticket renewal interval in minutes. |
verbose | bool | False | Whether to print verbose output. |
Breaking change in 1.14: password files
In 1.13 and earlier, password could hold either the password or the path to a password file. From 1.14, a password value that names an existing file is rejected; move the path to password_file.
Kerberos credentials and cache
Set at most one of keytab_path, password, and password_file; setting more than one is a configuration error. If none is set, BaseModel reads the password from the KRB5_PASSWORD environment variable or the password-file path from KRB5_PASSWORD_FILE; set only one of the two.
BaseModel keeps the Kerberos credential cache in the directory named by the MONAD_KRB5_CACHE_DIR environment variable or, if it is not set, in monad_kerberos_ccaches_<uid> under the system temporary directory. BaseModel creates the directory if it does not exist. The directory must be owned by the user that runs BaseModel, must not be a symbolic link, and must not be accessible to group or other (chmod 700); otherwise the run fails with a configuration error.
# DSN-based configuration
data_location:
database_type: hive
connection_params:
hive_params:
DSN: "MyHiveDSN"
ini_file: "/etc/odbc.ini"
kerberos_params:
user: "hive_user@REALM"
kinit_realm: "REALM"
kerberos_host: "10.0.0.1"
kerberos_service_name: "hive"
kerberos_fqdn: "hive-server.example.com"
keytab_path: "/etc/security/keytabs/hive.keytab"
table_name: transactions
ClickHouse
| Parameter | Type | Default | Description |
|---|---|---|---|
host | str | required | Hostname or IP address of the ClickHouse server (e.g., clickhouse.example.com). This is not a connection URI. |
Supported versions
BaseModel checks the ClickHouse server version when it connects. Versions 22.8 through 25.8 are supported (any patch release); other versions fail with Clickhouse version: <version> is currently unsupported. In 1.13 and earlier, the CLICKHOUSE_VERSION environment variable had to be set; from 1.14 it is no longer read.
data_location:
database_type: clickhouse
connection_params:
host: "${CLICKHOUSE_HOST}"
table_name: transactions
Synapse
| Parameter | Type | Default | Description |
|---|---|---|---|
server_name | str | required | Azure Synapse server name. |
database_name | str | required | Dedicated SQL pool name. |
user | str | required | Username. Supports environment variables. |
password | str \| None | None | Password. Supports environment variables. |
data_location:
database_type: synapse
connection_params:
server_name: "my-synapse.sql.azuresynapse.net"
database_name: "my_pool"
user: "${SYNAPSE_USER}"
password: "${SYNAPSE_PASSWORD}"
table_name: transactions
Storage Configuration
The storage configuration file holds credentials and options for the cloud filesystems that BaseModel reads and writes through fsspec. Pass the file with --storage-config-path on the command line, or as storage_config_path in the Python API (for example in pretrain). It does not configure DuckDB: for Parquet path values on S3, see Reading Parquet from S3. It also does not let you keep the pretraining configuration YAML on object storage: the path passed to pretrain or --config-path must be a local or mounted file, because BaseModel copies it into the output directory with a plain file copy before reading it.
The file is YAML. Its top-level keys are storage protocols; include only the ones you use:
| Key | Storage | Keys inside the block |
|---|---|---|
s3 | Amazon S3 or S3-compatible storage | AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY (set both or neither); AWS_S3_ENDPOINT_URL (only for S3-compatible storage such as MinIO); anon (anonymous access to a public bucket). |
gs | Google Cloud Storage | token (path to a service-account JSON file, a gcsfs token identifier, or the token contents); project. |
az, abfs | Azure Blob Storage, Azure Data Lake Gen2 | One of: AZURE_STORAGE_CONNECTION_STRING; AZURE_STORAGE_ACCOUNT_NAME, optionally with AZURE_STORAGE_ACCOUNT_KEY, AZURE_STORAGE_SAS_TOKEN, or ANON; or all of AZURE_STORAGE_TENANT_ID, AZURE_STORAGE_CLIENT_ID, and AZURE_STORAGE_CLIENT_SECRET (service principal). They are tried in that order. AZURE_STORAGE_SAS_TOKEN is new in 1.14 and is used when no account key is given. |
adl | Azure Data Lake Gen1 | AZURE_STORAGE_TENANT_ID, AZURE_STORAGE_CLIENT_ID, AZURE_STORAGE_CLIENT_SECRET, and store_name. |
- Every key in the table except
store_namefalls back to the environment variable of the same name when it is missing from the block or written without a value. ${VAR}references in string values are replaced with the value of the environment variable when the file is loaded, as inconnection_params.- Any other key in a block is passed to the underlying
fsspecfilesystem unchanged.
s3:
AWS_ACCESS_KEY_ID: ${AWS_ACCESS_KEY_ID}
AWS_SECRET_ACCESS_KEY: ${AWS_SECRET_ACCESS_KEY}
az:
AZURE_STORAGE_ACCOUNT_NAME: my-account
AZURE_STORAGE_SAS_TOKEN: ${AZURE_STORAGE_SAS_TOKEN}
Breaking change in 1.14: Azure service-principal keys
In 1.13 and earlier, the Azure service-principal keys were AZURE_TENANT_ID, AZURE_CLIENT_ID, and AZURE_CLIENT_SECRET. From 1.14 they are AZURE_STORAGE_TENANT_ID, AZURE_STORAGE_CLIENT_ID, and AZURE_STORAGE_CLIENT_SECRET. Rename these keys in existing storage configuration files: an Azure block that has only the old names fails with an error that lists the new ones.
JSON Schema
From 1.14, the image ships a JSON Schema for this file, storage-config.schema.json, which editors that support JSON Schema for YAML can use for autocompletion and validation. To extract it from the image:
docker run --rm -e MONAD_LOGGING_DIRECTORY=/tmp/monad-logs \
--entrypoint python <monad-image> \
-c "import importlib.resources as r; print(r.files('monad.config').joinpath('schemas', 'storage-config.schema.json').read_text())" \
> storage-config.schema.json
Both options are needed: --entrypoint python skips the image's startup banner, which would otherwise end up in the output file, and MONAD_LOGGING_DIRECTORY points logging at a writable directory. BaseModel does not validate the storage configuration file against the schema at run time; the schema is an aid for writing the file.
To check a file from Python, load it with StorageConfig.from_yaml(), which raises a validation error if the file does not match the schema — for example, an unknown top-level key, or AWS_ACCESS_KEY_ID without AWS_SECRET_ACCESS_KEY:
from pathlib import Path
from monad.config import StorageConfig
StorageConfig.from_yaml(Path("storage-config.yaml"))
Column Selection
Control which columns are included in training using allowed_columns or disallowed_columns on any data source.
| Parameter | Type | Description |
|---|---|---|
allowed_columns | list[str] | Only these columns will be used. All others are excluded. |
disallowed_columns | list[str] | These columns are excluded. All others are used. |
Note
Do not place the same name in both lists. Otherwise, allowed_columns and disallowed_columns may be combined when an explicit fitted shortlist is used and a raw field must remain available only through data_params.extra_columns. Use disallowed_columns to drop PII, unique IDs, or columns that would cause data leakage.
# Drop columns that add no signal
disallowed_columns: ["order_id", "internal_id"]
# Or explicitly select columns
allowed_columns: ["customer_id", "article_id", "t_dat", "price", "sales_channel_id"]
Date Column Formats
The date_column block specifies the event timestamp column and its format.
| Format String | Example Value |
|---|---|
"%Y-%m-%d" | 2024-01-15 |
"%Y-%m-%d %H:%M:%S" | 2024-01-15 14:30:00 |
"%d/%m/%Y" | 15/01/2024 |
"s" | 1705312200 (UNIX seconds since epoch) |
"ms" | 1705312200000 (UNIX milliseconds since epoch) |
"ns" | 1705312200000000000 (UNIX nanoseconds since epoch) |
Date format depends on database engine
The format string syntax differs by database:
- Snowflake uses SQL-standard tokens:
YYYY-MM-DD,YYYY-MM-DD HH24:MI:SS - All other engines (Parquet, BigQuery, Databricks, ClickHouse, Hive) use Python strftime:
%Y-%m-%d,%Y-%m-%d %H:%M:%S
Using the wrong syntax causes silent data filtering failures or Can't parse date errors at training time.
Joining Tables
Use joined_data_sources on event data sources to join attribute (dimension) tables.
# On the event data source
joined_data_sources:
- name: articles # Name of the attribute data source
join_on:
- [article_id, article_id] # [event_column, attribute_column]
Multiple joins are supported:
joined_data_sources:
- name: articles
join_on:
- [article_id, article_id]
- name: stores
join_on:
- [store_id, store_id]
Tip
The attribute data source must be defined separately in data_sources with type: attribute. The join_on pairs map [event_table_column, attribute_table_column].