Skip to content

Data Connectors

BaseModel reads data from multiple database types. This page documents the connection_params block for each supported database within the YAML configuration.

Overview

Database database_type Required fields
Parquet parquet path
Snowflake snowflake user, password, account, warehouse, database, db_schema
BigQuery bigquery filename
Databricks databricks host, warehouse_id, (token or client_id + client_secret)
Hive hive hive_params
ClickHouse clickhouse host
Synapse synapse server_name, database_name, user

Parquet

Read Parquet files from a local or mounted filesystem, or directly from S3. Supports single files, directories, and glob patterns.

Parameter Type Default Description
path Path \| str required Path to the Parquet file or directory. Glob patterns supported (e.g., *.parquet). From 1.14, an s3:// URI is accepted too — see Reading Parquet from S3.
cache_path Path \| None None Database cache path. Can be reused between trainings for faster restarts.
config_overrides dict {} DuckDB connection overrides (e.g., {"max_memory": "100GB"}). See DuckDB configuration reference and DuckDB temporary files.
data_location:
  database_type: parquet
  connection_params:
    path: "/data/transactions.parquet"
    cache_path: "/basemodel/db_cache/"
  table_name: transactions

Glob pattern example:

connection_params:
  path: "/data/transactions/*.parquet"
  cache_path: "/basemodel/db_cache/"

Reading Parquet from S3

From 1.14, path accepts an s3:// URI. BaseModel passes the URI to DuckDB unchanged, and DuckDB reads the objects through its httpfs extension.

data_location:
  database_type: parquet
  connection_params:
    path: "s3://my-bucket/transactions/*.parquet"
  table_name: transactions
  • URI form. Write the URI as a plain string with two slashes after the scheme: s3://bucket/key. A value in which the double slash has collapsed to one (s3:/bucket/key, which is what wrapping the URI in a Python pathlib.Path produces) is rejected with a validation error. gs:// and az:// URIs are kept as they are and passed to DuckDB in the same way; this section covers S3 only.
  • Credentials. BaseModel does not pass S3 credentials to DuckDB. The httpfs extension reads them from the environment variables AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN, and AWS_DEFAULT_REGION. config_overrides is passed to DuckDB unchanged, so DuckDB S3 settings such as s3_endpoint (for S3-compatible storage) can be set there. The storage configuration file does not apply to Parquet path.
  • Caching. Without cache_path, queries read straight from the bucket. With cache_path, an s3:// source is loaded into the local cache once and reused on later runs; BaseModel does not check the bucket for changes, so new or modified objects are ignored. Delete the cache_path directory to force a reload.

Reading s3:// paths requires outbound internet access

The BaseModel image does not include the DuckDB httpfs extension. DuckDB downloads it from the DuckDB extension repository (extensions.duckdb.org) the first time it needs it. Without outbound access to that repository, reading an s3:// path fails because the extension cannot be installed; setting a DuckDB S3 option in config_overrides fails the same way. In such environments, mount the bucket as a filesystem and point path at the mounted files.

DuckDB temporary files

When a query needs more memory than its limit allows, DuckDB writes temporary (spill) files to disk. From 1.14, each DuckDB instance that BaseModel opens gets its own subdirectory for these files, so concurrent instances do not overwrite each other's files. The subdirectory is created under the first of:

  1. temp_directory in config_overrides, if set;
  2. the MONAD_DUCKDB_TEMP_DIRECTORY environment variable, if set;
  3. monad_duckdb in the system temporary directory ($TMPDIR, usually /tmp).

By default, spill files therefore go to the system temporary directory. If the container's /tmp is small, point temp_directory or MONAD_DUCKDB_TEMP_DIRECTORY at a volume with enough free space. Set temp_directory: "" in config_overrides to disable spilling.


Snowflake

Parameter Type Default Description
user str required Login name of the user. Supports environment variables (e.g., ${SNOWFLAKE_USER}).
password str required Password for the user. Supports environment variables.
account str required Snowflake account identifier.
warehouse str required Virtual warehouse to use.
database str required Default database.
db_schema str required Default schema (alias: schema).
role str "PUBLIC" Snowflake role to use for the session.
data_location:
  database_type: snowflake
  connection_params:
    user: "${SNOWFLAKE_USER}"
    password: "${SNOWFLAKE_PASSWORD}"
    account: "${SNOWFLAKE_ACCOUNT}"
    warehouse: "${SNOWFLAKE_WAREHOUSE}"
    database: "${SNOWFLAKE_DATABASE}"
    schema: "${SNOWFLAKE_SCHEMA}"
    role: "${SNOWFLAKE_ROLE}"
  table_name: transactions

Set role explicitly

role defaults to PUBLIC, which in most Snowflake accounts holds no grants on the warehouse, database, or schema you need — the session then fails with permission errors. Set it to a role that already has the required privileges (for example via ${SNOWFLAKE_ROLE}).

OAuth Token Authentication (Snowflake Container Services)

For Snowflake Container Services, use the OAuth token-based configuration. This variant reads the token from /snowflake/session/token and uses SNOWFLAKE_ACCOUNT and SNOWFLAKE_HOST environment variables. Fields: warehouse, database, db_schema, authenticator (set to "oauth").


BigQuery

Parameter Type Default Description
filename Path required Path to the service account JSON file.
project_id str \| None None BigQuery project ID. Only needed if different from the one in the service account.
data_location:
  database_type: bigquery
  connection_params:
    filename: "/secrets/bigquery-user.json"
    project_id: "my-bigquery-project"
  table_name: transactions
  schema_name: my_dataset

Databricks

Parameter Type Default Description
host str required Server hostname for your Databricks cluster or SQL warehouse.
warehouse_id str required SQL warehouse ID for your Databricks SQL warehouse.
token str \| None None Personal access token. Mutually exclusive with client_id/client_secret.
client_id str \| None None Service principal client ID. Requires client_secret.
client_secret str \| None None Service principal client secret. Requires client_id.
catalog str \| None None Unity Catalog name.
db_schema str "default" Schema name (alias: schema).
extra_connect_params dict {} Additional kwargs forwarded verbatim to databricks.sql.connect. Use for options such as _socket_timeout, use_cloud_fetch, session_configuration, or http_headers. Keys that collide with the explicit connection kwargs (server_hostname, http_path, access_token, catalog, schema, credentials_provider) raise a TypeError at connection time.

Warning

token and client_id/client_secret are mutually exclusive. Use either personal access token or service principal authentication, not both.

Connection timeout

Set a socket-level timeout (in seconds) for the Databricks connection via extra_connect_params._socket_timeout:

data_location:
  database_type: databricks
  connection_params:
    host: "${DATABRICKS_SERVER_HOSTNAME}"
    warehouse_id: "${DATABRICKS_WAREHOUSE_ID}"
    token: "${DATABRICKS_TOKEN}"
    extra_connect_params:
      _socket_timeout: 300   # seconds, passed straight to databricks.sql.connect
  table_name: transactions
# Personal access token authentication
data_location:
  database_type: databricks
  connection_params:
    host: "${DATABRICKS_SERVER_HOSTNAME}"
    warehouse_id: "${DATABRICKS_WAREHOUSE_ID}"
    token: "${DATABRICKS_TOKEN}"
  table_name: transactions
# Service principal authentication
data_location:
  database_type: databricks
  connection_params:
    host: "${DATABRICKS_SERVER_HOSTNAME}"
    warehouse_id: "${DATABRICKS_WAREHOUSE_ID}"
    client_id: "${DATABRICKS_CLIENT_ID}"
    client_secret: "${DATABRICKS_CLIENT_SECRET}"
  table_name: transactions

Hive

Parameter Type Default Description
hive_params HiveParamsConfig required Connection parameters. Configure via DSN or Driver + Port + HiveServerType.
ini_file str \| None $ODBCINI env var Path to ODBC .ini file. Required when using DSN.
kerberos_params KerberosParamsConfig \| None None Kerberos authentication parameters. Required for Kerberos-secured clusters.

HiveParamsConfig

Parameter Type Default Description
DSN str \| None None ODBC Data Source Name.
Driver str \| None None ODBC driver path. Required if DSN is not set.
Port int \| None None Hive server port. Required if DSN is not set.
HiveServerType int \| None None Hive server type (e.g., 2). Required if DSN is not set.

KerberosParamsConfig

Parameter Type Default Description
user str required Kerberos principal name.
kinit_realm str required Kerberos realm.
kerberos_host str required Kerberos service host IP.
kerberos_service_name str required Kerberos service name (e.g., "hive").
kerberos_fqdn str required Fully qualified domain name.
keytab_path str \| None None Path to the keytab file. Must be readable by BaseModel; a keytab that group or other can access logs a warning.
krb5_config_path str "/etc/krb5.conf" Path to krb5.conf.
password str \| None None Password in plain text. A value that names an existing file is rejected; use password_file instead.
password_file str \| None None Path to a file that holds the password. Works only with the Heimdal Kerberos client. The file must be readable by BaseModel and not accessible to group or other (chmod 600; on Kubernetes, defaultMode: 0400 on the secret volume).
kerberos_renewal_interval_minutes int 540 Ticket renewal interval in minutes.
verbose bool False Whether to print verbose output.

Breaking change in 1.14: password files

In 1.13 and earlier, password could hold either the password or the path to a password file. From 1.14, a password value that names an existing file is rejected; move the path to password_file.

Kerberos credentials and cache

Set at most one of keytab_path, password, and password_file; setting more than one is a configuration error. If none is set, BaseModel reads the password from the KRB5_PASSWORD environment variable or the password-file path from KRB5_PASSWORD_FILE; set only one of the two.

BaseModel keeps the Kerberos credential cache in the directory named by the MONAD_KRB5_CACHE_DIR environment variable or, if it is not set, in monad_kerberos_ccaches_<uid> under the system temporary directory. BaseModel creates the directory if it does not exist. The directory must be owned by the user that runs BaseModel, must not be a symbolic link, and must not be accessible to group or other (chmod 700); otherwise the run fails with a configuration error.

# DSN-based configuration
data_location:
  database_type: hive
  connection_params:
    hive_params:
      DSN: "MyHiveDSN"
    ini_file: "/etc/odbc.ini"
    kerberos_params:
      user: "hive_user@REALM"
      kinit_realm: "REALM"
      kerberos_host: "10.0.0.1"
      kerberos_service_name: "hive"
      kerberos_fqdn: "hive-server.example.com"
      keytab_path: "/etc/security/keytabs/hive.keytab"
  table_name: transactions

ClickHouse

Parameter Type Default Description
host str required Hostname or IP address of the ClickHouse server (e.g., clickhouse.example.com). This is not a connection URI.

Supported versions

BaseModel checks the ClickHouse server version when it connects. Versions 22.8 through 25.8 are supported (any patch release); other versions fail with Clickhouse version: <version> is currently unsupported. In 1.13 and earlier, the CLICKHOUSE_VERSION environment variable had to be set; from 1.14 it is no longer read.

data_location:
  database_type: clickhouse
  connection_params:
    host: "${CLICKHOUSE_HOST}"
  table_name: transactions

Synapse

Parameter Type Default Description
server_name str required Azure Synapse server name.
database_name str required Dedicated SQL pool name.
user str required Username. Supports environment variables.
password str \| None None Password. Supports environment variables.
data_location:
  database_type: synapse
  connection_params:
    server_name: "my-synapse.sql.azuresynapse.net"
    database_name: "my_pool"
    user: "${SYNAPSE_USER}"
    password: "${SYNAPSE_PASSWORD}"
  table_name: transactions

Storage Configuration

The storage configuration file holds credentials and options for the cloud filesystems that BaseModel reads and writes through fsspec. Pass the file with --storage-config-path on the command line, or as storage_config_path in the Python API (for example in pretrain). It does not configure DuckDB: for Parquet path values on S3, see Reading Parquet from S3. It also does not let you keep the pretraining configuration YAML on object storage: the path passed to pretrain or --config-path must be a local or mounted file, because BaseModel copies it into the output directory with a plain file copy before reading it.

The file is YAML. Its top-level keys are storage protocols; include only the ones you use:

Key Storage Keys inside the block
s3 Amazon S3 or S3-compatible storage AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY (set both or neither); AWS_S3_ENDPOINT_URL (only for S3-compatible storage such as MinIO); anon (anonymous access to a public bucket).
gs Google Cloud Storage token (path to a service-account JSON file, a gcsfs token identifier, or the token contents); project.
az, abfs Azure Blob Storage, Azure Data Lake Gen2 One of: AZURE_STORAGE_CONNECTION_STRING; AZURE_STORAGE_ACCOUNT_NAME, optionally with AZURE_STORAGE_ACCOUNT_KEY, AZURE_STORAGE_SAS_TOKEN, or ANON; or all of AZURE_STORAGE_TENANT_ID, AZURE_STORAGE_CLIENT_ID, and AZURE_STORAGE_CLIENT_SECRET (service principal). They are tried in that order. AZURE_STORAGE_SAS_TOKEN is new in 1.14 and is used when no account key is given.
adl Azure Data Lake Gen1 AZURE_STORAGE_TENANT_ID, AZURE_STORAGE_CLIENT_ID, AZURE_STORAGE_CLIENT_SECRET, and store_name.
  • Every key in the table except store_name falls back to the environment variable of the same name when it is missing from the block or written without a value.
  • ${VAR} references in string values are replaced with the value of the environment variable when the file is loaded, as in connection_params.
  • Any other key in a block is passed to the underlying fsspec filesystem unchanged.
s3:
  AWS_ACCESS_KEY_ID: ${AWS_ACCESS_KEY_ID}
  AWS_SECRET_ACCESS_KEY: ${AWS_SECRET_ACCESS_KEY}

az:
  AZURE_STORAGE_ACCOUNT_NAME: my-account
  AZURE_STORAGE_SAS_TOKEN: ${AZURE_STORAGE_SAS_TOKEN}

Breaking change in 1.14: Azure service-principal keys

In 1.13 and earlier, the Azure service-principal keys were AZURE_TENANT_ID, AZURE_CLIENT_ID, and AZURE_CLIENT_SECRET. From 1.14 they are AZURE_STORAGE_TENANT_ID, AZURE_STORAGE_CLIENT_ID, and AZURE_STORAGE_CLIENT_SECRET. Rename these keys in existing storage configuration files: an Azure block that has only the old names fails with an error that lists the new ones.

JSON Schema

From 1.14, the image ships a JSON Schema for this file, storage-config.schema.json, which editors that support JSON Schema for YAML can use for autocompletion and validation. To extract it from the image:

docker run --rm -e MONAD_LOGGING_DIRECTORY=/tmp/monad-logs \
  --entrypoint python <monad-image> \
  -c "import importlib.resources as r; print(r.files('monad.config').joinpath('schemas', 'storage-config.schema.json').read_text())" \
  > storage-config.schema.json

Both options are needed: --entrypoint python skips the image's startup banner, which would otherwise end up in the output file, and MONAD_LOGGING_DIRECTORY points logging at a writable directory. BaseModel does not validate the storage configuration file against the schema at run time; the schema is an aid for writing the file.

To check a file from Python, load it with StorageConfig.from_yaml(), which raises a validation error if the file does not match the schema — for example, an unknown top-level key, or AWS_ACCESS_KEY_ID without AWS_SECRET_ACCESS_KEY:

Python
from pathlib import Path
from monad.config import StorageConfig

StorageConfig.from_yaml(Path("storage-config.yaml"))

Column Selection

Control which columns are included in training using allowed_columns or disallowed_columns on any data source.

Parameter Type Description
allowed_columns list[str] Only these columns will be used. All others are excluded.
disallowed_columns list[str] These columns are excluded. All others are used.

Note

Do not place the same name in both lists. Otherwise, allowed_columns and disallowed_columns may be combined when an explicit fitted shortlist is used and a raw field must remain available only through data_params.extra_columns. Use disallowed_columns to drop PII, unique IDs, or columns that would cause data leakage.

# Drop columns that add no signal
disallowed_columns: ["order_id", "internal_id"]

# Or explicitly select columns
allowed_columns: ["customer_id", "article_id", "t_dat", "price", "sales_channel_id"]

Date Column Formats

The date_column block specifies the event timestamp column and its format.

Format String Example Value
"%Y-%m-%d" 2024-01-15
"%Y-%m-%d %H:%M:%S" 2024-01-15 14:30:00
"%d/%m/%Y" 15/01/2024
"s" 1705312200 (UNIX seconds since epoch)
"ms" 1705312200000 (UNIX milliseconds since epoch)
"ns" 1705312200000000000 (UNIX nanoseconds since epoch)
date_column:
  name: t_dat
  format: "%Y-%m-%d"

Date format depends on database engine

The format string syntax differs by database:

  • Snowflake uses SQL-standard tokens: YYYY-MM-DD, YYYY-MM-DD HH24:MI:SS
  • All other engines (Parquet, BigQuery, Databricks, ClickHouse, Hive) use Python strftime: %Y-%m-%d, %Y-%m-%d %H:%M:%S

Using the wrong syntax causes silent data filtering failures or Can't parse date errors at training time.


Joining Tables

Use joined_data_sources on event data sources to join attribute (dimension) tables.

# On the event data source
joined_data_sources:
  - name: articles           # Name of the attribute data source
    join_on:
      - [article_id, article_id]  # [event_column, attribute_column]

Multiple joins are supported:

joined_data_sources:
  - name: articles
    join_on:
      - [article_id, article_id]
  - name: stores
    join_on:
      - [store_id, store_id]

Tip

The attribute data source must be defined separately in data_sources with type: attribute. The join_on pairs map [event_table_column, attribute_table_column].