Skip to main content

Docker (CLI)

Docker CLI can be installed and executed using the official Docker images. Each supported source driver has a dedicated image available on the OLake DockerHub page.

Prerequisites​

  • Docker installed and running on the host system
  • Internet access to pull Docker images

In the subsequent commands, replace [SOURCE-TYPE] with the value corresponding to the required driver.

Driver Name[SOURCE-TYPE]
Postgrespostgres
MySQLmysql
Oracleoracle
MongoDBmongodb
MSSQLmssql
Kafkakafka
DB2db2
S3s3

With the driver identified, the next step is configuring the source system.

Sources​

Each source driver requires specific configuration. Detailed setup instructions can be found in the following guides:

What's Next : Once the source has been configured, the next step is to define the destination.

Destination​

Destinations specify where ingested data will be written. OLake Go supports multiple destinations. The following guides provide setup of:

What's Next : Once both source and destination are configured, the next step is to discover available streams/tables.

Discover Command​

  • To discover the available streams from the source, run the following command:

    discover command
    docker run --pull=always \
    -v "[PATH_OF_CONFIG_FOLDER]:/mnt/config" \
    olakego/source-[SOURCE-TYPE]:latest \
    discover \
    --config /mnt/config/source.json

    Running this command generates a streams.json file, which contains the list of available streams along with their metadata.

    Available flags​

    --config required
    --config /mnt/config/source.json

    Description:

    --streams optional
    --streams /mnt/config/streams.json

    Description:

    Specifies the path to the streams.json file. This file is generated after the discover command. When used during discovery, this flag updates the existing streams.json:

    • Keeps prior manual changes.
    • Adds new streams detected in the source database.
    • Allows selecting which columns should be synced for each table.
    • Allows updating the destination database name for each stream.

    ➡️ You must learn about streams.json configuration. Refer to the Streams Config guide.

    --destination-database-prefix optional
    --destination-database-prefix olake

    Description:

    • Adds a custom prefix to the database name created in the destination.
    • Example:

      If the source database is sales-db and the driver is mysql:

      • Default (Normalized) → mysql_sales_db
      • With prefix (Normalized) → olake_sales_db
    --timeout optional
    --timeout 600

    Description:

    • Applies only to the discover command.
    • Overrides the default timeout of 300 seconds (5 minutes).
    • This is helpful when working with large datasets or slower networks where the operation may need extra time to complete.
    --max-discover-threads optional
    --max-discover-threads 100

    Description:

    • Applies only to the discover command.
    • Sets the maximum number of parallel threads used for discovering table schemas in the database.
    • Value: Integer (mandatory when the flag is used). Minimum value is 1 (must be greater than 0).
    • Default: If the flag is not provided, the value defaults to 50.
    --difference optional
    --streams /path/old_streams.json \
    --difference /path/new_streams.json

    Description:

    • Used with the discover command to compare differences between two streams, specifically an old stream and a new stream.
    • Must be used together with --streams /path/old_streams.json to specify the path to the old streams file.
    • The --difference flag specifies the path to the new streams file that will be compared against the old one.
    • Running this command generates a difference_streams.json file containing the differences between the old and new streams.
    Internal testing command

    This command is primarily used internally for testing purposes to compare differences between old and new streams. It is not typically required for production use. However, if you want to check the differences between streams, you can use this command.

Use version specific command

Using a version specific image is recommended instead of using latest in the commands for production use case.

What's Next : Understanding the streams config file for stream/table level modifications

Streams Config​

The streams.json file is organized into two main sections:

  • selected_streams: Lists the streams that have been chosen for processing. These are grouped by namespace.
  • streams: Contains an array of stream definitions. Each stream holds details about its schema, supported synchronization modes, primary keys, and other metadata.

1. Selected Streams​

The selected_streams section groups streams by their namespace. For example, the configuration might look like this:

streams.json
"selected_streams": {
"my_db": [
{
"partition_regex": "/{dropoff_datetime, year}",
"stream_name": "table1",
"normalization": true,
"update_type": "eq",
"append_mode": false,
"selected_columns": {
"columns": [
"_cdc_timestamp",
"_olake_id",
"_cdc_lsn",
"_op_type",
"id",
"updated_at",
"dropoff_datetime",
"_olake_timestamp"
],
"sync_new_columns": true
},
"filter": "UPDATED_AT >= \"08-JUN-25 07.19.23.690870000 AM\""
},
{
"partition_regex": "",
"stream_name": "table2",
"normalization": false,
"append_mode": false,
"selected_columns": {
"columns": [
"_cdc_timestamp",
"_olake_id",
"_cdc_lsn",
"_op_type",
"id",
"city",
"_olake_timestamp"
],
"sync_new_columns": true
},
"filter": "city = \"London\""
}
]
}

Details about all the fields mentioned in selected streams​

ComponentTypeExample ValueDescription
namespacestringmy_dbGroups streams that belong to a specific database or logical category
stream_namestring"table1", "table2"Identifier for the stream. Must match the stream name defined in the Streams configurations.
partition_regexstring"/{dropoff_datetime, year}"A pattern defining how to partition the data. To read more, refer the Partition Regex for S3 or Partition Regex for iceberg
normalizationbooleantrueConverts top-level JSON fields into columns, while nested JSON objects are stored as string values.
update_typestring"eq" or "pos" or "dv"Specifies the format in which delete files are written during sync.
"eq" writes equality delete files.
"pos" writes positional delete files.
"dv" writes deletion vector files.
append_modebooleanfalseDisables upserts in Iceberg when set to true. (In upsert mode dedupe happens via equality delete)
selected_columns.columnsarray["_cdc_timestamp", "id", "updated_at"]List of columns to be included in the sync for each stream.
selected_columns.sync_new_columnsbooleantrueIf true, newly discovered columns are automatically included in future syncs. If false, new columns must be manually added to selected_columns.columns to be synced.
filterstring"UPDATED_AT >= \"08-JUN-25 07.19.23.690870000 AM\""Ensures that only data matching the specified condition is synced.

2. Streams​

The streams section is an array where each element is an object that defines a specific data stream. For example, one stream definition looks like this:

streams.json
{
"stream": {
"name": "table1",
"namespace": "my_db",
"type_schema": {
"properties": {
"_id": {
"type": ["string"],
"destination_column_name":"_id"
},
"name": {
"type": ["string"],
"destination_column_name":"name"
},
"marks": {
"type": ["integer"],
"destination_column_name":"marks"
},
"updated_at": {
"type": ["timestamp"],
"destination_column_name":"updated_at"
},
...
}
},
"supported_sync_modes": ["full_refresh", "cdc", "incremental", "strict_cdc"],
"source_defined_primary_key": ["_id"],
"available_cursor_fields": ["_id", "name", "marks", "updated_at"],
"sync_mode": "incremental",
"cursor_field": "updated_at",
"destination_database": "prefix:my_db",
"destination_table": "table1"
}
}

2.1 Stream Configuration Elements​

ComponentExample ValueDescription & Possible Values
name"stream_8"Unique identifier for the stream. Each stream must have a unique name.
namespace"olake_db"The grouping or database name that the stream belongs to.
type_schema(JSON object with properties)Defines the columns in streams or tables, including the destination column name. The source column name is normalized by replacing special characters and spaces with underscores (_) and is used as the final column name in the destination table.
supported_sync_modes["full_refresh", "cdc", "incremental","strict_cdc"]Lists the sync modes the stream supports. Includes "full_refresh", "cdc", "strict_cdc" and "incremental". To read more about sync modes read
source_defined_primary_key["_id"]Specifies the field(s) that is set as a primary key in the source.
available_cursor_fields["_id", "name", "marks", "updated_at"]Lists fields that can be used to track synchronization progress in incremental sync mode.
sync_mode"incremental"Indicates the active sync mode. Possible values are defined in supported_sync_modes.
cursor_field"updated_at"Defines the cursor field used to track incremental sync. To read more about incremental sync read this
destination_database"prefix:my_db"Specifies the destination database name where data will be stored. By default, the driver name is used as the prefix, but you can override it using the --destination-database-prefix flag. The source namespace is always appended to the database name. All special characters and spaces are replaced with (_), and the database is automatically created if it doesn’t exist.
destination_table"table1"Specifies the name of the destination table where the data will be stored. The table name is derived from the source stream name and normalized by replacing special characters and spaces with underscores (_).

Whats's Next: Once the streams.json file is configured as required, sync can be initiated.

Sync​

The Sync section describes how to start the sync process using the configuration files. It also explains how to run sync with a state file when operating in Incremental (Bookmark/cursor) or CDC mode.

Sync command​

This command is executed when a sync needs to be run in Full Refresh mode or when running a sync for the first time (i.e., before a state file has been created).

sync command
docker run --pull=always \
-v "[PATH_OF_CONFIG_FOLDER]:/mnt/config" \
olakego/source-[SOURCE-TYPE]:latest \
sync \
--config /mnt/config/source.json \
--streams /mnt/config/streams.json \
--destination /mnt/config/destination.json

Available flags​

--config required
--config /mnt/config/source.json

Description:

--streams required
--streams /mnt/config/streams.json

Description:

Specifies the path to the streams.json file. This file is generated after the discover command. When used during discovery, this flag updates the existing streams.json:

  • Keeps prior manual changes.
  • Adds new streams detected in the source database.
  • Allows selecting which columns should be synced for each table.
  • Allows updating the destination database name for each stream.

➡️ You must learn about streams.json configuration. Refer to the Streams Config guide.

--destination required
--state optional
--state /mnt/config/state.json

Description:

  • Specifies the path to the state file.
  • The state file contains metadata (such as offsets and positions) that enables:
    • Resuming interrupted syncs.
    • Continuing incremental or CDC syncs without restarting from scratch.
    • Storing the version in the state file for maintaining backward compatibility.

The state.json file is organized into two main sections:

1. Global State

The global section contains global state information that applies to all streams and driver-specific replication metadata that tracks the overall position in the source database's change log. The structure varies by database driver:

MySQL uses binlog position for global state tracking to maintain the replication position across all streams.

state.json (MySQL)
{
"type": "STREAM",
"version": 1,
"global": {
"state": {
"server_id": 261398335,
"state": {
"position": {
"Name": "mysql-bin.000070",
"Pos": 811746
}
}
},
"streams": [
"my_db.decimal_test",
"my_db.incr_test"
]
}
}

2. Streams State

The streams section is an array where each element tracks the synchronization state for a specific stream. The structure varies by database driver:

Each stream state object contains:

state.json
{
"stream": "table1",
"namespace": "my_db",
"sync_mode": "",
"state": {
"chunks": []
}
}

State Configuration Elements

ComponentTypeExample ValueDescription
versioninteger1 or 0Version 0 enables legacy, lenient handling for backward compatibility
Version 1 and above enforce stricter validation and fail-fast behavior for newly created state
globalobjectFor postgres :
"global": { "state" : { "lsn": "BD7/650015C8" }, "streams": [ "public.sample_data", "public.employees" ] }
Contains global replication metadata. Structure varies by driver: MySQL uses binlog position, PostgreSQL uses LSN, MongoDB does not have a global section.
global.state.server_idinteger261398335(MySQL only) The MySQL server ID used for replication tracking.
global.state.state.positionobject{"Name": "mysql-bin.000070", "Pos": 811746}(MySQL only) Tracks the current binlog file name and position for CDC replication.
global.state.lsnstring"BD7/650015C8"(PostgreSQL only) Log Sequence Number (LSN) that tracks the position in the PostgreSQL write-ahead log (WAL) for CDC replication.
global.streamsarray["public.decimal_test", "public.incr_test"]List of all streams that are being tracked in this state file.
streamstring"decimal_test", "incr_test"The name of the stream being tracked. Must match the stream name defined in streams.json.
namespacestring"my_db", "public"The namespace (database/schema) that the stream belongs to.
sync_modestring"", "incremental", "cdc"The synchronization mode being used for this stream. May be empty if not explicitly set.
state.chunksarray[]Array tracking data chunks that have been processed. Used for resuming partial syncs and managing large data transfers.
state._datastring"82696F0837000000012B0429296E1404"(MongoDB only) Resume token used to track the position in MongoDB's change stream for CDC replication.

What's Next: The state file is automatically created and updated during sync operations. You can manually specify a state file using the --state flag to resume from a previous synchronization point.

--discover-schema optional
--discover-schema

Description:

  • Applies only to the sync command.
  • By default, sync skips source schema discovery to keep runs fast and trusts the catalog passed (--streams /path/to/streams.json) from a prior discover run.
  • When this flag is passed, sync re-discovers source schema and validates configured streams against the source.

Whats's Next: Monitor sync progress and performance using the stats.json file.

Stats​

The stats.json file is created as soon as the sync starts. The file looks like this:

stats.json
{
"Writer Threads": 21,
"Synced Records": 14236868,
"Bytes Read": 9876543210,
"Speed": "76542.20 rps",
"Seconds Elapsed": "186.00",
"Estimated Remaining Time": "1642.00 s",
"Memory Usage Bytes": 2335248384,
"CPU Utilization": 0.9348
}
ComponenetExample ValueDescription
Estimated Remaining Time1642.00A rough estimation in seconds, indicating when the sync is expected to complete
Memory Usage Bytes2335248384System memory currently in use, reported in bytes
Writer Threads21Number of parallel writer threads actively syncing data
Seconds Elapsed186.00Time in seconds elapsed since the sync started
Speed76542.20 rpsHow many rows per second are being synced in this sync.
Synced Records14236868Total number of records which have been synced since the sync started.
Bytes Read9876543210Total source bytes read and committed to the destination since the sync started
CPU Utilization0.9348CPU usage during the sync, reported as a decimal ratio (for example, 0.9348 is about 93.48%)

The stats.json file remains available after sync completion.

What's Next: Similarly logs get saved for every command that is executed. And those can be reviewed for additional information.

Logs​

During a sync, OLake Go automatically creates an olake.log file inside a folder named sync_<TIMESTAMP_IN_UTC> (for example: sync_2025-02-17_11-41-40). This folder is created in the same directory as the configuration files.

  • The olake.log file contains a complete record of all logs generated while the command is running.

VS Code window displaying OLake Go logs for MongoDB sync and schema production events

  • Console logs are displayed in real time with color formatting for easier identification.

The log shows a fatal error: &quot;snappy: corrupt input&quot; occurred while reading Parquet records, indicating a decompression or data corruption issue

"debug": "\033[36m", // Cyan
"info": "\033[32m", // Green
"warn": "\033[33m", // Yellow
"error": "\033[31m", // Red
"fatal": "\033[31m", // Red

With this, the sync setup using the OLake CLI is complete, from configuring source and destination to reviewing key outputs such as stats and logs.

info

For a complete reference of all available commands and their flags (used to configure stream-level properties) see the Commands and Flags guide.




💡 Join the OLake Community!

Got questions, ideas, or just want to connect with other data engineers?
👉 Join our Slack Community to get real-time support, share feedback, and shape the future of OLake together. 🚀

Your success with OLake is our priority. Don’t hesitate to contact us if you need any help or further clarification!