Docker (CLI)
Docker CLI can be installed and executed using the official Docker images. Each supported source driver has a dedicated image available on the OLake DockerHub page.
Prerequisites
- Docker installed and running on the host system
- Internet access to pull Docker images
In the subsequent commands, replace [SOURCE-TYPE] with the value corresponding to the required driver.
| Driver Name | [SOURCE-TYPE] |
|---|---|
| Postgres | postgres |
| MySQL | mysql |
| Oracle | oracle |
| MongoDB | mongodb |
| MSSQL | mssql |
| Kafka | kafka |
| DB2 | db2 |
| S3 | s3 |
With the driver identified, the next step is configuring the source system.
Sources
Each source driver requires specific configuration. Detailed setup instructions can be found in the following guides:
- Working with MongoDB
- Working with MySQL
- Working with Oracle
- Working with Postgres
- Working with S3
- Working with DB2
- Working with MSSQL
- Working with Kafka
What's Next : Once the source has been configured, the next step is to define the destination.
Destination
Destinations specify where ingested data will be written. OLake Go supports multiple destinations. The following guides provide setup of:
- Working with AWS S3
- Working with Iceberg Glue Catalog
- Working with Iceberg Hive Catalog
- Working with Iceberg JDBC Catalog
- Working with Iceberg REST Catalog
- Working with Iceberg Lakekeeper Catalog
- Working with Iceberg Nessie Catalog
- Working with Iceberg S3 Tables Catalog
- Working with Iceberg Unity Catalog
- Working with Iceberg Polaris Catalog
- Working with Iceberg BigLake Catalog
- Working with Iceberg Horizon Catalog
What's Next : Once both source and destination are configured, the next step is to discover available streams/tables.
Discover Command
-
To discover the available streams from the source, run the following command:
discover commanddocker run --pull=always \
-v "[PATH_OF_CONFIG_FOLDER]:/mnt/config" \
olakego/source-[SOURCE-TYPE]:latest \
discover \
--config /mnt/config/source.jsonRunning this command generates a streams.json file, which contains the list of available streams along with their metadata.
Available flags
--configrequired--config /mnt/config/source.jsonDescription:
- Specifies the path to the source configuration file.
- For details about configuration files for different sources, see:
--streamsoptional--streams /mnt/config/streams.jsonDescription:
Specifies the path to the streams.json file. This file is generated after the discover command. When used during discovery, this flag updates the existing streams.json:
- Keeps prior manual changes.
- Adds new streams detected in the source database.
- Allows selecting which columns should be synced for each table.
- Allows updating the destination database name for each stream.
➡️ You must learn about
streams.jsonconfiguration. Refer to the Streams Config guide.--destination-database-prefixoptional--destination-database-prefix olakeDescription:
- Adds a custom prefix to the database name created in the destination.
- Example:
If the source database is
sales-dband the driver ismysql:- Default (Normalized) →
mysql_sales_db - With prefix (Normalized) →
olake_sales_db
- Default (Normalized) →
--timeoutoptional--timeout 600Description:
- Applies only to the discover command.
- Overrides the default timeout of 300 seconds (5 minutes).
- This is helpful when working with large datasets or slower networks where the operation may need extra time to complete.
--max-discover-threadsoptional--max-discover-threads 100Description:
- Applies only to the discover command.
- Sets the maximum number of parallel threads used for discovering table schemas in the database.
- Value: Integer (mandatory when the flag is used). Minimum value is 1 (must be greater than 0).
- Default: If the flag is not provided, the value defaults to 50.
--differenceoptional--streams /path/old_streams.json \
--difference /path/new_streams.jsonDescription:
- Used with the discover command to compare differences between two streams, specifically an old stream and a new stream.
- Must be used together with
--streams /path/old_streams.jsonto specify the path to the old streams file. - The
--differenceflag specifies the path to the new streams file that will be compared against the old one. - Running this command generates a
difference_streams.jsonfile containing the differences between the old and new streams.
Internal testing command
This command is primarily used internally for testing purposes to compare differences between old and new streams. It is not typically required for production use. However, if you want to check the differences between streams, you can use this command.
Using a version specific image is recommended instead of using latest in the commands for production use case.
What's Next : Understanding the streams config file for stream/table level modifications
Streams Config
The streams.json file is organized into two main sections:
selected_streams: Lists the streams that have been chosen for processing. These are grouped by namespace.streams: Contains an array of stream definitions. Each stream holds details about its schema, supported synchronization modes, primary keys, and other metadata.
1. Selected Streams
The selected_streams section groups streams by their namespace. For example, the configuration might look like this:
"selected_streams": {
"my_db": [
{
"partition_regex": "/{dropoff_datetime, year}",
"stream_name": "table1",
"normalization": true,
"update_type": "eq",
"append_mode": false,
"selected_columns": {
"columns": [
"_cdc_timestamp",
"_olake_id",
"_cdc_lsn",
"_op_type",
"id",
"updated_at",
"dropoff_datetime",
"_olake_timestamp"
],
"sync_new_columns": true
},
"filter": "UPDATED_AT >= \"08-JUN-25 07.19.23.690870000 AM\""
},
{
"partition_regex": "",
"stream_name": "table2",
"normalization": false,
"append_mode": false,
"selected_columns": {
"columns": [
"_cdc_timestamp",
"_olake_id",
"_cdc_lsn",
"_op_type",
"id",
"city",
"_olake_timestamp"
],
"sync_new_columns": true
},
"filter": "city = \"London\""
}
]
}
Details about all the fields mentioned in selected streams
| Component | Type | Example Value | Description |
|---|---|---|---|
namespace | string | my_db | Groups streams that belong to a specific database or logical category |
stream_name | string | "table1", "table2" | Identifier for the stream. Must match the stream name defined in the Streams configurations. |
partition_regex | string | "/{dropoff_datetime, year}" | A pattern defining how to partition the data. To read more, refer the Partition Regex for S3 or Partition Regex for iceberg |
normalization | boolean | true | Converts top-level JSON fields into columns, while nested JSON objects are stored as string values. |
update_type | string | "eq" or "pos" or "dv" | Specifies the format in which delete files are written during sync. "eq" writes equality delete files. "pos" writes positional delete files. "dv" writes deletion vector files. |
append_mode | boolean | false | Disables upserts in Iceberg when set to true. (In upsert mode dedupe happens via equality delete) |
selected_columns.columns | array | ["_cdc_timestamp", "id", "updated_at"] | List of columns to be included in the sync for each stream. |
selected_columns.sync_new_columns | boolean | true | If true, newly discovered columns are automatically included in future syncs. If false, new columns must be manually added to selected_columns.columns to be synced. |
filter | string | "UPDATED_AT >= \"08-JUN-25 07.19.23.690870000 AM\"" | Ensures that only data matching the specified condition is synced. |
2. Streams
The streams section is an array where each element is an object that defines a specific data stream. For example, one stream definition looks like this:
{
"stream": {
"name": "table1",
"namespace": "my_db",
"type_schema": {
"properties": {
"_id": {
"type": ["string"],
"destination_column_name":"_id"
},
"name": {
"type": ["string"],
"destination_column_name":"name"
},
"marks": {
"type": ["integer"],
"destination_column_name":"marks"
},
"updated_at": {
"type": ["timestamp"],
"destination_column_name":"updated_at"
},
...
}
},
"supported_sync_modes": ["full_refresh", "cdc", "incremental", "strict_cdc"],
"source_defined_primary_key": ["_id"],
"available_cursor_fields": ["_id", "name", "marks", "updated_at"],
"sync_mode": "incremental",
"cursor_field": "updated_at",
"destination_database": "prefix:my_db",
"destination_table": "table1"
}
}
2.1 Stream Configuration Elements
| Component | Example Value | Description & Possible Values |
|---|---|---|
name | "stream_8" | Unique identifier for the stream. Each stream must have a unique name. |
namespace | "olake_db" | The grouping or database name that the stream belongs to. |
type_schema | (JSON object with properties) | Defines the columns in streams or tables, including the destination column name. The source column name is normalized by replacing special characters and spaces with underscores (_) and is used as the final column name in the destination table. |
supported_sync_modes | ["full_refresh", "cdc", "incremental","strict_cdc"] | Lists the sync modes the stream supports. Includes "full_refresh", "cdc", "strict_cdc" and "incremental". To read more about sync modes read |
source_defined_primary_key | ["_id"] | Specifies the field(s) that is set as a primary key in the source. |
available_cursor_fields | ["_id", "name", "marks", "updated_at"] | Lists fields that can be used to track synchronization progress in incremental sync mode. |
sync_mode | "incremental" | Indicates the active sync mode. Possible values are defined in supported_sync_modes. |
cursor_field | "updated_at" | Defines the cursor field used to track incremental sync. To read more about incremental sync read this |
destination_database | "prefix:my_db" | Specifies the destination database name where data will be stored. By default, the driver name is used as the prefix, but you can override it using the --destination-database-prefix flag. The source namespace is always appended to the database name. All special characters and spaces are replaced with (_), and the database is automatically created if it doesn’t exist. |
destination_table | "table1" | Specifies the name of the destination table where the data will be stored. The table name is derived from the source stream name and normalized by replacing special characters and spaces with underscores (_). |
Whats's Next: Once the streams.json file is configured as required, sync can be initiated.
Sync
The Sync section describes how to start the sync process using the configuration files. It also explains how to run sync with a state file when operating in Incremental (Bookmark/cursor) or CDC mode.
Sync command
This command is executed when a sync needs to be run in Full Refresh mode or when running a sync for the first time (i.e., before a state file has been created).
docker run --pull=always \
-v "[PATH_OF_CONFIG_FOLDER]:/mnt/config" \
olakego/source-[SOURCE-TYPE]:latest \
sync \
--config /mnt/config/source.json \
--streams /mnt/config/streams.json \
--destination /mnt/config/destination.json
Available flags
--config required
--config /mnt/config/source.json
Description:
- Specifies the path to the source configuration file.
- For details about configuration files for different sources, see:
--streams required
--streams /mnt/config/streams.json
Description:
Specifies the path to the streams.json file. This file is generated after the discover command. When used during discovery, this flag updates the existing streams.json:
- Keeps prior manual changes.
- Adds new streams detected in the source database.
- Allows selecting which columns should be synced for each table.
- Allows updating the destination database name for each stream.
➡️ You must learn about streams.json configuration. Refer to the Streams Config guide.
--destination required
--destination /mnt/config/destination.json
Description:
- Specifies the path to the destination configuration file.
- For details about destination configuration files, see:
- AWS Glue Catalog configuration
- Generic REST Catalog configuration
- Lakekeeper Catalog configuration
- Nessie Catalog configuration
- S3 Tables Catalog configuration
- Unity Catalog configuration
- Apache Polaris Catalog configuration
- BigLake Catalog configuration
- JDBC Catalog configuration
- Hive Catalog configuration
- Parquet configuration
--state optional
--state /mnt/config/state.json
Description:
- Specifies the path to the state file.
- The state file contains metadata (such as offsets and positions) that enables:
- Resuming interrupted syncs.
- Continuing incremental or CDC syncs without restarting from scratch.
- Storing the version in the state file for maintaining backward compatibility.
The state.json file is organized into two main sections:
1. Global State
The global section contains global state information that applies to all streams and driver-specific replication metadata that tracks the overall position in the source database's change log. The structure varies by database driver:
- MySQL
- PostgreSQL
- MongoDB
MySQL uses binlog position for global state tracking to maintain the replication position across all streams.
{
"type": "STREAM",
"version": 1,
"global": {
"state": {
"server_id": 261398335,
"state": {
"position": {
"Name": "mysql-bin.000070",
"Pos": 811746
}
}
},
"streams": [
"my_db.decimal_test",
"my_db.incr_test"
]
}
}
PostgreSQL uses LSN (Log Sequence Number) for global state tracking to maintain the replication position across all streams.
{
"type": "STREAM",
"version": 1,
"global": {
"state": {
"lsn": "BD7/650015C8"
},
"streams": [
"public.sample_data",
"public.employees"
]
}
}
MongoDB does not use a global state section. The state is maintained at the stream level only.
2. Streams State
The streams section is an array where each element tracks the synchronization state for a specific stream. The structure varies by database driver:
- MySQL & PostgreSQL
- MongoDB
Each stream state object contains:
{
"stream": "table1",
"namespace": "my_db",
"sync_mode": "",
"state": {
"chunks": []
}
}
MongoDB stream state includes a _data field that stores the resume token:
{
"stream": "users",
"namespace": "public",
"sync_mode": "",
"state": {
"_data": "82696F0837000000012B0429296E1404",
"chunks": []
}
}
State Configuration Elements
| Component | Type | Example Value | Description |
|---|---|---|---|
version | integer | 1 or 0 | Version 0 enables legacy, lenient handling for backward compatibility Version 1 and above enforce stricter validation and fail-fast behavior for newly created state |
global | object | For postgres : "global": { "state" : { "lsn": "BD7/650015C8" }, "streams": [ "public.sample_data", "public.employees" ] } | Contains global replication metadata. Structure varies by driver: MySQL uses binlog position, PostgreSQL uses LSN, MongoDB does not have a global section. |
global.state.server_id | integer | 261398335 | (MySQL only) The MySQL server ID used for replication tracking. |
global.state.state.position | object | {"Name": "mysql-bin.000070", "Pos": 811746} | (MySQL only) Tracks the current binlog file name and position for CDC replication. |
global.state.lsn | string | "BD7/650015C8" | (PostgreSQL only) Log Sequence Number (LSN) that tracks the position in the PostgreSQL write-ahead log (WAL) for CDC replication. |
global.streams | array | ["public.decimal_test", "public.incr_test"] | List of all streams that are being tracked in this state file. |
stream | string | "decimal_test", "incr_test" | The name of the stream being tracked. Must match the stream name defined in streams.json. |
namespace | string | "my_db", "public" | The namespace (database/schema) that the stream belongs to. |
sync_mode | string | "", "incremental", "cdc" | The synchronization mode being used for this stream. May be empty if not explicitly set. |
state.chunks | array | [] | Array tracking data chunks that have been processed. Used for resuming partial syncs and managing large data transfers. |
state._data | string | "82696F0837000000012B0429296E1404" | (MongoDB only) Resume token used to track the position in MongoDB's change stream for CDC replication. |
What's Next: The state file is automatically created and updated during sync operations. You can manually specify a state file using the --state flag to resume from a previous synchronization point.
--discover-schema optional
--discover-schema
Description:
- Applies only to the sync command.
- By default, sync skips source schema discovery to keep runs fast and trusts the catalog passed (
--streams /path/to/streams.json) from a prior discover run. - When this flag is passed, sync re-discovers source schema and validates configured streams against the source.
Whats's Next: Monitor sync progress and performance using the stats.json file.
Stats
The stats.json file is created as soon as the sync starts. The file looks like this:
{
"Writer Threads": 21,
"Synced Records": 14236868,
"Bytes Read": 9876543210,
"Speed": "76542.20 rps",
"Seconds Elapsed": "186.00",
"Estimated Remaining Time": "1642.00 s",
"Memory Usage Bytes": 2335248384,
"CPU Utilization": 0.9348
}
| Componenet | Example Value | Description |
|---|---|---|
| Estimated Remaining Time | 1642.00 | A rough estimation in seconds, indicating when the sync is expected to complete |
| Memory Usage Bytes | 2335248384 | System memory currently in use, reported in bytes |
| Writer Threads | 21 | Number of parallel writer threads actively syncing data |
| Seconds Elapsed | 186.00 | Time in seconds elapsed since the sync started |
| Speed | 76542.20 rps | How many rows per second are being synced in this sync. |
| Synced Records | 14236868 | Total number of records which have been synced since the sync started. |
| Bytes Read | 9876543210 | Total source bytes read and committed to the destination since the sync started |
| CPU Utilization | 0.9348 | CPU usage during the sync, reported as a decimal ratio (for example, 0.9348 is about 93.48%) |
The stats.json file remains available after sync completion.
What's Next: Similarly logs get saved for every command that is executed. And those can be reviewed for additional information.
Logs
During a sync, OLake Go automatically creates an olake.log file inside a folder named sync_<TIMESTAMP_IN_UTC> (for example: sync_2025-02-17_11-41-40).
This folder is created in the same directory as the configuration files.
- The olake.log file contains a complete record of all logs generated while the command is running.

- Console logs are displayed in real time with color formatting for easier identification.

"debug": "\033[36m", // Cyan
"info": "\033[32m", // Green
"warn": "\033[33m", // Yellow
"error": "\033[31m", // Red
"fatal": "\033[31m", // Red
With this, the sync setup using the OLake CLI is complete, from configuring source and destination to reviewing key outputs such as stats and logs.
For a complete reference of all available commands and their flags (used to configure stream-level properties) see the Commands and Flags guide.