Skip to main content
When RisingWave streams data into Iceberg, it generates many small data files, fragmented manifests, and frequent snapshots. Over time, this can degrade query performance and increase storage costs. To maintain healthy tables, RisingWave provides both automatic and manual maintenance features for compaction, manifest rewrite, and snapshot expiration.
  • Compaction: Merges small data files and delete files into larger, optimized files to improve read performance.
  • Manifest Rewrite: Merges fragmented data manifests into fewer, better-sized manifests to reduce metadata overhead during planning.
  • Snapshot Expiration: Removes old, unneeded snapshots and their associated data files to reclaim storage space.

Automatic maintenance

Version notes
  • RisingWave introduced Iceberg automatic maintenance (default) in v2.5.0 and added compaction.type with the small-files and files-with-delete compaction types in v2.7.0.
  • Newly created internal Iceberg tables (CREATE TABLE ... ENGINE = iceberg) using Iceberg format V2 enable enable_manifest_rewrite by default. External Iceberg sinks remain opt-in, and Iceberg format V3 does not support manifest rewrite.
  • compaction.write_parquet_compression and compaction.write_parquet_max_row_group_rows were added in v2.9.0. The compaction.target_file_size_mb parameter now also controls the output file size for sink writes (previously only applied to compaction).
  • Parameters prefixed with compaction are currently in technical preview stage and may change in future releases.
You can enable automatic maintenance to run periodically in the background for your Iceberg sinks and internal tables. Periodic manifest rewrite and snapshot expiration run in the meta-node maintenance loop, rather than using the per-sink compaction_interval_sec parameter. Configure the interval with iceberg_gc_interval_sec in the [meta] section of risingwave.toml (default: 3600 seconds).
Dedicated compactor required for data file compactionAutomatic data file compaction requires a dedicated compactor node. Before enabling enable_compaction = true, ensure your cluster has at least one compactor node deployed. Manifest rewrite and snapshot expiration run on the meta node and do not require a dedicated compactor. For deployment instructions and sizing guidelines, see Deploy a dedicated Iceberg compactor.

Compaction triggers

When automatic data file compaction is enabled, a new compaction becomes eligible when either condition is met:
  • The number of pending snapshot commits reaches compaction.trigger_snapshot_count, even if the time interval has not elapsed.
  • The compaction_interval_sec interval has elapsed and there is at least one pending snapshot commit, even if the snapshot-count threshold has not been reached.
The snapshot-count threshold allows an early trigger; the interval limits how long the scheduler waits for that threshold when there are pending commits. It is not a minimum delay between compactions, and elapsed time alone does not trigger a new compaction without pending commits. Actual execution depends on scheduling and compactor availability. For example, with compaction_interval_sec = 600 and compaction.trigger_snapshot_count = 10, ten pending commits can trigger compaction before 600 seconds elapse. If only one commit is pending when the interval elapses, compaction is still eligible. Setting the threshold to 1 makes each new commit sufficient to trigger compaction without waiting for the interval.

Compaction types

RisingWave supports three compaction types for Iceberg tables. You can specify the type using the compaction.type parameter.
The small-files and files-with-delete compaction types are only supported in Merge-on-Read mode. Copy-on-Write mode only supports the full compaction type.

Parameters

Configure automatic maintenance by specifying the following parameters in the WITH clause of a CREATE SINK or CREATE TABLE ... ENGINE = iceberg statement.

General parameters

Manifest rewrite parameters

Compaction parameters

Deprecated storage config: The node-level storage config iceberg_compaction_write_parquet_max_row_group_rows is deprecated. Use the sink-level parameter compaction.write_parquet_max_row_group_rows instead.

Examples

Full compaction (default)

The following example enables automatic compaction with the default full compaction type:

Manifest rewrite

For Iceberg sinks with fragmented metadata, enable manifest rewrite to periodically merge data manifests:
Manifest rewrite is not supported for Iceberg format version 3 tables because the current rewrite path cannot preserve row lineage.

Small files compaction

For Merge-on-Read tables with many small files, use the small-files compaction type to only compact files smaller than a threshold:

Files with delete compaction

For Merge-on-Read tables with accumulated delete files, use the files-with-delete compaction type to only compact data files that have associated delete files:

Manual maintenance

In addition to automatic background maintenance, you can trigger maintenance manually using the VACUUM command. VACUUM rewrites eligible data manifests, then expires eligible snapshots. Manifest rewrite runs only when enable_manifest_rewrite = true and the configured rewrite thresholds select an eligible batch. Snapshot expiration runs only when enable_snapshot_expiration = true. VACUUM FULL first compacts data files, then performs the same manifest rewrite and snapshot expiration steps, subject to their enable flags and thresholds. The manifest rewrite and snapshot expiration steps do not override their corresponding maintenance settings.
VACUUM FULL waits for the manual compaction task to finish before returning. If the target already has an automatic or manual compaction task queued or running, the new VACUUM FULL request returns an error instead of scheduling another task.