Skip to content

Latest commit

 

History

History

README.md

DataHub

Artifact Hub

A Helm chart for DataHub

Install DataHub

Run the following command to install datahub with default configuration.

helm repo add datahub https://helm.datahubproject.io
helm install datahub datahub/datahub

If the default configuration is not applicable, you can update the values listed below in a values.yaml file and run

helm install datahub datahub/datahub --values <<path-to-values-file>>

Chart Values

Key Type Default Description
datahub-frontend.enabled bool true Enable Datahub Front-end
datahub-frontend.hpa.enabled bool false Set to true to enable the Horizontal Pod Autoscaler (HPA) for the deployment. Default is false. Note: Enabling HPA may introduce instability during scaling events, such as delays in waiting for nodes to become ready to allocate new pods.
datahub-frontend.hpa.minReplicas int 2 The minimum number of replicas that the HPA can scale down to. Default is 2.
datahub-frontend.hpa.maxReplicas int 3 The maximum number of replicas that the HPA can scale up to. Default is 3.
datahub-frontend.hpa.behavior object {} Configuration for scaling behavior of the HPA. Leave empty for default behavior. If not provided, the default behavior is as follows:
Default Scale Down Behavior:
- Stabilization Window: 300 seconds (5 minutes) - The amount of time the HPA will wait before allowing further scaling down.
- Scale Down Policy: 1 pod every 180 seconds (3 minutes). This ensures that the HPA does not scale down too quickly.
Default Scale Up Behavior:
- Stabilization Window: 0 seconds - No waiting period before scaling up. The HPA can immediately increase replicas if needed.
- Scale Up Policy: 1 pod every 60 seconds. This defines how frequently the HPA can scale up.
datahub-frontend.hpa.targetCPUUtilizationPercentage int 70 The target average CPU utilization (as a percentage) for scaling. Default is 70.
datahub-frontend.hpa.targetMemoryUtilizationPercentage int 70 The target average memory utilization (as a percentage) for scaling. Default is 70.
datahub-frontend.image.repository string "acryldata/datahub-frontend-react" Image repository for datahub-frontend
datahub-frontend.image.tag string "v0.11.0" Image tag for datahub-frontend
datahub-frontend.image.pullPolicy string "IfNotPresent" Image pull policy for datahub-frontend
datahub-frontend.image.args list [] Override the image's args. Used to configure custom startup or shutdown behavior
datahub-frontend.image.command list [] Override the image's command. Used to configure custom startup or shutdown behavior
datahub-gms.enabled bool true Enable GMS
datahub-gms.image.repository string "acryldata/datahub-gms" Image repository for datahub-gms
datahub-gms.image.tag string "v0.11.0" Image tag for datahub-gms
datahub-gms.image.pullPolicy string "IfNotPresent" Image pull policy for datahub-gms
datahub-gms.image.args list [] Override the image's args. Used to configure custom startup or shutdown behavior
datahub-gms.image.command list [] Override the image's command. Used to configure custom startup or shutdown behavior
datahub-mae-consumer.image.repository string "acryldata/datahub-mae-consumer" Image repository for datahub-mae-consumer
datahub-mae-consumer.image.tag string "v0.11.0" Image tag for datahub-mae-consumer
datahub-mae-consumer.image.pullPolicy string "IfNotPresent" Image pull policy for datahub-mae-consumer
datahub-mae-consumer.image.args list [] Override the image's args. Used to configure custom startup or shutdown behavior
datahub-mae-consumer.image.command list [] Override the image's command. Used to configure custom startup or shutdown behavior
datahub-mce-consumer.image.repository string "acryldata/datahub-mce-consumer" Image repository for datahub-mce-consumer
datahub-mce-consumer.image.tag string "v0.11.0" Image tag for datahub-mce-consumer
datahub-mce-consumer.image.pullPolicy string "IfNotPresent" Image pull policy for datahub-mce-consumer
datahub-mce-consumer.image.args list [] Override the image's args. Used to configure custom startup or shutdown behavior
datahub-mce-consumer.image.command list [] Override the image's command. Used to configure custom startup or shutdown behavior
datahub-ingestion-cron.enabled bool false Enable cronjob for periodic ingestion
datahubUpgrade.podSecurityContext object {} Pod security context for datahubUpgrade jobs
datahubUpgrade.securityContext object {} Container security context for datahubUpgrade jobs
datahubUpgrade.podAnnotations object {} Pod annotations for datahubUpgrade jobs
datahubUpgrade.restoreIndices.resources object '{}' Kube Resource definitions for the datahub upgrade job 'restore indices'
datahubUpgrade.restoreIndices.extraSidecars list [] Add additional sidecar containers to the job pod
datahubUpgrade.restoreIndices.concurrencyPolicy string Allow, Forbid, Replace Add concurrencyPolicy for the restoreIndicies cron job
datahubUpgrade.restoreIndices.image.args list [] Override the image's args. Used to configure custom startup or shutdown behavior
datahubUpgrade.restoreIndices.image.command list [] Override the image's command. Used to configure custom startup or shutdown behavior
elasticsearchSetupJob.enabled bool true Enable setup job for elasicsearch
elasticsearchSetupJob.image.repository string "acryldata/datahub-elasticsearch-setup" Image repository for elasticsearchSetupJob
elasticsearchSetupJob.image.tag string "v0.11.0" Image repository for elasticsearchSetupJob
elasticsearchSetupJob.image.pullPolicy string "IfNotPresent" Image pull policy for elasticsearchSetupJob
elasticsearchSetupJob.image.args list [] Override the image's args. Used to configure custom startup or shutdown behavior
elasticsearchSetupJob.image.command list [] Override the image's command. Used to configure custom startup or shutdown behavior
elasticsearchSetupJob.resources object '{}' Kube Resource definitions for elasticsearchSetupJob
elasticsearchSetupJob.podSecurityContext object {"fsGroup": 1000} Pod security context for elasticsearchSetupJob
elasticsearchSetupJob.securityContext object {"runAsUser": 1000} Container security context for elasticsearchSetupJob
elasticsearchSetupJob.podAnnotations object {} Pod annotations for elasticsearchSetupJob
elasticsearchSetupJob.extraSidecars list [] Add additional sidecar containers to the job pod
kafkaSetupJob.enabled bool true Enable setup job for kafka
kafkaSetupJob.image.repository string "acryldata/datahub-kafka-setup" Image repository for kafkaSetupJob
kafkaSetupJob.image.tag string "v0.11.0" Image repository for kafkaSetupJob
kafkaSetupJob.image.pullPolicy string "IfNotPresent" Image pull policy for kafkaSetupJob
kafkaSetupJob.image.args list [] Override the image's args. Used to configure custom startup or shutdown behavior
kafkaSetupJob.image.command list [] Override the image's command. Used to configure custom startup or shutdown behavior
kafkaSetupJob.resources object '{}' Kube Resource definitions for kafkaSetupJob
kafkaSetupJob.podSecurityContext object {"fsGroup": 1000} Pod security context for kafkaSetupJob
kafkaSetupJob.securityContext object {"runAsUser": 1000} Container security context for kafkaSetupJob
kafkaSetupJob.podAnnotations object {} Pod annotations for kafkaSetupJob
kafkaSetupJob.extraSidecars list [] Add additional sidecar containers to the job pod
mysqlSetupJob.enabled bool false Enable setup job for mysql
mysqlSetupJob.image.repository string "acryldata/datahub-mysql-setup" Image repository for mysqlSetupJob
mysqlSetupJob.image.tag string "v0.11.0" Image repository for mysqlSetupJob
mysqlSetupJob.image.pullPolicy string "IfNotPresent" Image pull policy for mysqlSetupJob
mysqlSetupJob.image.args list [] Override the image's args. Used to configure custom startup or shutdown behavior
mysqlSetupJob.image.command list [] Override the image's command. Used to configure custom startup or shutdown behavior
mysqlSetupJob.resources object '{}' Kube Resource definitions for mysqlSetupJob
mysqlSetupJob.podSecurityContext object {"fsGroup": 1000} Pod security context for mysqlSetupJob
mysqlSetupJob.securityContext object {"runAsUser": 1000} Container security context for mysqlSetupJob
mysqlSetupJob.podAnnotations object {} Pod annotations for mysqlSetupJob
mysqlSetupJob.extraSidecars list [] Add additional sidecar containers to the job pod
postgresqlSetupJob.enabled bool false Enable setup job for postgresql
postgresqlSetupJob.image.repository string "acryldata/datahub-postgres-setup" Image repository for postgresqlSetupJob
postgresqlSetupJob.image.tag string "v0.11.0" Image repository for postgresqlSetupJob
postgresqlSetupJob.image.pullPolicy string "IfNotPresent" Image pull policy for postgresqlSetupJob
postgresqlSetupJob.image.args list [] Override the image's args. Used to configure custom startup or shutdown behavior
postgresqlSetupJob.image.command list [] Override the image's command. Used to configure custom startup or shutdown behavior
postgresqlSetupJob.resources object '{}' Kube Resource definitions for postgresqlSetupJob
postgresqlSetupJob.podSecurityContext object {"fsGroup": 1000} Pod security context for mysqlSetupJob
postgresqlSetupJob.securityContext object {"runAsUser": 1000} Container security context for mysqlSetupJob
postgresqlSetupJob.podAnnotations object {} Pod annotations for mysqlSetupJob
postgresqlSetupJob.extraSidecars list [] Add additional sidecar containers to the job pod
datahubSystemUpdate.extraSidecars list [] Add additional sidecar containers to the job pod
global.podLabels object {} Custom labels to add to all pods created by the chart
global.strict_mode boolean true Enables validations in helm charts to ensure features work as expected. Recommended NOT TO CHANGE.
global.datahub_standalone_consumers_enabled boolean false Enable standalone consumers for kafka
global.datahub_analytics_enabled boolean true Enable datahub usage analytics
global.datahub.appVersion string "1.0" App version for annotation
global.datahub.gms.protocol string "http" Protocol of GMS service
global.datahub.gms.host string `"datahub-datahub-gms" Host of GMS service
global.datahub.gms.port string "8080" Port of GMS service
global.datahub.monitoring.portName string jmx Name of Kube port for monitoring
global.elasticsearch.host string "elasticsearch-master" Elasticsearch host name (endpoint)
global.elasticsearch.port string "9200" Elasticsearch port
global.kafka.bootstrap.server string "prerequisites-broker:9092" Kafka bootstrap servers (with port)
global.kafka.zookeeper.server string "prerequisites-zookeeper:2181" Kafka zookeeper servers (with port)
global.kafka.consumer.stopContainerOnDeserializationError boolean true Determines whether or not to halt progress when encountering a deserialization error, halting prevents data loss but prevents progress until fixed
global.kafka.consumer_groups.datahub_upgrade_history_kafka_consumer_group_id.gms string "<<release-name>>-duhe-consumer-job-client-gms" Consumer group id for consuming DataHub Upgrade history events by the GMS
global.kafka.consumer_groups.datahub_upgrade_history_kafka_consumer_group_id.mae-consumer string "<<release-name>>-duhe-consumer-job-client-mcl" Consumer group id for consuming DataHub Upgrade history events by the MAE consumer
global.kafka.consumer_groups.datahub_upgrade_history_kafka_consumer_group_id.mce-consumer string "<<release-name>>-duhe-consumer-job-client-mcp" Consumer group id for consuming DataHub Upgrade history events by the MCE consumer
global.kafka.consumer_groups.datahub_actions_doc_propagation_consumer_group_id string "datahub_doc_propagation_action" Consumer group id for consuming events by DataHub Actions doc propagation action pipeline
global.kafka.consumer_groups.datahub_actions_ingestion_executor_consumer_group_id string "ingestion_executor" Consumer group id for consuming events by DataHub Actions ingestion executor pipeline
global.kafka.consumer_groups.datahub_actions_slack_consumer_group_id string "datahub_slack_action" Consumer group id for consuming events by DataHub Actions slack action pipeline
global.kafka.consumer_groups.datahub_actions_teams_consumer_group_id string "datahub_teams_action" Consumer group id for consuming events by DataHub Actions teams action pipeline
global.kafka.consumer_groups.datahub_usage_event_kafka_consumer_group_id string "datahub-usage-event-consumer-job-client" Consumer group id for consuming DataHub Usage events by the GMS or MAE consumer
global.kafka.consumer_groups.metadata_change_log_kafka_consumer_group_id string "generic-mae-consumer-job-client" Consumer group id for consuming Metadata Change Log events by the GMS or MAE consumer
global.kafka.consumer_groups.platform_event_kafka_consumer_group_id string "generic-platform-event-job-client" Consumer group id for consuming Platform events by the GMS or MAE consumer
global.kafka.consumer_groups.metadata_change_event_kafka_consumer_group_id string "mce-consumer-job-client" Consumer group id for consuming Metadata Change events by the GMS or MCE consumer
global.kafka.consumer_groups.metadata_change_proposal_kafka_consumer_group_id string "generic-mce-consumer-job-client" Consumer group id for consuming Metadata Change Proposal events by the GMS or MCE consumer
global.kafka.topics.metadata_change_event_name string "MetadataChangeEvent_v4" Kafka topic name for Metadata Change Events (deprecated)
global.kafka.topics.failed_metadata_change_event_name string "FailedMetadataChangeEvent_v4" Kafka topic name for Failed Metadata Change events (deprecated)
global.kafka.topics.metadata_audit_event_name string "MetadataAuditEvent_v4" Kafka topic name for Metadata Audit events (deprecated)
global.kafka.topics.datahub_usage_event_name string "DataHubUsageEvent_v1" Kafka topic name for DataHub Usage events
global.kafka.topics.metadata_change_proposal_topic_name string "MetadataChangeProposal_v1" Kafka topic name for Metadata Change Proposal events
global.kafka.topics.failed_metadata_change_proposal_topic_name string "FailedMetadataChangeProposal_v1" Kafka topic name for Failed Metadata Change Proposal events
global.kafka.topics.metadata_change_log_versioned_topic_name string "MetadataChangeLog_Versioned_v1" Kafka topic name for Versioned Metadata Change Log events
global.kafka.topics.metadata_change_log_timeseries_topic_name string "MetadataChangeLog_Timeseries_v1" Kafka topic name for Timeseries Metadata Change Log events
global.kafka.topics.platform_event_topic_name string "PlatformEvent_v1" Kafka topic name for Platform events
global.kafka.topics.datahub_upgrade_history_topic_name string "DataHubUpgradeHistory_v1" Kafka topic name for DataHub Upgrade History events
global.kafka.schemaregistry.url string `` URL to kafka schema registry if using KAFKA type
global.neo4j.host string "prerequisites-neo4j:7474" Neo4j host address (with port)
global.neo4j.uri string "bolt://prerequisites-neo4j" Neo4j URI
global.neo4j.database string "graph.db" Neo4J database
global.neo4j.username string "neo4j" Neo4j user name
global.neo4j.password.secretRef string "neo4j-secrets" Secret that contains the Neo4j password
global.neo4j.password.secretKey string "neo4j-password" Secret key that contains the Neo4j password
global.sql.datasource.driver string "com.mysql.cj.jdbc.Driver" Driver for the SQL database
global.sql.datasource.host string "prerequisites-mysql:3306" SQL database host (with port)
global.sql.datasource.hostForMysqlClient string "prerequisites-mysql" SQL database host (without port)
global.sql.datasource.port string "3306" SQL database port
global.sql.datasource.database string "datahub" SQL database name
global.sql.datasource.url string "jdbc:mysql://prerequisites-mysql:3306/datahub?verifyServerCertificate=false\u0026useSSL=true" URL to access SQL database
global.sql.datasource.username string "root" SQL user name
global.sql.datasource.username.secretRef string "mysql-secrets" Secret that contains the MySQL username
global.sql.datasource.username.secretKey string "mysql-username" Secret key that contains the MySQL username
global.sql.datasource.password.secretRef string "mysql-secrets" Secret that contains the MySQL password
global.sql.datasource.password.secretKey string "mysql-password" Secret key that contains the MySQL password
global.sql.datasource.password.value string "mysql-password" Alternative to using the secret above, uses raw string value instead
global.graph_service_impl string elasticsearch One of elasticsearch or neo4j. Determines which backend to use for the GMS graph service. Elasticsearch is recommended for a simplified deployment.

Optional Chart Values

Key Type Default Description
datahub-gms.sql.datasource.username string root SQL username for GMS (overrides global value)
datahub-gms.sql.datasource.username.secretRef string "mysql-secrets" Secret that contains the GMS SQL username (overrides global value)
datahub-gms.sql.datasource.username.secretKey string "mysql-username" Secret key that contains the GMS SQL username (overrides global value)
datahub-gms.sql.datasource.password.secretRef string "mysql-secrets" Secret that contains the GMS SQL password (overrides global value)
datahub-gms.sql.datasource.password.secretKey string "mysql-password" Secret key that contains the GMS SQL password (overrides global value)
datahub-gms.sql.datasource.password.value string "mysql-password" Alternative to using the secret above, uses raw string value for GMS SQL login (overrides global value)
mysqlSetupJob.username string root SQL username for mysqlSetupJob (overrides global value)
mysqlSetupJob.password.secretRef string "mysql-secrets" Secret that contains the mysqlSetupJob SQL password (overrides global value)
mysqlSetupJob.password.secretKey string "mysql-password" Secret key that contains the mysqlSetupJob SQL password (overrides global value)
mysqlSetupJob.password.value string "mysql-password" Alternative to using the secret above, uses raw string value for mysqlSetupJob SQL login (overrides global value)
mysqlSetupJob.image.registry string `` Image registry override to be used by the job.
postgresqlSetupJob.username string root SQL username for postgresqlSetupJob (overrides global value)
postgresqlSetupJob.password.secretRef string "mysql-secrets" Secret that contains the postgresqlSetupJob SQL password (overrides global value)
postgresqlSetupJob.password.secretKey string "mysql-password" Secret key that contains the postgresqlSetupJob SQL password (overrides global value)
postgresqlSetupJob.password.value string "mysql-password" Alternative to using the secret above, uses raw string value for postgresqlSetupJob SQL login (overrides global value)
postgresqlSetupJob.image.registry string `` Image registry override to be used by the job.
acryl-datahub-actions.ingestionSecretFiles.name string "" Name of the k8s secret that holds any secret files (e.g., SSL certificates and private keys) that are used in your ingestion recipes. The keys in the secret will be mounted as individual files under /etc/datahub/ingestion-secret-files
acryl-datahub-actions.ingestionSecretFiles.defaultMode string "" The permission mode for the volume that mounts k8s secret under /etc/datahub/ingestion-secret-files, default value is 0444 which allows read access by owner, group, and other users
global.credentialsAndCertsSecrets.name string "" Name of the secret that holds SSL certificates (keystores, truststores)
global.credentialsAndCertsSecrets.path string "/mnt/certs" Path to mount the SSL certificates
global.credentialsAndCertsSecrets.secureEnv map {} Map of SSL config name and the corresponding value in the secret
global.springKafkaConfigurationOverrides map {} Map of configuration overrides for accessing kafka
global.elasticsearch.useSSL bool false Whether to enable SSL for accessing elasticsearch
global.elasticsearch.auth.username string "" Elasticsearch username
global.elasticsearch.auth.password.secretRef string "" Secret that contains the elasticsearch password
global.elasticsearch.auth.password.secretKey string "" Secret key that contains the elasticsearch password
global.elasticsearch.auth.password.value string "" Alternative to using the secret above, uses raw string value instead
global.kafka.precreateTopics boolean true URL to kafka schema registry if using KAFKA type
global.kafka.schemaregistry.type string "INTERNAL" Type of schema registry (INTERNAL, KAFKA, or AWS_GLUE)
global.kafka.schemaregistry.glue.region string "" Region of the AWS Glue schema registry
global.kafka.schemaregistry.glue.registry string "" Name of the AWS Glue schema registry
global.kafka.schemaregistry.configureCleanupPolicy boolean null Whether to configure clean up policy on schema registry. By default, a suitable default is chosen based on global.kafka.schemaregistry.type (true only of type is KAFKA). Set this to have explicit control.
datahub.metadata_service_authentication.enabled bool true Whether Metadata Service Authentication is enabled.
elasticsearchSetupJob.image.registry string `` Image registry override to be used by the job.
kafkaSetupJob.image.registry string `` Image registry override to be used by the job.
datahubUpgrade.image.registry string `` Image registry override to be used by the job.
datahubSystemUpdate.image.registry string `` Image registry override to be used by the job.
global.datahub.metadata_service_authentication.systemClientId string "__datahub_system" The internal system id that is used to communicate with DataHub GMS. Required if metadata_service_authentication is 'true'.
global.datahub.metadata_service_authentication.systemClientSecret.secretRef string datahub-auth-secrets The reference to a secret containing the internal system secret that is used to communicate with DataHub GMS. If a secret reference is not provided, a random one will be generated for you in a Kubernetes secret called datahub-auth-secrets.
global.datahub.metadata_service_authentication.systemClientSecret.secretKey string system_client_secret The key of a secret containing the internal system secret that is used to communicate with DataHub GMS. If a secret reference is not provided, a random one will be generated for you in a Kubernetes secret value named system_client_secret within a secret named datahub-auth-secrets.
global.datahub.metadata_service_authentication.tokenService.signingKey.secretRef string datahub-auth-secrets The reference to a secret containing the internal system secret that is used to sign JWT auth tokens issued by DataHub GMS. If a secret reference is not provided, a random one will be generated for you in a Kubernetes secret called datahub-auth-secrets.
global.datahub.metadata_service_authentication.tokenService.signingKey.secretKey string token_service_signing_key The key of a secret containing the internal system secret that is used to sign JWT auth tokens issued by DataHub GMS. If a secret reference is not provided, a random one will be generated for you in a Kubernetes secret value named token_service_signing_key within a secret named datahub-auth-secrets.
global.datahub.metadata_service_authentication.tokenService.salt.secretRef string datahub-auth-secrets The reference to a secret containing the internal system secret that is used to salt JWT auth tokens signatures issued by DataHub GMS that is part of the metadata graph. If a secret reference is not provided, a random one will be generated for you in a Kubernetes secret called datahub-auth-secrets.
global.datahub.metadata_service_authentication.tokenService.salt.secretKey string token_service_salt The key of a secret containing the internal system secret that is used to salt JWT auth tokens signatures issued by DataHub GMS that is part of the metadata graph. If a secret reference is not provided, a random one will be generated for you in a Kubernetes secret value named token_service_salt within a secret named datahub-auth-secrets.
global.datahub.metadata_service_authentication.provisionSecrets.enabled bool true Whether auth secrets (system client secret, token signing key & token service salt) should be provisioned on the first deployment for you. Set this to false if you are overriding global.datahub.metadata_service_authentication.tokenService.signingKey.secretRef or global.datahub.metadata_service_authentication systemClientSecret.secretRef.
global.datahub.metadata_service_authentication.provisionSecrets.autoGenerate bool true Whether auth secrets (token signing key, system client secret & token service salt) should be provisioned on the first deployment for you with a random seed on the first deployment for you. Set this to false and use global.datahub.metadata_service_authentication.provisionSecrets.secretValues.* if you would like to specify the secret values directly.
global.datahub.encryptionKey.provisionSecrets.secretValues.secret string `` The system client secret key value to be used if specified directly.
global.datahub.encryptionKey.provisionSecrets.secretValues.signingkey string `` The system signing key value to be used if specified directly.
global.datahub.encryptionKey.provisionSecrets.secretValues.salt string `` The token service salt value to be used if specified directly.
global.datahub.managed_ingestion.enabled bool true Whether or not UI-based ingestion experience is enabled.
global.datahub.encryptionKey.secretRef string datahub-encryption-secrets The reference to a secret containing an alpha-numeric encryption key, which is used to encrypt Secrets on DataHub. If a secret reference is not provided, a random one will be generated for you in a Kubernetes secret named datahub-encryption-secrets.
global.datahub.encryptionKey.secretKey string encryption_key_secret The key of a secret containing an alpha-numeric encryption key, which is used to encrypt Secrets on DataHub. If a secret reference is not provided, a random one will be generated for you in a Kubernetes secret value named encryption_key_secret within a secret named datahub-encryption-secrets.
global.datahub.encryptionKey.callerGuardMode string ENFORCE SECRET_SERVICE_CALLER_GUARD_MODE: ENFORCE (default), AUDIT, or DISABLED. v1.7+ ENFORCE blocks browser/user-PAT secret decrypt; use datahub-actions with system client credentials.
global.datahub.i18n.enabled bool true I18N_ENABLED on GMS. v1.7+ defaults on (multi-language UI). Set false for English-only.
global.datahub.objectStorage.uri string `` Preferred object storage URI (s3://, gs://, file:///). Sets DATAHUB_OBJECT_STORAGE_URI. Prefer over legacy global.datahub.gms.s3.
global.datahub.managed_ingestion.defaultCliVersion string `` 0.11.0 This is the version of the DataHub CLI to use for UI ingestion, by default.
global.datahub.encryptionKey.provisionSecret.enabled bool true Whether an encryption key secret should be provisioned on the first deployment for you. Set this to false if you are overriding global.datahub.encryptionKey.secretRef.
global.datahub.encryptionKey.provisionSecret.autoGenerate bool true Whether an encryption key secret should be provisioned for you with a random seed on the first deployment for you. Set this to false and use global.datahub.encryptionKey.provisionSecret.secretValues.encryptionKey if you would like to specify the secret values directly.
global.datahub.encryptionKey.provisionSecret.secretValues.encryptionKey string `` The encryption key value to be used if specified directly.
global.datahub.enable_retention bool false Whether or not to enable retention on local DB
global.sql.datasource.hostForpostgresqlClient string "" SQL database host (without port) when using postgresqlSetupJob
global.imageRegistry string "docker.io" Default docker image registry to be used by all services.

Enabling Semantic Search (Beta)

⚠️ Beta Feature: Semantic search is currently in beta. Only the document entity type is officially supported. Other entity types may work but your mileage may vary (YMMV).

Semantic search (vector similarity search) allows finding entities based on semantic meaning rather than just keyword matching.

Prerequisites

  • Elasticsearch/OpenSearch cluster with k-NN plugin support
  • Documents must have embeddings generated
  • Embedding provider credentials (OpenAI API key, AWS credentials, or Cohere API key)

Configuration by Provider

Choose one of the following configurations based on your embedding provider:

Option 1: OpenAI (Recommended for getting started)

1. Create a secret with your OpenAI API key:

kubectl create secret generic openai-secret --from-literal=api-key=sk-your-api-key-here

2. Configure in values.yaml:

global:
  datahub:
    semantic_search:
      enabled: true
      enabledEntities: "document"
      vectorDimension: 3072  # For text-embedding-3-large

      provider:
        type: "openai"
        openai:
          apiKey:
            secretRef: "openai-secret"
            secretKey: "api-key"
          model: "text-embedding-3-large"

Option 2: AWS Bedrock (Cohere)

1. Configure AWS credentials (via IAM role, service account, or secret)

2. Configure in values.yaml:

global:
  datahub:
    semantic_search:
      enabled: true
      enabledEntities: "document"
      vectorDimension: 1024  # For Cohere embed-english-v3

      provider:
        type: "aws-bedrock"
        bedrock:
          modelId: "cohere.embed-english-v3"
          awsRegion: "us-west-2"

Note: AWS Bedrock uses AWS SDK default credentials chain (IAM roles, environment variables, etc.)

Option 3: Cohere Direct

1. Create a secret with your Cohere API key:

kubectl create secret generic cohere-secret --from-literal=api-key=your-cohere-key

2. Configure in values.yaml:

global:
  datahub:
    semantic_search:
      enabled: true
      enabledEntities: "document"
      vectorDimension: 1024  # For embed-english-v3.0

      provider:
        type: "cohere"
        cohere:
          apiKey:
            secretRef: "cohere-secret"
            secretKey: "api-key"
          model: "embed-english-v3.0"

Vector Dimensions by Model

⚠️ Critical: The vectorDimension must match your embedding model output:

Provider Model Vector Dimension
OpenAI text-embedding-3-small 1536
OpenAI text-embedding-3-large 3072
AWS Bedrock cohere.embed-english-v3 1024
Cohere embed-english-v3.0 1024

Deploy

helm upgrade datahub datahub/datahub --values values.yaml

For more details, see the DataHub Semantic Search documentation.