Stackable Operator for Apache HBase
Apache HBase® on the
Stackable Data Platform
Run HBase natively on Kubernetes
Real-time read/write access to massive datasets
When you have billions of rows and require a specific one back in milliseconds, you need HBase – a distributed, scalable big-data store for real-time read/write access to huge datasets. The Stackable Operator deploys and operates HBase the Kubernetes-native way, automating cluster setup, scaling and monitoring, and integrating with the rest of your platform for ingestion, querying and visualization. Run it on-prem, air-gapped or in a sovereign cloud – no vendor lock-in.
Why Apache HBase in the Stackable data platform (SDP)?
Apache HBase is a distributed, column-oriented NoSQL database with Bigtable-like capabilities, built for real-time read/write access to huge, sparse datasets. It runs on top of HDFS – which on Kubernetes is backed by persistent volumes – and integrates with the wider Stackable Data Platform.
Linear, modular scalability
Scales horizontally, so you add capacity as data or traffic grows.
Strict consistency
Strictly consistent reads and writes help avoid stale or inconsistent data.
Automatic, configurable sharding (Regions)
Tables split and distribute across RegionServers automatically to balance load.
Automatic failover
If a RegionServer fails, the system fails over automatically to preserve availability.
Real-time query performance
Block cache, Bloom filters, and server-side filters optimize lookup latency and predicate push-downs; Prometheus metrics keep cluster health observable.
Multiple APIs & access methods
Java API plus REST and Thrift, with multiple data encoding options.
What does the Stackable Operator add for Apache HBase?
Scale masters and RegionServers independently
Change the replica count in the HBaseCluster CRD and the operator adds or removes RegionServer pods and rebalances regions automatically, absorbing higher load or shrinking footprint without manual intervention.
SQL and platform access to a NoSQL store
Query through the HBase API, or run SQL on HBase via Phoenix and a Trino catalog. Ingest through NiFi, Kafka or Spark, making a fast NoSQL store reachable for analysts and AI workloads.
Bulk-load big tables during migration
Load large datasets by bypassing the normal write path, migrating big tables onto the platform without overwhelming RegionServers.
What do all Stackable operators have in common?
One operator framework (Rust)
Every operator follows the same CRD pattern with Roles, RoleGroups and ConfigOverrides. Learn one operator, and you find your way around the next one instantly.
Infrastructure-as-Code
Every data app is a YAML CRD that lives in Git – reviewable, lintable and CI/CD-ready. The same definition works on Dev, Test and Prod.
Lifecycle management (Day-2 operations)
Deployment, restarts, certificate rotation and rolling upgrades are handled automatically – including pod restarts when configs or secrets change. And because you need to see what’s running, monitoring and logging come built in: Prometheus, Vector and OpenTelemetry feed ready-made Grafana dashboards – one observability setup for the whole platform.
Security by default
TLS, Kerberos for HDFS, OIDC login and OPA-based fine-grained authorization, all configurable per CRD. Daily vulnerability scans, SBOMs and cosign-signed images keep the whole platform patched – so you don’t track a dozen upstream projects yourself.
Kubernetes-native and modular
Runs on-prem, in any cloud or on a laptop, with no vendor lock-in – and explicit air-gapped support. Install only the operators you need, and add more later without touching the rest.
OpenShift-certified
All operators are Red Hat–certified and install directly from the Red Hat Certified Operators Catalog (OperatorHub) – compatible with Security Context Constraints and RBAC. Certified operation via a Stackable subscription.
Use Cases for Apache HBase
Real-time data capture
Store and query time-series and event data at scale, for workloads that need both high throughput and millisecond latency.
Consistent storage for fast-changing data
Manage dynamic datasets such as user profiles or session states while keeping consistency and availability.
Behavioral & clickstream data
E-commerce and digital platforms store billions of user interactions in HBase for personalization, recommendations and analytics.
FAQ - Frequently Asked Questions about HBase and Stackable
Stackable follows the stable upstream HBase releases and tests them with the operator. The versions currently supported in production (including any pinned patch levels) are listed on the ‘Supported product versions’ page of the Stackable documentation. The version is selected via the image in the HBaseCluster resource, or a custom image can be used if needed.
Click here for the current list.
Simply change the desired replica count in your HBaseCluster CRD and apply it. The operator adds or removes RegionServer pods and rebalances regions automatically, so the cluster can absorb higher load or reduce footprint without manual intervention.
Perform rolling upgrades by bumping the version in the CRD. The operator restarts components in a safe order (RegionServers first, then Masters), keeping the cluster available while moving to the new version – subject to any upstream upgrade notes.
Yes. Authentication is handled via Kerberos; with Stackable you store keytabs and principals as Kubernetes Secrets (managed by the Secret Operator), and the operator wires the JAAS configuration into the HBase components. Combine this with TLS to encrypt traffic between clients, RegionServers and Masters.
Yes. Superset supports fine-grained role-based acceHow do I secure HBase?ss control (RBAC). Permissions can be assigned to roles, which you then map to users or groups. With Stackable’s operator, these policies can be managed declaratively and version-controlled, which makes it easier to enforce security boundaries across multiple teams.
Yes. The operator supports HA Masters, automatic Region re-assignment on failures, and Kubernetes best practices like PodDisruptionBudgets, anti-affinity and persistent volumes. This gives you failover and high availability with predictable day-2 operations (upgrades, scaling, monitoring).
HBase is designed to run on HDFS. In Kubernetes you typically run HDFS (backed by persistent volumes) as part of the platform. Azure Data Lake Storage (ADLS) is also supported as a full replacement for HDFS. S3 (via Hadoop S3A) can be used for snapshot export, backups, and bulk data movement.
Use snapshots and the built-in ExportSnapshot utilities for fast, consistent backups of tables or namespaces. You can ship snapshots to external storage (e.g. HDFS or S3-compatible object storage) and complement this with replication for DR scenarios.
Yes. Metrics are exposed for Prometheus, and ready-made Grafana dashboards show cluster health, RPC latency, block cache effectiveness, compactions, region states and more. This gives operators real-time visibility and alerting out of the box.
Use NiFi or Kafka for ingestion and streaming pipelines, then query HBase data via Trino (through the Phoenix connector) or Apache Phoenix directly for SQL access. For dashboards, connect Superset to Trino/Phoenix to visualize results while HBase serves low-latency reads and writes underneath.
Yes. Spark can read from and write to HBase using the HBase client/connector (or via Phoenix for SQL semantics). This is ideal for batch processing, feature generation for ML and complex transformations that complement HBase’s real-time access patterns.
Use the HBase shell or client APIs to create or alter tables and column families. Many changes can be applied online; some properties require a brief disable/enable cycle. In GitOps setups, keep your table definitions and policies version-controlled so schema changes are reviewed and repeatable.
Yes, HBase stores byte arrays and handles sparse, wide tables well. For very large blobs, a common pattern is to store the object in an object store and keep a pointer or metadata in HBase; this keeps cell sizes efficient and read paths fast.
HBase is optimized for low-latency random reads/writes and operational/time-series workloads (OLTP-style). For interactive SQL analytics, pair HBase with Trino (via Phoenix) or Spark. This gives you the best of both worlds: fast operational access in HBase and powerful analytical queries via the SQL/compute layer.
The operator has been tested on major managed and self-hosted Kubernetes platforms: EKS, AKS, GKE, OpenShift, IONOS and K3s. This flexibility lets you deploy HBase reliably across cloud providers or on-premise infrastructure, while still benefiting from declarative management and operator automation.
Resources - Learn how to use the Stackable Operator for HBase
HBase Operator Docs
Dive into the official Stackable HBase Operator documentation to learn how to deploy and manage Apache HBase clusters on Kubernetes. The guide covers everything from installation and configuration to Kerberos & TLS security, scaling RegionServers, monitoring with Prometheus and Grafana, and performing rolling upgrades – all with GitOps-friendly, declarative cluster management.
Benchmarking HBase on Kubernetes
Does moving from bare-metal to Kubernetes cost you performance? A customer with a latency-critical use case wanted to know for sure – so we benchmarked the HBase and HDFS stack on Kubernetes against bare-metal. The result: no significant penalty compared to bare-metal. Details in the blog post.
GitHub Repository
Explore the official Stackable GitHub organization to access all open-source components of the Stackable Data Platform. Here you’ll find the source code for operators (including HBase), example configurations, Helm charts and issue trackers – the central hub for code, community contributions and release notes.
Demo
The hbase-hdfs-load-cycling-data demo copies a public cycling dataset from S3 into HDFS, turns it into HFiles and bulk-loads them into an HBase table you can then query from the hbase shell – deployable with a single stackablectl command.
Subscribe to our Newsletter
With the Stackable newsletter, you’ll always stay up to date on the latest from Stackable!
Newsletter
Subscribe to the newsletter
With the Stackable newsletter you’ll always be up to date when it comes to updates around Stackable!