Expand description
§FastPFor for Rust
Fast integer compression for Rust — both a pure-Rust implementation and a wrapper around the C++ FastPFor library. Supports 32-bit (and for some codecs 64-bit) integers. Based on the Decoding billions of integers per second through vectorization, 2012 paper.
The Rust decoder is about 29% faster than the C++ version. The Rust implementation is safe code: its only unsafe is the call into the AVX2 kernels of the opt-in Simd codecs, made after runtime CPU feature detection. The crate has #![deny(unsafe_code)]; the only other exemption is the generated FFI bridge of the optional cpp feature.
§Usage
§Rust Implementation (default)
The simplest way is FastPFor256 — a composite codec that handles any input
length by compressing aligned 256-element blocks with FastPForBlock256 and encoding any
leftover values with VariableByte.
use fastpfor::{AnyLenCodec, FastPFor256};
let mut codec = FastPFor256::default();
let input: Vec<u32> = (0..1000).collect();
let mut encoded = Vec::new();
codec.encode(&input, &mut encoded).unwrap();
let mut decoded = Vec::new();
codec.decode(&encoded, &mut decoded, None).unwrap();
assert_eq!(decoded, input);For block-aligned inputs you can use the lower-level BlockCodec API:
use fastpfor::{BlockCodec, FastPForBlock256, slice_to_blocks};
let mut codec = FastPForBlock256::default();
let input: Vec<u32> = (0..512).collect(); // exactly 2 blocks of 256
let (blocks, remainder) = slice_to_blocks::<FastPForBlock256>(&input);
assert_eq!(blocks.len(), 2);
assert!(remainder.is_empty());
let mut encoded = Vec::new();
codec.encode_blocks(blocks, &mut encoded).unwrap();
let mut decoded = Vec::new();
codec.decode_blocks(&encoded, Some(u32::try_from(blocks.len() * 256).expect("block count fits in u32")), &mut decoded).unwrap();
assert_eq!(decoded, input);§64-bit integers (u64)
The FastPForWide128 / FastPForWide256 codecs compress u64 values.
They implement AnyLenCodec (with Elem = u64) for native use, and BlockCodec64
(encode64 / decode64) for comparison against the C++ codecs.
The wire format is byte-compatible with the C++ CppFastPFor128 / CppFastPFor256 64-bit paths.
use fastpfor::{AnyLenCodec, FastPForWide256};
let mut codec = FastPForWide256::default();
let input: Vec<u64> = (0..600).map(|i| i * 1_000_000_000).collect();
let mut encoded = Vec::new();
codec.encode(&input, &mut encoded).unwrap();
let mut decoded = Vec::new();
codec.decode(&encoded, &mut decoded, None).unwrap();
assert_eq!(decoded, input);§SIMD kernels
The FastPForSimd* codecs (FastPForSimd128, FastPForSimd256, FastPForSimdWide128, FastPForSimdWide256,
and the matching FastPForSimdBlock* block codecs) are drop-in replacements for the codecs above.
They produce byte-identical output and decode each other’s streams, so encoders and decoders can be mixed freely.
x86_64: AVX2 kernels, selected at runtime; CPUs without AVX2 use the scalar kernels.aarch64: NEON kernels.u64values wider than 32 bits use the scalar kernels.- Other targets: the scalar kernels.
use fastpfor::{AnyLenCodec, FastPFor256, FastPForSimd256};
let input: Vec<u32> = (0..1000).collect();
let mut encoded = Vec::new();
FastPForSimd256::default().encode(&input, &mut encoded).unwrap();
let mut scalar_encoded = Vec::new();
FastPFor256::default().encode(&input, &mut scalar_encoded).unwrap();
assert_eq!(encoded, scalar_encoded);
let mut decoded = Vec::new();
FastPFor256::default().decode(&encoded, &mut decoded, None).unwrap();
assert_eq!(decoded, input);Note that the C++ CppSimdFastPFor* codecs use a different, interleaved bit layout and are not
compatible with either the Rust codecs or the C++ CppFastPFor* codecs.
§C++ Wrapper (cpp feature)
Enable the cpp feature in Cargo.toml:
fastpfor = { version = "0.9", features = ["cpp"] }All C++ codecs implement the same AnyLenCodec trait (encode / decode), so
the usage pattern is identical to the Rust examples above — just swap the codec type,
e.g. cpp::CppFastPFor128::new().
Thread safety: C++ codec instances have internal state and are not thread-safe. Create one instance per thread or synchronize access externally.
§Crate Features
| Feature | Default | Description |
|---|---|---|
rust | yes | Pure-Rust implementation — safe code, no build dependencies |
cpp | no | C++ wrapper via CXX — requires a C++14 compiler with SIMD support |
cpp_portable | no | Enables cpp, compiles C++ with SSE4.2 baseline (runs on any x86-64 from ~2008+) |
cpp_native | no | Enables cpp, compiles C++ with -march=native for maximum throughput on the build machine |
The FASTPFOR_SIMD_MODE environment variable (portable or native) can override the SIMD mode at build time.
Recommendation: Use cpp_portable (not cpp_native) for distributable binaries.
§Supported Algorithms
§Rust (rust feature)
Rust block codecs require block-aligned input. CompositeCodec chains a block codec with a tail codec (e.g. VariableByte) to handle arbitrary-length input. FastPFor256/FastPFor128 (for u32) and FastPForWide256/FastPForWide128 (for u64) are type aliases for such composites.
| Codec | Description |
|---|---|
FastPFor256 | CompositeCodec of FastPForBlock256 + VariableByte (u32) |
FastPFor128 | CompositeCodec of FastPForBlock128 + VariableByte (u32) |
FastPForWide256 | CompositeCodec of FastPForBlockWide256 + VariableByte (u64) |
FastPForWide128 | CompositeCodec of FastPForBlockWide128 + VariableByte (u64) |
VariableByte | Variable-byte encoding, MSB is opposite to protobuf’s varint |
JustCopy | No compression; useful as a baseline |
FastPForBlock256 | FastPFor with 256-element u32 blocks; block-aligned input only |
FastPForBlock128 | FastPFor with 128-element u32 blocks; block-aligned input only |
FastPForSimd* | Same as the codec without Simd, using SIMD kernels; byte-identical output |
§C++ (cpp feature)
All C++ codecs are composite (any-length) and implement AnyLenCodec only.
u64-capable codecs (CppFastPFor128, CppFastPFor256, CppVarInt) also implement BlockCodec64 with encode64 / decode64.
| Codec | Notes |
|---|---|
CppFastPFor128 | FastPFor + VByte composite, 128-element blocks. Also supports u64. |
CppFastPFor256 | FastPFor + VByte composite, 256-element blocks. Also supports u64. |
CppSimdFastPFor128 | SIMD-optimized 128-element variant |
CppSimdFastPFor256 | SIMD-optimized 256-element variant |
CppBP32 | Binary packing, 32-bit blocks |
CppFastBinaryPacking8 | Binary packing, 8-bit groups |
CppFastBinaryPacking16 | Binary packing, 16-bit groups |
CppFastBinaryPacking32 | Binary packing, 32-bit groups |
CppSimdBinaryPacking | SIMD-optimized binary packing |
CppPFor | Patched frame-of-reference |
CppSimplePFor | Simplified PFor variant |
CppNewPFor | PFor with improved exception handling |
CppOptPFor | Optimized PFor |
CppPFor2008 | Reference implementation from original paper |
CppSimdPFor | SIMD PFor |
CppSimdSimplePFor | SIMD SimplePFor |
CppSimdNewPFor | SIMD NewPFor |
CppSimdOptPFor | SIMD OptPFor |
CppSimple16 | 16 packing modes in 32-bit words |
CppSimple9 | 9 packing modes |
CppSimple9Rle | Simple9 with run-length encoding |
CppSimple8b | 8 packing modes in 64-bit words |
CppSimple8bRle | Simple8b with run-length encoding |
CppSimdGroupSimple | SIMD group-simple encoding |
CppSimdGroupSimpleRingBuf | SIMD group-simple with ring buffer |
CppVByte | Standard variable-byte encoding |
CppMaskedVByte | SIMD masked variable-byte |
CppStreamVByte | SIMD stream variable-byte |
CppVarInt | Standard varint. Also supports u64. |
CppVarIntGb | Group varint |
CppCopy | No compression (baseline) |
§Benchmarks
§Decoding
Using Linux x86-64 running just bench::cpp-vs-rust-decode native. The values below are time measurements; smaller values indicate faster decoding.
| name | cpp (ns) | rust (ns) | % faster |
|---|---|---|---|
clustered/1024 | 643.24 | 392.93 | 38.91% |
clustered/4096 | 1986 | 1414.8 | 28.76% |
sequential/1024 | 653.69 | 396.02 | 39.42% |
sequential/4096 | 2106 | 1476.2 | 29.91% |
sparse/1024 | 428.8 | 352.38 | 17.82% |
sparse/4096 | 1114 | 1179.5 | -5.88% |
uniform_large_value_distribution/1024 | 286.74 | 153.06 | 46.62% |
uniform_large_value_distribution/4096 | 748.19 | 558.05 | 25.41% |
uniform_small_value_distribution/1024 | 606.4 | 405.44 | 33.14% |
uniform_small_value_distribution/4096 | 2017.3 | 1403.7 | 30.42% |
Rust encoding has not yet been fully optimized or verified.
§Build Requirements
- Rust feature (
rust, the default): no additional dependencies. - C++ feature (
cpp): requires a C++14-capable compiler with SIMD intrinsics. See FastPFor C++ requirements.
§Linux
The default GitHub Actions runner has all needed dependencies.
For local development:
# This list may be incomplete
sudo apt-get install build-essentiallibsimde-dev is optional. On ARM/aarch64, the C++ build fetches SIMDe via CMake
and the CXX bridge reuses that include path automatically.
§macOS
On Apple Silicon, SIMDe installation is usually not required — the C++ build fetches it via CMake.
If you prefer a Homebrew fallback:
brew install simde
export CXXFLAGS="-I/opt/homebrew/include"
export CFLAGS="-I/opt/homebrew/include"§Development
This project uses just as a task runner:
cargo install just # install once
just # list available commands
just test # run all tests§License
Licensed under either of
- Apache License, Version 2.0 (LICENSE-APACHE or https://www.apache.org/licenses/LICENSE-2.0)
- MIT license (LICENSE-MIT or https://opensource.org/licenses/MIT) at your option.
§Contribution
Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in the work by you, as defined in the Apache-2.0 license, shall be dual-licensed as above, without any additional terms or conditions.
Modules§
- cpp
cpp - Rust wrapper for the
FastPFORC++ library C++ codec wrappers — see the crate-level documentation for usage and codec selection.
Structs§
- Composite
Codec rust - Combines a block-oriented codec with an arbitrary-length tail codec.
- FastP
For rust - Type-safe block codec with block size encoded in the type. Type-safe block codec with block size encoded in the type. Type-safe block codec with block size encoded in the type. Fast Patched Frame-of-Reference (FastPFOR) codec.
- Just
Copy rust - Pass-through codec — implements
AnyLenCodec. Pass-through codec — implementsAnyLenCodec. Pass-through codec — implementsAnyLenCodec. A no-op codec that copies data without compression. - Scalar
rust - Portable scalar kernels.
- Simd
rust - SIMD kernels producing byte-identical output to
Scalar: AVX2 onx86_64when detected at runtime, NEON onaarch64, andScalarotherwise. - Variable
Byte rust - Variable-byte codec — implements
AnyLenCodec. Variable-byte codec — implementsAnyLenCodec. Variable-byte codec — implementsAnyLenCodec. Variable-byte encoding codec, generic over element widthT(u32oru64).
Enums§
- FastP
ForError - Errors that can occur when using the
FastPForcodecs.
Traits§
- AnyLen
Codec - Compresses and decompresses an arbitrary-length
&[u32]slice. - Block
Codec - Compresses and decompresses fixed-size blocks of
u32values. - Block
Codec64 - Codec that supports compressing 64-bit integers into a 32-bit word stream.
- Kernels
rust - Bit-packing kernels used by
FastPFor:ScalarorSimd. Sealed. - Pod
- Marker trait for “plain old data”.
Functions§
- slice_
to_ blocks - Split a flat
&[u32]into(&[Blocks::Block], &[u32])without copying.
Type Aliases§
- FastP
For128 rust - Any-length
FastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthFastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthFastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthu32FastPFORcodec with 128-value blocks. - FastP
For256 rust - Any-length
FastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthFastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthFastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthu32FastPFORcodec with 256-value blocks. - FastP
ForBlock128 rust - Type alias for
FastPForwith 128-elementu32blocks. - FastP
ForBlock256 rust - Type alias for
FastPForwith 256-elementu32blocks. - FastP
ForBlock Wide128 rust - Type alias for
FastPForwith 128-elementu64blocks. - FastP
ForBlock Wide256 rust - Type alias for
FastPForwith 256-elementu64blocks. - FastP
ForResult - Alias for the result type of
FastPForoperations. - FastP
ForSimd128 rust - Any-length
FastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthFastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthFastPFORcodecs:FastPFor*foru32,FastPForWide*foru64.FastPFor128usingSimdkernels; byte-compatible with it. - FastP
ForSimd256 rust - Any-length
FastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthFastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthFastPFORcodecs:FastPFor*foru32,FastPForWide*foru64.FastPFor256usingSimdkernels; byte-compatible with it. - FastP
ForSimd Block128 rust FastPForBlock128usingSimdkernels; byte-compatible with it.- FastP
ForSimd Block256 rust FastPForBlock256usingSimdkernels; byte-compatible with it.- FastP
ForSimd Block Wide128 rust FastPForBlockWide128usingSimdkernels; byte-compatible with it.- FastP
ForSimd Block Wide256 rust FastPForBlockWide256usingSimdkernels; byte-compatible with it.- FastP
ForSimd Wide128 rust - Any-length
FastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthFastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthFastPFORcodecs:FastPFor*foru32,FastPForWide*foru64.FastPForWide128usingSimdkernels; byte-compatible with it. - FastP
ForSimd Wide256 rust - Any-length
FastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthFastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthFastPFORcodecs:FastPFor*foru32,FastPForWide*foru64.FastPForWide256usingSimdkernels; byte-compatible with it. - FastP
ForWide128 rust - Any-length
FastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthFastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthFastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthu64FastPFORcodec with 128-value blocks. - FastP
ForWide256 rust - Any-length
FastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthFastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthFastPFORcodecs:FastPFor*foru32,FastPForWide*foru64. Any-lengthu64FastPFORcodec with 256-value blocks.