In finance, a wrong number in the wrong place costs millions. That makes "mostly right" a useless standard for an AI answer. "Traceable to the filing" is the one that counts.
LucyData builds data infrastructure for financial AI: 2.8M+ SEC EDGAR and Korea DART filings, 27 report types, 10 years of history, preprocessed so financial tables keep their original row and column structure and every chunk links back to its source document.
The retrieval layer runs on Qdrant. Every question carries structure (company, CIK, form type, period, document ID), and those payload filters apply inside the query rather than as a post-filter, so the search is scoped to the exact filing before ranking starts. Four lanes each run the same dense-plus-sparse hybrid query, fused with reciprocal rank fusion.
What that produced:
- 94.7% Recall@10 on FinanceBench
- +0.160 answer accuracy from the retrieval layer alone, 0.792 to 0.952, with the answer model held fixed
- An open-weight model on Lucy RAG context at 0.863, ahead of frontier models on web search
- TurboQuant cut the 10-K index from 52 GB to 10.7 GB, with hybrid retrieval quality effectively unchanged
One detail worth stealing: LucyData first kept every SEC filing type in a single collection, with metadata prefixed into the chunk text. As coverage grew, latency rose 2x to 3x and recall fell. Splitting into collections by filing type and moving that metadata into indexed payload fields fixed both.
Twelve collections, roughly 52M dense and 52M sparse vectors, self-hosted, run by fewer than 10 engineers.
"Finance is complex, and AI models can sometimes return incorrect answers to specific financial questions. We believe those answers need a source of truth, and for public-company financial data, that source is the original SEC filing." - Jihoi Park, Co-Founder, LucyData
https://lnkd.in/gDbi5M-K