Qdrant’s cover photo
Qdrant

Qdrant

Software Development

Berlin, Berlin 64,620 followers

Composable high-performance vector search

About us

Powering the next generation of AI applications with advanced and high-performant vector similarity search technology. Qdrant is an open-source vector search engine. It deploys as an API service providing a search for the nearest high-dimensional vectors. With Qdrant, embeddings or neural network encoders can be turned into full-fledged applications for matching, searching, recommending, and much more. Make the most of your Unstructured Data!

Industry
Software Development
Company size
51-200 employees
Headquarters
Berlin, Berlin
Type
Privately Held
Founded
2021
Specialties
Deep Tech, Search Engine, Open-Source, Vector Search, Rust, Vector Search Engine, Vector Similarity, Artificial Intelligence , Machine Learning, and Vector Database

Products

Locations

Employees at Qdrant

Updates

  • View organization page for Qdrant

    64,620 followers

    What happens when engineers are given a controversial topic and told to debate it? We’re bringing something a little different to San Francisco Tech Week. Join Qdrant × Neo4j for Hard Negatives: Engineers Debate Night on October 9. We’ll put engineers, AI researchers, and builders into 1v1 debates on spicy topics like: “Open weights should be banned.” The audience votes live on who wins each round. Plus, your RSVP includes food and one drink. Venue: Manny’s, San Francisco October 9 | 6:30–9:00 PM PDT Spots are limited and registration requires approval. Link: https://lnkd.in/gwRBggGx

    • No alternative text description for this image
  • Qdrant reposted this

    Got a connection request this weekend with a remarkable claim: a "zero-dep C99 vector engine" with 33k QPS, 31 ns IPC, and 512x less RAM than FAISS. Reproducible in 60 seconds. 🤔 So I read the code. Ok, ok, I let Claude read the code, and here's what's under the hood: - 512x less RAM: every 512-dim vector stored as a single 32-bit hash. That's not compression, that's amnesia. - 31 ns IPC: one process copying 64 bytes into its own memory and reading them back. Inter-process communication, minus the second process. - 33k QPS: a popcount loop that finds the true nearest neighbor 0.1% of the time. Random guessing gets 0.002%, so to be fair, it's 50x better than random. Then I asked about recall. The answer: recall@10 = 1.0, exact search, same 33k QPS, now on 1M vectors. On a 6-vCPU VPS. For 768-dim embeddings, that's ~50 TFLOP/s of brute force. An H100 would like to know your secret. Bonus: the same "benchmark stand" validates cold fusion calorimetry with assert 2.3 == 2.3 and prints "OK". Link to the repo in comments, if someone is interested. But be aware, there is more brainfuck included, even some smart contracts... 🤯

    • No alternative text description for this image
  • View organization page for Qdrant

    64,620 followers

    In finance, a wrong number in the wrong place costs millions. That makes "mostly right" a useless standard for an AI answer. "Traceable to the filing" is the one that counts. LucyData builds data infrastructure for financial AI: 2.8M+ SEC EDGAR and Korea DART filings, 27 report types, 10 years of history, preprocessed so financial tables keep their original row and column structure and every chunk links back to its source document. The retrieval layer runs on Qdrant. Every question carries structure (company, CIK, form type, period, document ID), and those payload filters apply inside the query rather than as a post-filter, so the search is scoped to the exact filing before ranking starts. Four lanes each run the same dense-plus-sparse hybrid query, fused with reciprocal rank fusion. What that produced: - 94.7% Recall@10 on FinanceBench - +0.160 answer accuracy from the retrieval layer alone, 0.792 to 0.952, with the answer model held fixed - An open-weight model on Lucy RAG context at 0.863, ahead of frontier models on web search - TurboQuant cut the 10-K index from 52 GB to 10.7 GB, with hybrid retrieval quality effectively unchanged One detail worth stealing: LucyData first kept every SEC filing type in a single collection, with metadata prefixed into the chunk text. As coverage grew, latency rose 2x to 3x and recall fell. Splitting into collections by filing type and moving that metadata into indexed payload fields fixed both. Twelve collections, roughly 52M dense and 52M sparse vectors, self-hosted, run by fewer than 10 engineers. "Finance is complex, and AI models can sometimes return incorrect answers to specific financial questions. We believe those answers need a source of truth, and for public-company financial data, that source is the original SEC filing." - Jihoi Park, Co-Founder, LucyData https://lnkd.in/gDbi5M-K

    • No alternative text description for this image
  • Qdrant reposted this

    We’re back this Friday 👀 AI Debate Night is coming back for round two with Qdrant and Neo4j. I’ll be hosting again, and we’ve got another set of 1 vs 1 debates coming. We’ll start announcing some of the matchups and topics soon. If you’re in SF, come hang out, vote on the debates, and grab some food and a drink. https://lnkd.in/gBEdA6Dg https://lnkd.in/gxRzqSuW

    • No alternative text description for this image
  • Qdrant reposted this

    Yes! You can swap query models without re-embedding! We did some research for you to make it work. Ok, ok.. slowly. So, a document vector gets reused across thousands of searches. A query vector gets used once and thrown away. Most retrieval stacks still run the same big model on both, and switching to something cheaper on the query side means re-embedding the whole collection, right? Nope, not anymore. Meet Constella, our new research preview, which splits the two. Stella, a 400M English embedding model, encodes your documents once. Then you pick the query encoder per request, and all three search the same Qdrant collection: - Full Stella: 38.95 ms per query on CPU - Nano Stella: 34.5M parameters, 3.13 ms - Zero Stella: 0.081 ms, 480x faster than Stella 😮 Ok, there is a tiny catch. Zero is a learned bag of tokens: look up, pool, normalize. That's the entire encoder. It has no idea about word order, so "dog bites man" and "man bites dog" are the same query. But paired with BM25, it gives you a hybrid search with no transformer in the query path, and this unlocks several use cases such as: offline search on low-power devices; hybrid search as you type, high-volume retrieval APIs, and agentic workflows. Isn’t it Stellar? We think so. Details in the preview blog by Dylan Couzon 👏 https://lnkd.in/d5kik22Q

    • No alternative text description for this image
  • View organization page for Qdrant

    64,620 followers

    What if you could change your query embedding model without re-embedding all your documents? That’s the idea behind Constella, a research preview from Qdrant. Normally, your documents are embedded with a model and stored as vectors in Qdrant. When a user searches, the query needs to be converted into a vector in the same vector space so Qdrant can compare it with the stored document vectors. The problem is that changing the embedding model usually means re-embedding your documents too. Constella takes a different approach here: Documents → Stella → document vectors → Qdrant Then, for queries, you can use different models: Query → Stella / Nano / Zero → query vector → Qdrant Nano and Zero are specifically trained to produce vectors compatible with Stella’s document embeddings. So you can switch between them depending on the trade-off you want between quality and query latency, without rebuilding the document index. For example, Zero is extremely lightweight, while Nano is a small transformer that gets much closer to Stella’s retrieval quality. Constella is still a research preview, so there’s plenty more to explore around the quality/latency trade-offs and real-world workloads. https://lnkd.in/dYKDyrq7

    • No alternative text description for this image
  • Qdrant reposted this

    🚀 Two simple Qdrant filters. Two very practical problems solved. In my latest blog, I explore two interesting filtering capabilities in Qdrant 1.19: Prefix Match and Slice Filter, using the MedQuAD dataset as a practical example. 🔍 What will you learn? 👉 How Prefix Match helps retrieve payload values based on exact prefixes, without relying on text tokenization. 👉 How Slice Filter deterministically divides a collection into stable partitions, making parallel processing and reproducible sampling much easier. 💡 Why do these matter? Because vector search systems are not only about similarity search. In production RAG systems, we also need better ways to filter, partition, process, re-embed, and evaluate large collections reliably. The blog walks through both concepts with architecture, code, and practical examples. 📖 Read the full article: https://lnkd.in/d5CQxx7q #Qdrant #VectorSearch #RAG #GenAI #VectorSearch #LLM

  • Qdrant reposted this

    A few milliseconds can hide a lot of engineering. While working with Qdrant at 100M vectors, I ran into a search performance bottleneck that wasn't obvious from the default configuration. So I started digging into the search path - measuring where the time was going, changing one variable at a time, and validating each change against recall. What started as “why is this query so slow?” eventually turned into a 200x+ throughput improvement. I documented the investigation and the lessons from tuning vector search at scale: 🔗 https://lnkd.in/ebYHEHqh #Qdrant #VectorSearch #PerformanceEngineering #DistributedSystems #AI

Similar pages

Browse jobs