New to Glean: podcasts 🎙️ Listen below to hear what’s new. Don't forget to like and subscribe for more from Glean 🤭
Glean
Software Development
San Francisco, CA 234,793 followers
The Work AI platform connected to all your company data. Find, create, and automate anything.
About us
Work AI for all. The Work AI platform connected to all your data. Find, create, and automate anything.
- Website
-
https://glean-it.com/glean
External link for Glean
- Industry
- Software Development
- Company size
- 501-1,000 employees
- Headquarters
- San Francisco, CA
- Type
- Privately Held
- Founded
- 2019
Products
Glean
Enterprise Search Software
The Work AI platform connected to all your company data. Find, create, and automate anything.
Locations
-
Primary
Get directions
San Francisco, CA 94107, US
Employees at Glean
Updates
-
Glean reposted this
Glean just published agent unit economics against Claude Cowork, and the numbers decompose beautifully!!!! $0.45 per task vs $1.84. A 4x gap. Most people will read the headline. I did the math underneath it, and that is where it gets interesting. Cost per task is just two variables: tokens consumed x blended rate per million tokens. Glean's advantage splits cleanly across both. Variable 1: token efficiency. Glean averaged 688K tokens per task. Cowork averaged 2.0M. That is a 2.9x delta on the same work. Across the full benchmark, 29.8M tokens vs 88.8M. This is not the model being verbose. It is retrieval precision. Glean's index feeds the agent tighter context per step, so prompts carry less dead weight and trajectories converge in fewer turns. Context engineering shows up directly on the meter Variable 2: blended rate. Work it out from the published numbers. $0.45 over 688K tokens implies roughly $0.65 per million tokens. $1.84 over 2.0M implies roughly $0.92. That is the 1.4x rate advantage, and it comes from routing policy. 2.9 x 1.4 = 4x. The gap is compounding, not additive. Now the part that breaks intuition. Glean ran Opus, the most expensive tier, on 29% of its token volume. Cowork ran it on 2.8%. The cheaper system used 10x more of the priciest model. That only makes sense if you optimize cost at the task level, not the token level. Glean routes commodity steps to cheap models, including Luna, which it says runs 10x cheaper than Claude Sonnet, and reserves frontier reasoning for the steps where it changes the outcome. Expensive tokens, deployed surgically, lower total cost. Compare the distributions. Cowork: 96% Sonnet family, essentially a single point on the capability-cost curve. Glean: 64% GPT-5.6 family, 29% Opus, 7% other, a portfolio spread across the frontier. One is a model strategy. The other is a routing strategy. The benchmark suggests routing wins on economics. Zoom out and this is the real story. For two years we benchmarked intelligence. Now we are benchmarking harnesses, because at 100K tasks a month this gap compounds to roughly $1.7M a year. That number moves platform decisions. Standard caveat applies, this is Glean's own benchmark, so pressure test it against your task mix. Glean says more results land at Glean:GO. I am excited to be there in person next to next week!!!! The routing layer is becoming the moat. If your agent stack runs one model for everything, you are leaving margin on the table in both directions. What does your cost per task look like? Genuinely curious what others are measuring. #data #ai #benchmark #tokens #glean #api #gleango #theravitshow
-
-
Glean reposted this
We've been investing heavily in our harness and routing capabilities, and we put them to the test benchmarking Glean's token costs against Claude Cowork. The results were striking: Glean is 4x more cost-effective, averaging $0.45 per task versus $1.84 for Claude Cowork. That 4x advantage comes from two things compounding: 2.9x lower token volume and a 1.4x cheaper blended rate per million tokens. Here's how: - Model family routing: Glean made use of Luna which is 10x cheaper than Claude Sonnet and widely capable. We’re able to strike the balance by routing between open and closed models. - Model tier routing: In Glean, Opus was used 10x more (29% vs 2.8%) but surgically for the right things and balanced by other models. - Better context: Glean’s harness and indexing capabilities result in fewer tokens consumed; Claude Cowork used 3x the tokens per query on average using 88.8M versus 29.8M in Glean. More results coming out at Glean:GO! Hit me up if you're still looking for an invite.
-
-
Glean reposted this
Looking forward to being back on the Glean:GO stage this year, I’ll be digging into a problem every organization is starting to feel. AI is spreading across the enterprise, and the stack is getting harder to manage. More models. More front doors. More tokens. The upside is enormous. So is the bill. AI is going to have many front doors. Enterprises need a control layer underneath them that keeps the stack open, routes each task to the right model, and gives leaders visibility into what they’re spending and what they’re getting back. At this year's event, I’ll get into why open models matter, how to match the right model to the job, and how to get more work done without burning through your token budget. Join us at Glean:GO in San Francisco or online Aug 26–27. https://lnkd.in/grtz5uvF
-
-
Glean reposted this
For an industry giant like General Motors, the real challenge of AI isn’t just identifying high-value use cases. It’s executing them across tens of thousands of employees. Following a successful proof of concept, GM selected Glean as its core AI platform and scaled from 0 to 60k users in 90 days, bringing AI into its mission-critical workflows. That kind of velocity requires more than a software rollout. It requires a connected AI architecture that works with the systems employees already use, giving teams the context and governance to improve existing workflows, identify new ones, and build new ones without creating fragmentation or AI sprawl. At Glean:GO, I’ll be on stage with Philip Luedtke from GM to discuss the decisions behind that approach, including how GM prioritized high-value workflows, enabled teams to build without creating fragmentation or sprawl, and maintained security and governance as adoption grew. Join us at Glean:GO on August 26–27 in San Francisco or virtually to hear directly from Philip and other leaders at some of the world’s most innovative companies about how they’re using Glean to transform how work gets done across their businesses: https://lnkd.in/eK74AJKF
-
-
🌟 Meet our marquee Glean:GO sponsor: Deloitte You’ll hear from Deloitte leaders Ashish Verma, US Chief Data and Analytics Officer, during our keynote, plus hear from Ajay Tripathi, Deputy Chief Data and Analytics Officer, and Rohit Balasubramanian, Managing Director, in breakout sessions on what it takes to build agents around real business processes. Expect practical conversations on AI-first enterprises, business processes, and building agents that can get real work done. There’s still time to sign up for #GleanGO: https://lnkd.in/gqfzDYKM
-
-
Glean reposted this
For two years, the AI market ran on a stable bargain: closed models won on quality, open models won on price. But that broke in one week at the end of July, when Moonshot's Kimi K3 matched frontier reasoning as an open weight model, and three days later OpenAI cut GPT-5.6 Luna's price by 80%, to $0.20/$1.20 per million tokens, cheaper than almost every open model you can rent. The interesting part is how OpenAI got there. They pointed Sol, their flagship model, at their own inference stack. Sol rewrote its production serving kernels, found compute that could be skipped or parallelized, and cut end-to-end serving costs by 20%. Then it redesigned its own speculative-decoding draft model, the small model that proposes tokens for the big one to verify, and improved token-generation efficiency by another 15%. Stacking those together the math it works out to roughly 30% lower cost. But Luna's price fell 80%. The other 50 points might be OpenAI eating into its own margin to own the low end of the market. A lab that owns both the frontier model and the stack underneath it can keep running this loop. "Open is cheaper" and "closed is better" both broke definitively last week. These days it seems like any confident take on the model landscape has a shelf life measured in days. I think this is exactly why the application layer wins here. Glean has been model-agnostic and harness-agnostic from the start, with real infrastructure for switching. When the input layer is genuinely agnostic, a price cut or a new open release is just a config change. We felt this ourselves the week Luna dropped: we reclassified it and passed the savings straight through to customers almost immediately, because the plumbing to do that already existed. Intelligence is starting to look like a commodity market. The value is migrating to whoever sits between that volatile supply and demand: understands the customer's workloads deeply, moves fastest when the ground shifts again next week, and passes the quality and price gains straight through.