I re-evaluated 15 of today's top VLMs on VLADBench, a driving scene VQA benchmark from 2025, and here is what I learned: • VLMs have gotten a lot better. The field mean improved across all tasks with one exception. • Frontier capability is going to cost you. GPT-6 Astra (83.6, $542.16) and Claude Opus 5.5 (82.7, $205.28) sit at the frontier of visual understanding. • Gemini 3.8 Flash (75.9, $16.93) is a goldilocks pick on cost-performance but fails to lead the field on any task. • Also, building benchmarks for VLMs is hard. We're releasing the results of this VLADBench re-evaluation with the full code and the dataset on Hugging Face. Link in the comments!
Daft
Technology, Information and Internet
San Francisco, CA 2,387 followers
Distributed query engine providing simple and reliable data processing for any modality and scale.
About us
Daft is a high-performance data engine providing simple and reliable data processing for any modality and scale, from local to petabyte-scale distributed workloads. The core engine is written in Rust and exposes both SQL and Python DataFrame interfaces as first-class citizens. Solving the fundamental challenge of working with multimodal data at scale and powering the next generation of AI applications, we are eliminating the traditional barriers and redefining how developers interact with multimodal data. Try Daft today: pip install daft
- Website
-
https://daft.ai/
External link for Daft
- Industry
- Technology, Information and Internet
- Company size
- 11-50 employees
- Headquarters
- San Francisco, CA
Locations
-
Primary
Get directions
San Francisco, CA, US
Employees at Daft
Updates
-
Some people think we're running out of data. Sammy from Eventual says it was never about volume though.... Use something like Claude today and you get output in the top caliber of software engineers on certain tasks. Pretty good. But the labs are pushing further. And further. Here's what nobody sees. As you climb the stack of quality, sourcing that data becomes exponentially harder. The last mile isn't a little harder. It's a different problem entirely. Which is exactly the problem Eventual exists to solve.
-
Some people think passion means 996. Sammy of Eventual defines it completely differently. You only have 24 hours in a day. Passion doesn't mean grinding around the clock. Passion means it's always running in the background thread of your brain. You're always thinking about it. You're immersed in the problem. And here's the distinction. If you're solving problems that come up every day, where it's about grinding them out? You don't necessarily need passion for that. But when you're solving problems that have never been solved, when you're on the frontier, those problems have to live in your mind constantly. And that only happens if you're genuinely passionate about it.
-
Why does an offline VLA eval need a WebSocket server and multiple Python runtimes? 🐮 It doesn’t. I ran OpenVLA + LIBERO across all 10 Spatial tasks in one Python 3.12 process and was surprised at how elegantly it composed. 🪶 Some of the apparent difficulty of scaling Physical AI evaluation is accidental complexity inherited from how research repositories are packaged. Now that I've got one rollout nailed down, adding new policies and LIBERO variants is a matter of stamping concrete implementations. The real work begins with scaling. I'll follow up with VLA-JEPA and LIBERO-Pro, then move into new scaling regimes for faster benchmarks. Can't wait to begin hill climbing CPU + GPU saturation until the cows come home. 🐮 link in the comments!
-
-
Some engineers are paid to write code. At Eventual our engineers are paid to find the problem. What separates them? Number one: intellectual curiosity. Good engineers are systems thinkers. They love going deep into a problem. Rabbit holing. Never afraid to go lower and lower level until they've uncovered the real problem. Number two, and this is becoming everything: Agency. Because the coding part of the job matters less and less now. What matters is uncovering the right problem to solve. That's the core of engineering. First find the right problem. Then solve it in a deep, critical way.
-
A teleoperated robot episode doesn't start when the recording does - the operator is still setting up while the camera rolls. The robot's own joint positions tell you which frames matter without decoding a single pixel. With this method, trimming all of DROID - 500 hours of robot data - took 32 seconds on a laptop. https://lnkd.in/gthAb6zN
-
-
Zuck is spending tens of billions on GPUs. Here's what actually happens. Someone understands the algorithms. Or someone understands the data. They train a model. And they're using 10 percent of the hardware. The rest of that utilization? Money lit on fire. So the engineers getting the big bucks understand how to build the algorithms, how to run the training, how to sort the data. And how to make all of it efficient. That gap is the reason Eventual exists. Everyone can build the model. Almost nobody can make the machine earn its cost.
-
Some engineers chase what's hot. Sammy, the founder of Eventual, says the rarest skill is perspective. Engineers get caught up in what's hot right now. What's hot for the next three months. They're not taking the long view. He sees it constantly. People look at a startup's revenue growth curve and jump on the bandwagon. RAG got hot, and suddenly every engineer he knew needed a RAG startup. That collectivist mentality. Most people follow the pack. The real work is digging in and asking: What am I actually passionate about? What do I want to be part of? Where is this industry going? The three-month view is crowded. The ten-year view is empty.
-
Some engineers spent a decade building web apps. Physical AI, the problem Eventual is built around, makes that look small. Software engineering became synonymous with building a web app. A full stack application. In physical AI, you're not building a web app. The product is physical intelligence. Systems that move and change things in the real world. Web engineering runs on server-grade hardware. Cloud instances. Comfortable. Physical AI runs somewhere else. NVIDIA hardware. Embedded systems. A completely different set of constraints. And that is very fun engineering. Any software engineer should look at it and think: Hell yes. That's the environment I want to work in. Because those are challenges you'd never get to touch anywhere else.
-
Some engineers optimize inside the box. Sammy built Eventual to make the box bigger. In software, constraints make sense. But in physical AI, with GPUs, that frame breaks. Because if you make something more efficient, you don't just save money. You buy room for higher quality. Train more efficiently, train a bigger model. Improve inference, put a bigger model on the robot. So why optimize within the constraints you're given? Why not make the constraints better? That's where young engineers fall short. They look at the box they're handed and interpret it as a box. The move is to open it. To ask, given every dimension I have, what can I actually exploit here?