A New Chapter in AI Evaluation
It’s about time someone decided to put their money where their mouth is. Snorkel AI is stepping up with a $3 million commitment through its Open Benchmarks Grants, aiming to tackle the elephant in the AI room: how to measure these ever-advancing AI systems when they’re light years ahead of our current tools. The program kicked off in February 2026 and has already attracted a flood of interest, with hundreds of project applications from researchers and engineers looking to push the envelope.
Who's on the Frontline?
Enter the initial wave of projects making the cut. First up, we've got Frontier-Bench, reimagined from Terminal-Bench 3.0. If you're wondering what makes it so special, it's all about the diverse domains and adversarial review approach.
“From complex environments and huge autonomy horizons to sophisticated outputs, these projects tackle some of the field's hardest evaluation challenges,” noted Fred Sala from the grant’s steering committee.
Next, there’s Agents' Last Exam, a mouthful for sure, but substantial in ambition. It's about as exhaustive as it gets, evaluating agents on 1,500 tasks in 55 different sub-industries, with a goal that's over three times bigger. Talk about setting the bar high!
- Frontier-Bench: Successor to Terminal-Bench 2.1, emphasizes diversity and continuous review.
- Agents' Last Exam: Spanning 55 sub-industries, aiming for 5,000 tasks, engaging 300+ experts.
- OSWorld 2.0: Evaluates across various platforms and apps, focusing on long-term workflows.
- Continual Learning Bench: Tests an agent's genuine improvement in sequential tasks.
- SlopCode Bench: Ponders the degradation of code quality over time.
Snorkel isn't just stopping at these benchmarks. They're already neck-deep in Terminal-Bench Science, broadening their scope to computational research in multiple scientific arenas. The breadth of their ambition is something else.
The Backers and Brainpower
Now, who’s feeding this machine? Heavyweights like Hugging Face, Prime Intellect, and others are in Snorkel’s corner, ensuring this whole venture doesn’t end up as mere chat among academics. The funds are rolling, and with project applications reviewed continuously, there's no stop to this train.
Behind the Scenes
Even beyond the grants, Snorkel's joint venture with Princeton and the University of Wisconsin–Madison isn't shabby. They’ve cooked up Senior SWE-Bench to set a high bar for coding agents. Imagine a setup where code isn’t just churned out but crafted to a standard—pinning down runtime bugs and adhering to coding conventions no less.
Strong emphasis is on collaboration with top researchers to ensure the benchmarks are not only relevant but also pushing the boundaries of what's achievable.
"Releases like Terminal-Bench 2.1 are still evolving," integrating corrections and validation with ferocity.
Why Investors Should Watch
What's the big deal for investors here? With AI applications blowing up and the stakes only set to grow, understanding where and how agents perform can lead to massive efficiencies and innovations. Snorkel AI’s early maneuvering positions them to be a crucial player in shaping future AI capabilities.
Whether it’s harnessing data or pioneering programmatic solutions, Snorkel’s commitment today might just dictate tomorrow’s market leaders in AI technology.
Final Thoughts
Dollars are talking, and Snorkel's setting itself up as a hard-to-ignore force. For those keeping tabs on AI's financial landscapes, here’s something that feels not just ambitious—but essential.