Protege Raises $10M to Tackle the AI Training Data Bottleneck
Protege has raised $10 million in a seed round to reshape how AI training data is shared and accessed. The goal is straightforward: remove the biggest hurdles that block model builders from finding and using the right data, while giving data holders simple, safe ways to participate. The result, if it works as intended, is a smoother path for both sides—data holders and AI teams—to collaborate with confidence.
Inside the Funding Round
Who Backed the Round
Investors in the round include CRV, SV Angel, Liquid 2 Ventures, and Bloomberg Beta, alongside notable figures such as Adam D'Angelo and Travis May. Their participation signals broad conviction that making access to quality training datasets simpler—and safer—could unlock significant progress across AI.
What the CEO Is Solving For
Bobby Samuels, Protege’s CEO, put it plainly: “The lack of availability of training data is the biggest bottleneck in AI today.” The platform is built to offer controlled access to valuable data resources, so model builders can train better systems without compromising how that data is governed or used.
Rewriting the AI Data Exchange
Today, trading training data is complicated. Deals drag on, often tied up in intellectual property and governance concerns. Protege’s approach is to streamline this workflow. With the right tools and guardrails, data holders can share assets securely, and model builders can source datasets responsibly, reducing the friction that slows progress.
Where the Team Is Aiming
Since its start in 2024, Protege has focused on proprietary data types that matter most for AI. Much of this data sits unused, despite its value. The founders aren’t narrowing in on a single industry; they see room across many sectors, where easier, compliant access to these datasets could lift AI performance and expand what’s possible.
Rising Demand for Training Data
As co?founder Travis May notes, the training data market has changed fast. Five years ago, there was barely a market at all. Now, every major language model—and a wide range of new AI applications—compete for the same finite supply. Reducing friction is essential. In his words, “Unlocking the seamless and controlled exchange of training data will help unleash the power and benefits of AI to the world.”
Why Investors Are Leaning In
Saar Gur of CRV called the training data opportunity one of the largest he’s seen in his career. Strong interest from experienced investors reinforces belief in Protege’s vision—and in the practical need to make trustworthy data access far simpler.
What Protege Does
Protege operates a platform that connects data holders with AI developers. The aim is a working relationship that supports thoughtful, innovative AI without losing sight of compliance. By enabling secure exchange, the company hopes to back the steady, responsible evolution of smarter technologies.
Learn More About Protege
If you want to dig deeper into how the platform works and what it offers, the official platform provides more detail. At its core, Protege’s mission is to bring data holders and AI innovators together so both sides can move faster—and do it the right way.
Frequently Asked Questions
What is Protege’s primary mission?
Protege’s mission is to make access to AI training data seamless and controlled, so data holders can share safely and model builders can source responsibly.
Who invested in Protege’s seed round?
Investors include CRV, SV Angel, Liquid 2 Ventures, and Bloomberg Beta, along with notable individuals such as Adam D'Angelo and Travis May.
What specific challenge is Protege addressing?
Protege targets the toughest part of AI development: exchanging training data. The current process is slowed by intellectual property and governance negotiations; Protege aims to streamline that with the right tools and guardrails.
When was Protege founded, and by whom?
Protege began in 2024. The team includes CEO Bobby Samuels and co?founder Travis May.
Why is training data so important for AI models?
Training data shapes how well AI models learn and perform. As demand grows across major language models and new AI applications, controlled, reliable access to high?quality data has become essential.