By now, we’re accustomed to the simplicity of the AI age.

Have a burning question that requires reading through a hundred web pages? No problem. Get an AI agent working on it. Hours of work condensed into seconds.

And so, the real commodity in the AI age is no longer time. It’s data.

It’s why companies like Scale AI and Mercor have multi-billion-dollar valuations, and why startups are selling their data to brokers for millions.

That’s just the surface of this week’s article. Enjoy.

Also: SF Tech Week is next week. Were hosting six events, including three firesides, two parties, and one dinner. Links below. We’ll see you soon.

  • Founding GPs VC Omakase (10/6, SF) - Join Unicorner and Mercury for a small and highly curated omakase dinner with a room of founding GPs. Spots are extremely limited and subject to approval.

  • The Art of Starting Again (10/8, SF) - Join YG Leboeuf, CEO & Founder of Deck, and Dori Yona, CEO & Founder of SimpleClosure, for an honest conversation about starting again.

  • What Your Company Data is Worth (10/8, SF) - Join SimpleClosure for a practical conversation on how company data gets valued, priced, and licensed to AI labs.

  • Zero to Unicorn with Bolt.new CEO (10/8, SF) - Hear directly from Eric Simons, CEO of Bolt.new, who went from an idea and an early team to building a billion-dollar company.

  • Frequency: A Tech Festival (10/8, SF) - We're hitting some new frequencies tonight. Join us for the tech festival defining SF Tech Week.

  • a16z x Unicorner: The Official Afterparty (10/9, SF) - Join us for the official afterparty of SF Tech Week, hosted by Unicorner, a16z, and Notion. Spots are extremely limited and subject to approval.

Arek and Ethan 🦄

Model builders are reaching the point where they need much more specialized data. A healthcare model, for example, may need longitudinal clinical data that reflects what happens at the point of care rather than another broad corpus of medical text. Protege connects AI companies looking for hard-to-access data with organizations that own it. It starts with the problem the model builder is trying to solve, then finds and prepares the data it needs. It also handles the work between access and delivery, including preparing the dataset and making sure the appropriate rights are in place. For data partners, Protege creates a way to license what they already have without having to build the infrastructure around it themselves.

Check it out: withprotege.ai

Protege helps model builders source data for a specific training or evaluation need, then works with data providers to license it for that use. 

For data owners, Protege creates a new revenue stream from data they already have while giving them clear terms around how that data can be used. While public pricing and take rates are not disclosed, Protege says its standard licensing frameworks are designed to give providers clear terms around permitted use.

  • Raised $65 million since founding, including a $30 million Series A extension led by Andreessen Horowitz in January 2026 following a $25 million Series A led by Footwork VC in August 2025 and a $10 million seed led by CRV

  • Works with the majority of the Magnificent Seven, alongside other AI model builders

  • Its provider network now spans hundreds of data sources across real-world domains

  • Grew its business 20x in 2025 and had already generated tens of millions of dollars in revenue for data partners by its Series A announcement

As models and compute improved, Bobby Samuels and Travis May saw that valuable information was quickly shifting from being available on the open internet to being locked inside private systems and proprietary archives. Both had spent years around privacy-first data environments and saw an opening to build the infrastructure between those data owners and AI companies.

Protege started in healthcare where the problem was especially clear. There was no shortage of real-world medical information, but much of it contained sensitive patient data and sat inside systems that weren’t built for AI development. The team also believed synthetic data could not fully recreate what happens at the point of care or inside healthcare operations.

Before it could be used, the data had to be de-identified and licensed under clear terms that protected patient privacy. That made privacy part of the product before there was much of a product at all. Protege completed a privacy review before writing its first line of code and built its licensing model around giving data owners control over how their information could be used.

❝

“We believe in a world where data holders are compensated, privacy is respected, and AI builders can get access to real-world data that reflects how people actually live and interact.”

Bobby Samuels

As model builders began asking for different kinds of real-world data, Protege followed that demand into new categories.

The public internet gave the first generation of modern AI an enormous amount of training material, but the data needed to improve models from here lives somewhere else. Consider medical records, which are locked inside healthcare systems. Or the years of proprietary content sitting in private media company archives that AI companies cannot easily scrape or train on. In many domains, much of the context behind real-world work was never published online at all. Protege is making that information usable for AI without stripping away the rights and protections around it.

The problem is bigger than access. Model builders are becoming much more specific about where their systems fail, which means handing them a massive dataset is often less useful than finding the narrow slice of data they directly need. Protege starts by understanding what the AI team is trying to improve, then works backward to figure out what kind of data would help the model perform better, sometimes combining sources from different providers or using its DataLab researchers to determine what information will improve the model.

Protege has to understand the model well enough to know what data the buyer should be asking for, then make that data usable. That’s a task that goes beyond the job description of a standard data broker. The largest AI labs have the money to negotiate directly with data owners, so Protege becomes harder to replace when going through it is easier than sourcing and preparing the same data internally.

Protege connects AI companies with organizations that own hard-to-access data, preparing and licensing it so models train on real-world information.

Half the battle is convincing organizations to make valuable information available at all. Protege gives data owners a way to license specific uses of what they own without transferring ownership of the underlying data. The terms can vary by deal, with Protege structuring the permitted use and revenue share around the dataset and the buyer’s needs. In media, for example, a license can allow content to be used for model training while still prohibiting redistribution or reproduction of the original material. In healthcare, a provider can license de-identified data for a defined model-development use while still restricting re-identification or redistribution.

❝

“We also take extensive steps to ensure that privacy, data protections, and de-identification is done correctly for each delivery, and validate the dataset(s) against the agreed requirements to then securely deliver it in the format that the specific AI builder or team requires. This can vary by domain, but the end goal is consistent across the board: we turn raw, fragmented real-world representative data into licensed, trustworthy, and AI-training or evaluation-ready datasets while protecting and compensating our trusted data providers."

Bobby Samuels

If model performance increasingly depends on information residing in private systems, owning the relationship between the people who need that data and the people willing to license it becomes a massive business.