Petrarch acquires proprietary datasets from distressed and failed AI startups — anonymizes them completely — and connects them with the labs and enterprises that need specialized training data.
Every year, hundreds of AI startups fail. When they do, years of proprietary data collection — voice recordings, medical images, robotics sensor logs, annotated corpora — goes with them. It's deleted, locked in bankruptcy proceedings, or simply abandoned.
Meanwhile, AI labs and enterprises are starving for specialized, high-quality training data that doesn't exist in public datasets. Petrarch sits at that intersection. We identify distressed companies early, acquire their data assets through structured agreements, run them through our anonymization pipeline, and sell them to buyers who need exactly what those companies spent years collecting.
We monitor 2,800+ failed and distressed AI startups continuously — EDGAR filings, court records, LinkedIn signals, GitHub activity — to find companies with valuable data assets before they're gone.
We acquire the data through structured agreements, run it through our agentic anonymization pipeline (differential privacy, PII redaction, compliance review), and prepare it for institutional sale.
We match datasets with the buyers who need them most — AI labs, research institutions, and enterprises — and handle all licensing, compliance paperwork, and deal negotiation.
Petrarch is backed by Y Combinator. The three founders dropped out of Harvard College to build the first specialized data marketplace.
Questions? Reach us at founders@petrarch.co
Active distress signals
These are companies still operating but showing strong distress signals — we move early to evaluate their data assets before competitors or liquidators get there.
Under-monetized data troves
These organizations often don't realize their data has market value — Petrarch evaluates and monetizes it on their behalf.