Proprietary Data Moat
As AI makes software near-free to produce, durable advantage migrates to data a model cannot reproduce on demand, especially free data that no one ever collected.
What ten thousand geniuses cannot do
vLex, a bootstrapped legal-data company, spent years buying up Spanish legal records reaching back centuries. Then AI demand for structured primary law spiked, and the investor Alex Rampell describes the company going from roughly $20M a year to roughly $100M a year inside a single year, because no language model can regenerate a corpus that had to be assembled one archive at a time.1 Rampell puts the general case as a thought experiment: imagine a "country of geniuses in a datacenter" able to write any code on demand. They can rebuild the software. They cannot run a query against the past.
His other worked examples share that shape. DomainTools runs a daily cron recording ownership of every internet site, building a historical record no one else has: the data was free at the time but, as he puts it, you cannot time-travel and run the query now. FlightAware's aircraft feeds and the aggregations behind FactSet and Bloomberg follow the same logic, mostly-free data that becomes valuable because the aggregation cannot be re-run from scratch. The unifying observation is that the advantage is not that the data is secret or for sale, but that someone started the boring time series years ago.
Ramp as a first-person instance
Eric Glyman described the same moat from inside Ramp two years before Rampell's framing, and named the mechanics.2 Because Ramp powers the card transaction and, for operational reasons, also collects the receipts and invoices around it, it accumulates a vendor-pricing dataset as a byproduct: Glyman says the company can tell a customer they are "paying too much before you pay a vendor like Salesforce." The dataset exists not because Ramp set out to sell data but because its position as the spend rail generates it.
Glyman also distinguishes two network effects that are often conflated. Data network effects run through the price intelligence, where more customers produce better benchmarks, offered on a give-get basis where a customer submits data to see the comparison. Direct network effects run through bill pay, where each payment connects two finance teams and vendor density helps validate that a counterparty is real. His third point is that combination is itself the moat: a single sliver of data cannot automate accounting, but transaction data joined to an org chart and ERP history turns receipt matching from OCR guesswork into a matching problem against raw records.
Where the principle holds and where it thins
The principle names the surviving moat on the other side of no moat in software and its buyer-side companion SaaS apocalypse: code commoditizes while irreproducible data compounds. It also sits alongside marketplace knockoff asymmetry as an argument that the durable edge is rarely the artifact and often the thing that cannot be copied on demand. The stated limits are real. Not all proprietary data is defensible, since anyone can start their own daily cron today, which means the moat is only the historical tail and that tail can erode in relevance over time. And the most valuable spend, location, and identity datasets are also the most exposed to regulatory and privacy contest, so the same accumulation that creates the advantage creates the liability.
Practiced by
Connections
Loading connections…
References
- 01
Ramp's Eric Glyman on How AI Is Changing Corporate Spending
Eric Glyman · podcast
- 02
AI at Ramp (Eric Glyman, MAD Podcast)
Eric Glyman · podcast
Related