Thesis

The AI CRM is the wrong bet

AI doesn't need a better CRM. It needs a system of its own.

By Kal PrinceFounder of Ghost. Close to ten years in enterprise go-to-market before going full engineer brain.19 min read

00The bet

Right now the best funded idea in sales tech is that the CRM should be rebuilt with AI inside it. Over the last eighteen months several of the most talked about companies in the category have launched on more or less that same mission.

The first AI-native CRM.

The only AI-native go-to-market platform.

Different teams, different architectures, and underneath all of it one bet.

It's a serious bet, made by people who know this market well, and the pain they're pointing at is real. I've felt it myself.

I still believe it's wrong.

And I think the reason it's wrong points at something a lot bigger than the CRM, which is that this technology wave doesn't need a better record for people. It needs a purpose-built system for AI to think and reason.

Machines needing their own systems isn't a new idea. It's the premise under Exa, Parallel and Browserbase, and it seems to be working. But look at what those companies built: a second system beside the human one, with Chrome and Google left running untouched. Nobody replaced anything. The AI CRM had the same insight and drew the opposite conclusion.

The difference is what the incumbent holds. The web exists whether or not Google does, so a machine-native search engine only has to build a new way in. Go-to-market doesn't work like that. Every account, every deal, every contact the rest of the company points at lives inside the CRM, and comp and forecasting and routing are wired straight into those records.

Browsers and search were unowned surfaces. Every wire in the company runs back to the CRM.

Two records side by side. The record of revenue holds fields with single values: amount, stage, close date, owner. The record of context, outlined in violet, holds a quoted claim from a champion scored 0.94 and marked extracted, plus an inferred claim marked held rather than written.
The record of revenue holds a value. The record of context holds a claim, with a source, a time and a score.Swipe the diagram to see all of it.

Now on to the thing that drives disruption, and it already has a name. People have been calling it a context graph since December and I don't have a better name for it, so I'm not going to pretend I invented the vernacular. The name is the least interesting part. Very few people are writing about what it takes to run one, or how it fits into a stack somebody already paid for. And after building one with customers who still use a CRM, my honest view is that the gap between the idea and the working system is enormous. That gap is where the category gets decided.

So the argument comes in two parts, and they're really the same argument. First, the reasons replacing the CRM fails right now, in roughly the order I came to believe them. Then the more interesting part, which is that every one of those reasons is also a reason a context graph wins. Every wall in front of the replacement is a tailwind for the newer layer.

01I thought I had my founder moment

It was early 2024 and I thought I had my founder moment.

I'd wanted to start a company since 2020. You hear the same thing over and over, which is that founders have some visceral pain or lived experience they can't get out of their head, and I'd been waiting around to feel it.

I had started a new sales job, I received my territory assignments for the year, and I immediately couldn't see a path to hitting my number.

At that point I had close to ten years of selling behind me. I'd carried a number as an individual contributor and hit it every time, and I'd also run a team, which meant running a forecast every week and owning compensation for the people on my team. So I'd seen these systems from both sides, as the person filling them in and as the person depending on whatever got filled in, and at this point in my career I had also become much more in tune with the realities of product-market fit for specific customer types. And when I received the territory plan, I started crunching the numbers on how I'd ended up with it. My conclusion was that the platform and the data were the problem. Bad segmentation, stale fields, an account model that didn't reflect anything real about the business or the opportunity to sell specific products.

So I left after one year in the job, and I built for that. The first version of the company was an ICP product, using AI and segmentation to fix the data underneath go-to-market and put it somewhere a non-technical person could actually use it. This was early days of AI coding products and we came across as a feature rather than a solution for our buyer type.

We quickly pivoted into an outbound product with sophisticated personalization levers rooted in context engineering. We thought that this was an easier distribution path and it pointed at the lack of quality we were seeing in AI generated emails. It got far enough to get people interested with a few pilots, which is its own kind of trap, because interest is not the same as need and it takes a while to learn the difference.

What actually moved me was one comment from a chief product officer on a networking call earlier this year, where I showed him an outbound automation I built. He told me that enterprise sellers don't care about meetings. They care about revenue.

I'd been in that seat for ten years and I still needed somebody to say it to me plainly. Every enterprise seller I ever worked with already had a calendar full of meetings. More of them was never the constraint. The constraint was always leveraging existing relationships to navigate to decision makers and showcasing how your product best aligns to their jobs to be done. Which meant every touchpoint needed to be a strategic chess move that got me closer to a deal. The higher up the decision-making chain you went, the more true that was.

And if you're working on existing customers, or prospects where a relationship already exists, outbound is simply not a strategic lever.

Once I'd heard it I couldn't stop applying it backwards.

There had been a conversation I had about a year prior, with someone who used to be C-suite at a company I'd worked at for several years. I'd had a large book there, north of $30M, at a company doing roughly $250M a year. Even though the existing business was large, my team was really only responsible for $4M to $5M in annual growth. Meanwhile there were a handful of partnerships driving somewhere between $50M and $100M a year, extremely volatile, with basically all of the company's upside residing there. He told me those were the only deals that were going to get the company to its number and that nothing else really mattered. He was right.

That's the same lesson coming from the other end. My territory didn't matter to the number. My meetings didn't matter to the number. What mattered were our most important existing relationships and the greenfield partnerships that looked exactly like them. This is a context problem meeting a consultative selling problem, not a CRM problem, and definitely not an enrichment or outbound problem.

So I'd spent a year building a product for the pain I personally felt instead of the outcome the people I was selling to were actually measured on. Ouch.

That's the inherent mistake that I believe comes with going after the incumbent CRMs.

Anybody who has carried a number knows the feeling. The interface is bad. The data entry is a tax. Nothing about opening a CRM feels connected to closing a deal other than having one single pane of glass.

The complaint isn't the problem. What gets built from it is. A CRM is bad at the job its users think it has and good at a job its users never see. The rep experiences a form. The company experiences the one place where territory routing and compensation and the forecast and renewals all agree with each other. Those are two different products wearing the same name, and the person best placed to describe the pain is the person least placed to see what the thing is actually there for.

Now add LLMs to it. What if the model did all the data entry? And once the entry is handled, what if it ran your territory plans, and built account plans every day of the year, and read third-party data to fill in the records nobody ever got around to filling in? Everyone wants that. I wanted that. It just isn't what the CRM was ever there to provide.

02Ninety days and the pressure of capitalism

The first reason is the one no product argument survives.

A company isn't an architecture diagram. It's a group of people who have ninety days to hit a number and then another ninety after that. Replacing the CRM doesn't help anybody hit it. It's a cost that has to be outweighed by revenue you can see this quarter, and what's being sold right now is time savings. You're selling to somebody who can't afford to wait until next quarter, and a change of that size going wrong is the kind of thing that gets you fired.

I'm not claiming enterprises never replace software. Obviously they do. They do it when the replacement produces revenue immediately or the ROI outweighs the switching cost. Administrative relief doesn't produce revenue inside a quarter, it produces a migration, and a migration is what you end up explaining to your board instead of explaining growth. It reads as unfocused, which is its own kind of expensive.

It's why the CRM keeps getting renewed by the same people who complain about it constantly. The complaint and the renewal aren't in conflict. They're both rational.

And that's just the top of the market. The bottom is closed for a different reason. There's no wiring to rip out at a fifty-person company, but there's already a cheap CRM that takes an afternoon to set up, built by a company whose entire business is winning exactly that customer, and it's running the same playbook as the enterprise incumbent, going full tilt on AI to make its own platform easier to stay on.

It's good enough down there for the same reason Salesforce is good enough up here, because holding the record is a solved problem at both ends. The enterprise is protected by switching costs. The long tail is protected by distribution and a price you can't undercut by enough to matter. Two different moats, and the AI CRM has to beat one of them before it gets to be a category.

03What the record is holding

Underneath all the complaining about the CRM, one thing gets lost. It was never built to drive revenue. It was built to track it, and that isn't a design failure, that is the design. Everything downstream is allowed to depend on it precisely because it is a boring, predictable place where numbers go to be counted.

The second reason comes down to determinism, the thing only deep operators who have engineered these systems talk about. In its simplest form it means X happens and you get Y. I run this report against these accounts and I get this forecast, every single time. Sellers feel the pain of the fields. They almost never see how many other systems are reading them, and the deeper I go as a developer, the more I realize how important this concept is for creating effective business systems.

It's worth being precise about which fields those are, because the CRM is two different things wearing one name here too. The revenue data in it is reliable. Amounts, stages, close dates, bookings. Not because the software is careful but because there's a close every month, people whose job is to make those numbers true, and a finance team downstream who will notice inside a day if they aren't. Money forces accuracy. Everything else in the record is optional, and optional data rots.

Nobody keeps Salesforce because they like the UI. They keep it because territory routing is wired to it, and so is compensation, and the forecast, and marketing automation, and the renewal triggers, and the numbers that go to the board every quarter. Most of the people who built those wires have left. Some of them are a script somebody wrote in 2019 that has run every night without anybody thinking about it again.

Take the record away and none of that fails loudly. It fails quietly, in the middle of a quarter, in ways nobody can list in advance because nobody knows the scope of impact. People who have been around enterprise software know this cold, which is exactly why they're so slow to move. The juice isn't worth the squeeze.

Here's the part that took me much longer to see. All of that wiring is held together by people. Integrations break weekly. Fields drift. A sync fails on a Tuesday and nobody notices until the following month. There are humans employed full time to keep the record connected to the work, and most RevOps teams are doing it with popsicle sticks and glue.

That's the labor this generation of models replaces. And it also makes the case stronger for the incumbent platforms (Claudeforce).

The other half of the record is a mess, and every senior go-to-market person knows it even though nobody writes it down. Contacts who left two years ago. Accounts segmented by somebody who isn't at the company anymore. A decision maker field that was accurate the week it was typed and never again. None of that shows up in a close, so nothing forces it to be true. Big companies deal with it by hiring people to clean records that go stale the moment somebody touches them. That's the job an LLM can finally do. Not a rules script from 2019, but something that reads the call and the email and the invoice and works out what the field should have said, continuously, without a ticket. The mess stops being something you manage and becomes something you fix, and once it's fixed you can build on top of it, which is the part almost nobody gets to today, because everything they'd build would sit on data they don't trust.

Which is really the shape of the whole argument. The part of the record that money forces to be accurate is the part nobody should be touching, and the part nobody audits is the part that can actually drive deal flow.

04The question nobody asked

Both of those reasons would be true if AI had never happened. But look at what they have in common. Nobody up-market can afford the disruption, nobody down-market needs the replacement, and everything downstream is wired to the record. Each one is an argument against replacing the record, and each one is an argument for putting something next to it, which is why I think the case against the AI CRM and the case for a context graph are the same case told from two directions.

The route to that system doesn't run through owning the record first. It runs through the stack buyers already have, because the only version of this a serious buyer says yes to is one they can turn on now, beside what they've already paid for, that makes all of it work better. A company that bet on replacing the record can't sell that way. Its own premise won't let it.

So the question was never how to make the tracker nicer. It points somewhere else, at something I don't see anyone in this debate actually asking.

Every field in a CRM is a fossil of a single constraint, which is what a human being could be made to type into a form between calls. That constraint decided everything downstream of it. It decided what was worth recording, which turned out to be whatever was cheap enough to capture and conducive to the compliance of a sales team. And it decided the shape of the record, which is a small number of boxes each holding one value.

Automating the typing changes who does the chore. It doesn't change what the system knows.

Every version of the rebuild-the-CRM argument assumes the job is to make a better record for people, with AI helping. Almost nobody stops to ask the other question.

What would a record look like if the thing reading it most often wasn't a person?

That's a different question with a different answer, and once you've asked it the CRM argument stops being the interesting one.

Because the reason none of this works yet isn't that the models aren't good enough. I've spent two years watching frontier models do things that would have been science fiction when I was carrying a bag. Intelligence has stopped being the bottleneck. Context is the bottleneck now. And context has a bottleneck of its own.

05What Foundation Capital got right

The best published answer to that question came from Foundation Capital last December, in an essay about context graphs, and I want to give them credit properly before I share my opinion. They put this conversation in front of the market. They named the gap, which I'll quote here: "Not that the data is dirty or siloed, but that the reasoning connecting data to action was never treated as data in the first place."

Tomasz Tunguz got to the neighborhood a couple of days earlier, writing that "leaders have recognized their companies need a new system of record for AI agents in the form of a context database."

Here's the part where I break from it.

If you read the essay, you'll find it never says how anybody would know whether a captured piece of context is true. Not who said it, not what it rests on, not what happens when two pieces of captured context flatly contradict each other. The closest it comes is their line about decision traces that show "who approved what," and an approval is not evidence. It tells you somebody signed off. It tells you nothing about which claim wins.

They write that the context graph "becomes the real source of truth for autonomy," and a source of truth that can't tell you why to believe any single entry in it isn't a source of truth. It's a very well-organized pile of assertions.

So while my disagreement is narrow, I believe it's significant, and it's something I noticed by being hands on building the exact system and measuring the outputs with customers.

Capturing context was never the hard part. Making context you can check and validate is the hard part.

Everything else I have to say comes down to that one condition, and until it's met, I don't believe anyone hands a system like this a decision that matters.

06What a record for AI needs

Take a deal that gets marked closed lost.

The CRM handles that part fine. There's a picklist, somebody picks Price, the stage moves, and the audit log will tell you who clicked it and when.

What the record can't hold is what actually happened. That the price objection showed up on the third call and not the first. That it came from procurement, while the champion was still fighting for you internally and told you so. That the same objection landed two quarters ago at a company that looks exactly like this one, and we handled it differently there and won.

A field in a context graph carries the claim, where it came from, when, the sentence somebody actually said, and how it ranks against a competing claim about the same thing. Price, according to procurement, on the third call, quoted. Not price, according to the champion, in an email a week later. Both are stored. One outranks the other and the system knows which.

A closed lost deal shown two ways. The CRM record lists stage Closed Lost and loss reason Price, with a footnote reading one value, no source, no why. Beside it the context record, outlined in violet, resolves the same field to Price, per procurement, above two competing quoted claims: procurement scored 0.93 and marked as governing, and the champion scored 0.87 and marked outranked but retained.
The same deal, both ways. Two claims cleared the threshold. Standing decided which one governs, and the loser is retained rather than erased.Swipe the diagram to see all of it.

For a person that's a nicety, because you already supply the missing half yourself. You know without being told that what a customer said on a call outranks what a colleague assumed in a pipeline review. The best sellers understand this to be leverage.

The record never had to write that down. The person was the error correction, and the schema got to stay thin because somebody was always there to fill in whatever it left out.

Take the person out and there's nobody left to fill the gaps. A model handed a bare value (i.e. an API call to the CRM) has no way to decide whether to trust it. No memory of the meeting, no read on who's reliable, no basis for choosing between two facts that contradict each other.

Evidence is what a person wants. Rank is what a machine needs.

None of the primitives here are new, and I want to be clear about that. Versioning, lineage, provenance, and authority have existed for decades. What's new is combining them with models capable of reading the unstructured work a company produces and turning it into a living semantic layer.

An email no longer has to remain a single event called "Bob sent an email." It can become a set of queryable claims:

{
  "event": "email_sent",
  "claims": {
    "motion": "warm follow-up",
    "relationship": "existing customer",
    "recipient": "CPG leader",
    "follows": "meeting - Aug 12",
    "intent": "schedule next step"
  },
  "source": "email - Aug 14",
  "authority": "firsthand - the sender wrote it",
  "confidence": 0.91,
  "standing": "governs intent",
  "superseded_by": "newer firsthand evidence"
}

A deal moving from $100,000 to $200,000 can become evidence about which stakeholder drove the expansion, what requirement changed, which conversation preceded it, and why the opportunity became more valuable.

That is the shift.

The system doesn't merely store what happened. It develops structured, evidence-backed opinions about what happened and why it matters. Every opinion carries its type, source, time, confidence, and standing, so an agent can inspect it instead of blindly trusting it.

The customer stating her renewal date outranks a rep's guess. A signed order form outranks both. The graph preserves all three, but it knows which one should govern the answer and it can explain why.

One source event can therefore become five, ten, or fifteen useful data points without severing the line back to reality. That reasoning used to live temporarily in someone's head. Now it compounds inside the system.

We watched this distinction matter in a pilot with a customer's live data. The details are anonymized here. The mechanics are not.

One account carried two incompatible descriptions of the relationship. An existing synthesized brief called it a customer. The authoritative contract fields said it was still a prospect. Both claims were plausible, and both were useful, but they were not equally authoritative.

The system did not delete the losing claim. It let the contract record govern relationship status while preserving the brief's richer account goal, risks, and active motion. Newer direct evidence could supersede either. So the answer became more precise than either source alone: prospect for the purpose of commercial status, with a specific active motion and blocker drawn from the surrounding evidence.

A wider test surfaced the inverse problem. Twenty records looked like commercial changes, but their lineage showed they had been created by a dataset reconciliation rather than an actual change in the CRM. The language was plausible. The provenance was wrong. Once the system required a connector-owned before-and-after event, all twenty claims disappeared.

That is what standing changes. The model is no longer choosing whichever sentence sounds most convincing. It is operating over claims that carry source, time, authority, and the conditions under which another claim can supersede them. If somebody asks why, the system can return the evidence. If the excerpt loses important nuance, it can trace back to the complete underlying source.

Multiplying one event into fifteen claims is easy. Preserving the line from each claim back to its source, while deciding which one should govern, is the actual context problem.

07What I know now

In early 2024 I thought the problem was my territory. I then built two products working my way toward the thing that was actually true, and the only reason I got there is that I kept being wrong out loud in front of people who were willing to correct me.

Here's where I've landed.

Your stack stays. All of it (for the most part). The CRM, the engagement tools, the enrichment, the reporting nobody likes and everybody depends on. Nothing about this new technology requires you to pull out a system that has an entire company wired to it, and anybody telling you otherwise is asking you to carry a risk that pays them and not you.

What changes is that something new gets built on top of it. A layer that reads everything the company already produces, the calls and the emails and the meetings and the invoices and the tickets, and turns it into a record a machine can act on without a person standing next to it. That's a context graph, and I don't think there's much argument left about whether one gets built. The argument is about what has to be true inside it.

This changes the ninety-day equation too. A context graph does not require a CRM migration before it produces value. It can begin with one narrow workflow, backfill what the company already knows, and make that context available inside the applications and agents people already use. The CRM keeps doing the deterministic job the company depends on. The new layer starts improving decisions beside it.

And the thing that has to be true is provenance. A context graph without it is just a faster way to be confidently wrong at scale. Every claim has to carry where it came from, when it was made, how confident the system is, and what it outranks, or you've built a very expensive pile of assertions.

And this is playing out inside almost every large company right now, whether or not anybody names it that way. They've spent real money on AI. It demos well and then it doesn't do much once it's wired into the stack, and the conclusion they usually land on is that it's a people problem. We don't have the right talent. It's a culture thing, nobody's adopting it. So they run enablement, hire an AI lead, reorg a team around it.

The model is usually fine and so are the people. What it got wired into wasn't built for it. You get different answers on the same data. That isn't a culture problem. Their systems are disconnected, and that disconnection is a data infrastructure problem sitting right in front of them.

The last tech wave asked what your team remembered to log. This one is going to ask what your company actually knows about its customers, and whether anybody can prove it.

Keep your CRM. It earned the record of revenue. Give your AI a system of its own.