Most chat models are tuned on human preference: two answers go to a rater, the rater picks the one they like better, and the model learns to write more like the winner. That works well for explaining how a lease works, but it falls apart when the job is selling one, because nobody reading a transcript can tell which reply would have brought the buyer in. More often than you’d expect, the reply that reads best is the one that lost the deal.

So we grade on something a rater can’t fake, which is what the buyer actually did next.

What a labelled conversation looks like

RevDesk agents reach a new lead within seconds of a web form, follow up across channels, and hand the conversation to a salesperson when it’s time. Whenever something comes of a conversation, the outcome is recorded against the call or message that produced it:

  • an appointment set
  • an appointment kept
  • a qualified lead
  • a unit sold, with the deal value when the dealer’s system has it

Each outcome also records who got it there, whether that was the agent on its own, the agent with a salesperson’s help, or a salesperson alone. Nobody sits and rates the replies, because the label is simply what the customer did.

Here’s how a thread and its label fit together:

Website lead, 9:41pm: Is the gray Tacoma still there? What would you give me for my 4Runner?

Agent, iMessage, 9:41pm: It’s here. Trade numbers need eyes on the truck, and our used car manager can do that in about ten minutes. Want to bring the 4Runner by Saturday at 10?

Lead, 9:44pm: make it 11

Replied in 3 min · Appointment set · Appointment kept · Sold on day 6 · Agent with salesperson

Mike sold the Tacoma on the lot six days later, so the thread gets credit for its part, and the label records that it had help getting there.

Waiting for the ending

Outcomes in our customers’ businesses tend to arrive late. A buyer who texts on a Tuesday might come in on Saturday and sign the following week, so if you label a conversation the moment it goes quiet, you’ll label most of them wrong.

That’s why a conversation can’t be used until it has been over for 14 days, and an outcome only counts toward it if it lands within 30 days of the conversation starting. A written thread is treated as finished after seven days of silence, which means a reply that shows up three weeks later starts a new conversation with its own ending.

14 daysA conversation must be over before it counts
30 daysFor an appointment or a sale to land
OffContribution, until an admin turns it on

How it’s graded

We grade the model the way our customers grade a salesperson, by what the conversations earned, which for a dealership means appointments that show and units that sell. A new version will have to beat the current one on that measure before it replaces it. The hope is that over enough conversations, the model stops merely sounding like a salesperson and starts getting a salesperson’s results.

It knows the person, inside one business

The other half of the data is context about the person. Before an agent writes anything, it knows whether this is a first inquiry or a returning customer, which channel they tend to answer on, and what they asked about last time, and the right message changes with all of it. A service customer who has been in four times doesn’t need a discount to book an oil change, while a first-time buyer comparing two stores on a Sunday night might need a real reason to pick yours.

That context stays inside the business it came from. The agent keeps it on the contact record and reads from the business’s own knowledge base and inventory, so it remembers without training on anything, and nothing one dealership’s agent knows about its customers ever reaches another dealership.

What the global model can learn from

The global model is built on foundation models, licensed and public material, and conversations from businesses that choose to contribute them. Before a contributed conversation is stored, a few things happen:

  • Names, phone numbers, emails, addresses, VINs, card numbers, and other identifiers are swapped for typed placeholders, so the model learns that a phone number was given without ever seeing which one. Both rules and a model do the redacting, the rules then run again on the result, and any conversation that still matches is dropped.
  • Anyone who has opted out on any channel, or who is on a do-not-call list, is left out entirely, along with minors and test conversations.
  • At launch, text messages are left out too, as are businesses outside the US, workspaces in healthcare, legal, and lending, and any workspace with HIPAA turned on.

If a business turns contribution off, any records that haven’t been used in training yet are deleted within 30 days, and records that have already been used are kept for up to 24 months for retraining and evaluation.

Where it is

Our agents are in production at dealerships and service businesses today, and every conversation they hold is recorded along with how it ended. For now, that record reaches each agent through its memory and through managers, who read transcripts and outcomes and adjust knowledge, playbooks, and prompts themselves. Agents don’t yet change their own behavior based on outcomes, and that’s the loop we’re building.

Contribution opened on September 30, 2026, and the first contributed conversations will become usable once their 14 days have passed.

What we haven’t figured out

Quite a lot. We don’t yet have a settled way to credit a thread for a sale that a salesperson closed on the lot, or to score a conversation that set an appointment the customer later regretted. The biggest open question is how to keep a model that’s rewarded for sales from spending the trust that brings the customer back for service. Most of our time goes into questions like these, and we’ll write about them here as we work through them.

If you’d like to help build this loop, we’re hiring in New York →