Most conversations about AI in revenue operations start with the tool and work backwards. I find it more useful to start with the work. RevOps teams spend a large share of their week on tasks that are repetitive, rule-bound and easy to check: cleaning records, filling missing fields, routing leads, writing up calls, explaining why the forecast moved. Those are good candidates for AI. They also spend time on judgement calls that carry real cost when wrong: which accounts to prioritise, what number to commit to the board, whether to disqualify a deal. Those are poor candidates for automation, though AI can still help prepare the decision. I am not going to quote vendor statistics on time saved, because most of them are self-reported and untestable. Instead, here is how I sort the use cases, with a simple rule for each: what the model does, what a person checks, and what never leaves human hands.
Data hygiene and enrichment
This is the safest place to start, because the output is easy to verify and the cost of a wrong suggestion is low when a person approves it. A model can flag likely duplicates that exact-match rules miss, such as the same company entered as an abbreviation and a full legal name. It can standardise free-text fields, like job titles into a seniority and function picklist, or messy country and state values into a clean list. It can read an email signature or a website and propose missing fields such as industry, employee range or location. And it can spot records that break your own data dictionary, for example closed-won deals with no amount or a lead source of "other". What I keep human is the merge itself and any change to fields that drive commission, territory or reporting. The pattern is suggest, review, apply: the model writes its proposal to a staging field or a review queue, a person accepts in bulk, and only then does the record change. That keeps an audit trail and avoids the worst failure, a confident model quietly rewriting thousands of records with plausible but wrong values.
Lead scoring and qualification
AI is useful here in two specific ways. It can read unstructured signals that rule-based scoring ignores: the free-text "how can we help" field, the pages someone visited, the content of a reply email. And it can summarise why a lead scored the way it did, which makes sales more likely to trust the score. A practical setup is a rule-based fit score built from firmographics, plus a model-generated intent note that a rep sees next to it. Where I hold back is letting the model set the final score or disqualify leads on its own, at least until you have tested it against closed revenue. A model trained on your historical wins inherits your historical biases, such as under-scoring segments you never sold to properly. Start by running the model's score in parallel with your current one for a quarter, compare both against what actually converted, and only then let it influence routing. I have written separately about building predictive scoring without a data science team, which covers the testing in more detail.
Routing and speed to lead
Routing is mostly a rules problem, and rules should stay in charge of it: territory, segment, round robin, ownership of existing accounts. AI earns its place at the edges. It can classify inbound requests that do not fit a form neatly, such as a partnership enquiry, a support issue sent to sales, or a job applicant who filled in the demo form, and send them to the right queue instead of a rep's inbox. It can match a new lead to an existing account when the email domain differs from the company domain. And it can draft the first reply for a rep to edit and send, which shortens response time without putting unreviewed text in front of a buyer. What stays human is the routing logic itself and any exception that moves a lead between reps, because those decisions affect pay and team trust. If you are unsure how much delay is costing you, model it before buying anything. Slow routing is often a process problem that no model will fix.
Call notes, CRM updates and deal hygiene
Call summarisation is the use case with the clearest everyday value. A model can turn a recorded call into a summary, next steps, stakeholders mentioned, and proposed updates to fields such as next step date, competitor and close date. Reps hate writing notes, so this improves CRM completeness in a way that years of reminders often do not. Two rules make it work. First, proposed field updates go to the rep for one-click approval, they are not written straight into the opportunity, because a model will occasionally hear "we might look at this in Q3" as a committed close date. Second, recording and transcription need consent and a policy that matches where your customers are. Rules on recording calls and processing personal data differ between India, US states and other markets, so get that settled with whoever owns privacy before you switch it on. Deal hygiene checks follow naturally: a weekly pass that flags opportunities with no activity, close dates in the past or stages that do not match the notes, sent to the owning rep, not fixed silently.
Forecasting: explain the number, do not set it
Forecasting is where I see the most overreach. A model can be very helpful in preparing a forecast call. It can summarise what changed in the pipeline since last week, which deals moved stage or date, which large deals have gone quiet, and where rep commits diverge from historical conversion by stage. It can write the first draft of the variance commentary a CFO wants to read. What it should not do is produce the committed number on its own. The forecast is a commitment made by people who will be held to it, and it relies on context a model rarely has, such as a buyer's budget freeze mentioned in a hallway conversation. A useful compromise is to show a model-generated or statistical forecast as one input alongside the rep roll-up and the manager call, and to track all three against actuals over time. If the model is consistently closer, that is evidence to give it more weight. Until then, it is a second opinion.
A simple rule for what to automate
When a new AI use case comes up, I ask four questions. Is the output checkable by a person in seconds? Is the cost of a wrong answer small and reversible? Does it run often enough that saving a few minutes each time adds up? And is there a clean place to put the output, such as a review queue or staging field, rather than straight into a system of record? Four yeses means automate with review. Two or fewer means keep it human and let AI prepare the material. Anything that commits money, sends a message to a customer without review, changes compensation-relevant data or alters the forecast stays with a person regardless. Finally, measure each use case against a baseline you recorded before switching it on, such as field completeness, time to first response or hours spent on forecast prep. Without a baseline, every AI project looks like a success and none of them can be defended at budget time.
Sources
This article contains no third-party statistics. For platform capabilities mentioned in passing: Salesforce Hosted MCP Servers (per-user authentication for agent actions): https://developer.salesforce.com/docs/platform/hosted-mcp-servers/guide . HubSpot MCP servers: https://developers.hubspot.com/mcp
FAQ
CRM data hygiene: duplicate detection, field standardisation and enrichment suggestions. The output is easy to check, mistakes are cheap when a person approves changes in bulk, and cleaner data improves every report and score built on top of it. Write suggestions to a review queue rather than directly into records.
For low-risk fields with high confidence, it can, once tested. For anything that affects reporting, territory, commission or the forecast, have the model propose the change and a person approve it. Proposed updates from call summaries in particular should go to the rep for one-click approval.
It should not set the committed number. Use it to explain pipeline movement, flag risky deals and draft variance commentary, and show a model forecast alongside the rep roll-up. Track both against actuals, and give the model more weight only if it proves more accurate over several periods.