Get Your RevOps Health Score Book a Free Assessment
Lead Scoring

Lead Scoring Your Sales Team Will Actually Use

Most scoring models get built in a workshop, rank enthusiastic browsers above real buyers, and get ignored within two quarters. We calibrate yours against your own closed-won and closed-lost history instead.

Lead Scoring Challenges

Junk MQLs

Sales rejects most MQLs because scoring rewards activity instead of buying intent.

No Fit Scoring

Every lead is treated the same regardless of company size, industry, or budget.

Stale Models

The scoring model was built two years ago and hasn't been recalibrated since.

Lead Scoring Services

Score smarter, convert faster

Scoring Model Design

Build fit + intent scoring models that align marketing and sales on what makes a quality lead.

  • ICP fit scoring
  • Behavioral scoring
  • Negative scoring
  • Threshold calibration

Lead Grading

Layer demographic and firmographic grading on top of scoring for complete lead qualification.

  • Firmographic criteria
  • Technographic signals
  • Company size/revenue
  • Industry match

Lifecycle Stages

Define clear stages from subscriber to customer with automated transitions.

  • Stage definitions
  • SLA agreements
  • Handoff automation
  • Recycling rules

Model Optimization

Continuously refine scoring based on closed-won/lost data and sales feedback.

  • Conversion analysis
  • Score decay
  • A/B testing
  • Quarterly review

Why most lead scoring models get ignored

Almost every B2B SaaS company has a scoring model. Far fewer have one their reps read. The failure is rarely technical, because building the model in HubSpot or Salesforce is the easy part. The failure is that the points were decided in a room instead of derived from the deals you already won and lost.

  • The points came from a workshop. Marketing and sales negotiated the weights in a meeting. Nobody checked them against closed-won data afterwards, so the model encodes what the room believed in 2024.
  • Activity gets confused with intent. Reading six blog posts scores higher than visiting the pricing page twice. The model promotes researchers, students and competitors, and reps learn to distrust it within a quarter.
  • There is no negative scoring. Nothing subtracts for a personal email domain, a job title that will never hold budget, a company well below your ICP floor, or a support contact from an existing account. Points only ever go up.
  • Scores never decay. Someone who was interested last March still looks hot, because activity accumulates and nothing ages out.
  • The threshold was set to hit a volume target. The MQL bar got tuned until the number of MQLs matched the marketing goal, which makes the threshold a reporting decision rather than a quality one.
  • Nothing happens differently at the top of the range. A 95 and a 45 land in the same queue with the same follow-up time. If the score does not change routing or priority, it is a number nobody has any reason to act on.

The tell is simple. Ask a rep what score a lead needs before they call it. If they cannot answer, or they say they just work the list in order, the model is decoration and has been for a while.

How the build actually runs

About six weeks to build, then a quarter running alongside the existing model before you cut over. The shadow period is not padding. It is the only honest test of whether the new model ranks your real winners higher than the old one did.

  1. Mine the history, weeks 1 to 2. We pull your closed-won and closed-lost records and work out which attributes and behaviors actually preceded a win, at what rate, and which ones look predictive but are not. This step regularly kills two or three weights the business was confident about, and finds one nobody had scored at all.
  2. Fit and intent as two axes, weeks 2 to 3. Firmographic fit and behavioral intent stay separate, because the two failure modes are opposite and a single blended number hides which one you are looking at. Negative scoring and decay rules go in here.
  3. Thresholds and handoff, weeks 3 to 4. The bar gets set against a quality target, then routing, follow-up SLA and recycling rules get wired so the score changes what happens to the lead. This is the step most often skipped, and it is the one that decides whether anyone uses the model.
  4. Backtest, week 4. We run the new model against the last two quarters of closed deals and check that it would have ranked the winners above the losers. If it would not, it goes back to step one rather than into production.
  5. Shadow run, weeks 5 to 6 and the following quarter. Both models score in parallel. Reps keep working the old one while we compare acceptance and win rates, then you cut over on evidence instead of on launch day optimism.
  6. Recalibration schedule. A quarterly review lands in someone's calendar with the query already written, because a model is a snapshot of last year's buyers and decays quietly as ICP, pricing and route to market change.

The one scoring engagement we can describe publicly is written up in our AI lead scoring case study: a growth-stage SaaS team with 10,000 leads a month and no way to rank them, where win rate moved from 5 percent to 13 percent over the following six months. Our production scoring models run at roughly 82 percent accuracy. We anonymize clients to industry level, so that is the level of detail we can give.

Where AI lead scoring fits

Predictive and AI lead scoring is worth doing. It is worth doing second. A model trained on your CRM learns whatever your stage and lifecycle definitions actually encode, so if qualified means whatever the rep felt that morning, the model learns rep behavior rather than buyer intent and hands it back with a probability attached. That is harder to argue with than a rules-based model and no more correct.

Three questions worth answering before you buy it:

  • Do you have the volume? Predictive scoring needs a real history of closed-won and closed-lost outcomes. A few hundred deals a year with inconsistent property data will not produce a model worth trusting, whatever the interface implies.
  • Is the underlying data consistent? Lifecycle stage, lead source and deal stage need to have meant the same thing for the period the model trains on. If stages were renamed last year, the training window is shorter than you think.
  • Can you explain a score to a rep? A rules-based model can be argued with, which is how it gets corrected. A black-box probability that reps cannot interrogate gets ignored in exactly the same way the old model did.

In practice we usually ship a calibrated rules-based model first, because it is explainable and it forces the definitions to get fixed. Predictive scoring goes on top once there is clean history to train against, and the rules model stays as the sanity check.

Worth checking the licence before you plan around a native predictive model. HubSpot's documentation on predictive lead scoring lists it under Marketing Hub Enterprise and Sales Hub Enterprise, and describes it as analyzing your existing customer data to estimate probability of closing. Note that the page does not publish a minimum data threshold, so whether you have enough history for a reliable model is a judgement call rather than something the tool will tell you.

What a lead scoring engagement costs

Published 2026 market benchmarks put project-based RevOps work at 10,000 to 150,000 dollars or more and a focused revenue operations audit at 2,500 to 10,000. A scoring build normally sits at the lower end of the project range, because the scope is one model rather than a system rebuild.

What moves the number is the state of your historical data, not the sophistication of the model. If closed-lost reasons are populated and lifecycle stages have been stable for two years, calibration is quick. If half your closed-lost records say "other" and lead source is free text, the first job is reconstruction, and that is where the hours go. It is the same reason two quotes for the same brief can differ by a factor of three.

Ranges as published in MergeYourData's 2026 RevOps pricing benchmarks. Market-wide figures, not our rate card. More detail in our guide to what RevOps consulting costs.

What you own when we leave

  • The scoring model itself, built inside your own HubSpot or Salesforce instance, editable by your team.
  • The analysis that produced the weights, so you can see why each attribute scores what it does and argue with it.
  • The recalibration query, written and documented, so the quarterly review is a task rather than a project.
  • The written definitions for each lifecycle stage and the MQL threshold, which is the document that stops the bar drifting the next time volume targets get tight.
  • The routing and SLA rules, in your instance, under your admin ownership.

No proprietary scoring engine you have to keep paying for, and no arrangement where recalibrating your own model requires a call with us.

Related reading: why your CRM data is costing you revenue covers the data-quality problem that sits underneath most scoring models that stopped working.

Frequently Asked Questions

What is the difference between lead scoring and lead grading?

Scoring measures behavioral intent, meaning what someone does. Grading measures firmographic fit, meaning who they are. The distinction matters because the two failure modes are opposite: a high score with a low grade is an enthusiastic browser who will never buy, and a high grade with a low score is a perfect-fit account that has not engaged yet. Collapsing both into one number hides which of those you are looking at, which is why we keep them as two axes.

Why does sales ignore our lead scores?

Usually because the model rewards activity rather than buying intent, so it promotes people who read a lot of blog posts. Common causes are points assigned in a workshop rather than derived from closed-won data, no negative scoring so students and competitors accumulate points, a threshold set to hit an MQL volume target instead of a quality bar, and no decay, so a lead who was interested a year ago still looks hot. Once reps have been burned a few times they stop reading the score, and after that the model is decoration.

How do you build a lead scoring model?

We start from your closed-won and closed-lost history rather than a workshop. Which attributes and behaviors actually preceded a win, at what rate, and which ones look predictive but are not. Then we build fit and intent as separate axes, add negative scoring and decay, set thresholds against a quality bar rather than a volume target, and wire routing and the follow-up SLA so the score changes what happens to the lead. A score that does not change routing is a number nobody acts on.

How long does a lead scoring build take?

Around six weeks to build, then a quarter of running it alongside the old model before you cut over. Scoring is one of the few RevOps projects where the shadow period is not optional, because the only honest test is whether the new model would have ranked your last two quarters of closed-won deals higher than the old one did.

Do we need AI lead scoring?

Only after the fundamentals are in place, and only if you have the volume. A predictive model needs a meaningful history of closed-won and closed-lost outcomes with consistent property data behind them, and it learns whatever your stage and lifecycle definitions actually encode. If qualified means whatever the rep felt that morning, the model learns rep behavior rather than buyer intent and returns the same bias with a probability attached. Platform-native predictive scoring also tends to sit on the top subscription tier, so there is a licence cost to weigh as well.

How often should we recalibrate the model?

Quarterly at minimum, and immediately whenever MQL acceptance rates drop, win rates diverge from what the score predicted, or you change ICP, pricing or route to market. A scoring model is a snapshot of which buyers converted under last year's conditions, so it decays quietly as the business changes.

How much does a lead scoring engagement cost?

Published 2026 market benchmarks put project-based RevOps work at 10,000 to 150,000 US dollars or more and a focused revenue operations audit at 2,500 to 10,000, with a scoring build normally sitting at the lower end of the project range. The cost driver is the state of your historical data rather than the complexity of the model, because a model is only as good as the closed-won and closed-lost records it is calibrated against.

Ready to Score Leads That Convert?

Build My Scoring Model