gtmjosh

How to build a lead scoring model Sales will trust

· 19 min read· Salesforce · HubSpot· Platform behavior verified August 24, 2026

Build a lead scorecard from fit and intent, define MQL vs. SQL acceptance, and implement Salesforce or HubSpot lead scoring with explainable AI advice.

On this page

A lead score has one job: help a team decide which people deserve attention. The math is the easy part. The hard part is producing a number Sales understands, trusts, and gets excited to see.

I keep fit and intent separate behind the scenes, then give Sales one combined score. The components let RevOps diagnose the model. The single number gives the team a shared operating signal. A theoretically better matrix is a poor system when the people using it cannot explain what its boxes mean.

This page shows how I build that system, from ICP evidence and enrichment through decay, qualification, stakeholder review, and ongoing calibration.

The model was working. The business was not.

I once helped run a heavily intent-weighted system that influenced account scoring and the target-account list. On paper, the engine looked better every month.

Marketing would swear up and down that engagement had never been stronger. More of the people responding came from target accounts, campaign KPIs were climbing, and the system was passing more leads to Sales. Sales could see that the people were engaged, but struggled to explain why so many of them still were not getting it. Executives looked at an engine that appeared healthier than ever and could not reconcile that story with deals stalling or being closed out.

Every team had data supporting its version of events. Nobody had a satisfying explanation for the whole thing.

The scoring model was not broken in the obvious sense. It was finding enthusiasm. The people reaching Sales really were more engaged and often more excited about the product. The system was doing exactly what we had weighted it to do.

It was answering the wrong question too confidently.

The companies were still wrong too often. Even when company revenue matched the target-account definition, industry sometimes did not. Seniority looked excellent while the department had no reason to buy. When we finally added person-level persona matching, we discovered that only about one in eight records in the database represented somebody likely to understand the product or want to buy it.

Intent had turned a targeting problem into a larger, faster targeting problem. It rewarded genuine engagement without asking whether the engagement came from a viable buyer. Marketing saw demand, Sales felt the mismatch, and leadership saw the gap between leading indicators and commercial outcomes.

That failure changed how I build scoring systems. The lesson was not to abandon intent. Intent is often the right place to begin and may deserve the stronger initial weight. But enthusiasm cannot substitute for account fit, persona fit, or a shared definition of a buyer Sales can actually help.

A scoring system exists to make that definition operational. It should expose where the evidence came from, preserve the disagreement when Sales overrides it, and give the whole company one signal whose meaning survives the handoff between dashboards and conversations. Everything that follows is in service of that.

Start with the decision, not the points

Lead scoring usually supports two decisions:

  • Qualification: should Marketing and Sales actively work this person?
  • Prioritization: among the people who qualify, who should receive attention first?

Routing, nurture, and SLA enforcement use the result. They are separate policies. That separation lets you change who works a lead without changing what the score means.

For example, a company might send its highest scores directly to an account executive, send the next band to an SDR for additional qualification, and keep a lower band in an automated marketing program. The score did not make three decisions. It supplied evidence to three workflows.

Write the qualification decision before assigning a point:

lead-scoring-decision.yaml
decision: Is this person worth active Sales or Marketing attention now?
owner: Revenue Operations
components:
  - fit
  - intent
output: combined_score
allowed_outcomes:
  - qualified
  - needs_more_qualification
  - disqualified
downstream_uses:
  - prioritization
  - sales_acceptance_sla
  - tiered_routing
  - nurture_eligibility

The Decision Card template is the fuller version of this contract. Use it when the score can trigger expensive or customer-facing work.

Build fit and intent separately

Fit and intent answer different questions.

Fit asks whether you can sell successfully to this company and person. Company size, industry, geography, business model, current technology, department, role, and persona can all belong here.

Intent asks whether this person is showing evidence of a buying motion now. A demo request carries more intent than an ebook download. Repeated useful visits can matter more than one shallow click. A meeting or direct conversation should outweigh an email open.

Keep the component scores visible to RevOps even when Sales only sees the combined number:

ComponentExample evidenceWhat it must not claim
Account fitRevenue band, industry, geography, target-account statusThat the person is a buyer
Persona fitDepartment, role, responsibility, likely influenceThat seniority alone creates relevance
IntentHandraiser event, direct request, repeated engagement, meetingThat enthusiasm fixes poor fit
Combined scoreAgreed weighting of fit and intentThat one number explains every component

Persona deserves its own line. A senior employee at a large target company can still be the wrong person. A senior creative leader at a global software company may have budget and authority, but neither helps if the product solves a technical search problem owned by another department.

This is where an intent-heavy model usually breaks. It finds people who are excited about the content and quietly assumes they are excited about buying the product.

Give Sales one number with a stable meaning

Matrices are useful during design. High fit and high intent should behave differently from high fit and low intent, and the matrix makes that obvious to the builder.

Most sales teams struggle to use the matrix consistently in live work. I collapse the components into one weighted score once the backend behavior is sound.

Sales does not need another dashboard that requires a legend and a training session before somebody can decide whom to call. RevOps needs the diagnostic detail. Sales needs the conclusion.

A first version might look like this:

Example combined score
combined score = (fit score × 0.40) + (intent score × 0.60)

That is an example, not a universal ratio. I often start with more weight on intent. Then I watch what Sales disqualifies. If high-intent records keep failing because the companies or personas are wrong, fit needs more influence or a hard gate needs to run earlier.

The number needs a sentence Sales can repeat. Something like:

A score above 75 means the person has shown current buying intent, the company fits the market we serve, and the role is likely to understand or influence the purchase.

If a rep cannot say why a high score is worth pursuing, the model is unfinished.

Put hard gates before weighting

A weighted score should not average away a condition the business considers non-negotiable.

A demo request is evidence of intent. It is not a hall pass around the ICP.

Suppose the company targets enterprises above $500 million in annual revenue. The team might decide that companies below $200 million are outside the sales motion. That belongs in an eligibility rule, not as a small negative weight a strong demo request can overcome.

Hard gates vary by company. Common candidates include:

  • a known spam or vendor solicitation;
  • a competitor or prohibited geography;
  • a company clearly below an agreed minimum size;
  • a product or industry exclusion the business has validated;
  • a role with no plausible connection to the problem, when persona is reliable;
  • a personal email address in an enterprise motion that requires a verifiable company.

Personal email is a business-model choice. I disqualify it in mid-market and enterprise motions where company identity is required. I accept it on GTM Josh because education is the primary goal and readers change employers. A long-term reader on a personal address can be more valuable here than one download from a famous company domain.

When a business accepts personal email, increase the qualification burden. Ask for more useful information on the form, use an AI-assisted email before assigning manual work, or hold the record for review. Weak evidence on paper needs more qualification, not a confident guess.

Treat missing data as unknown

A blank revenue field does not mean the company has no revenue. Missing industry does not mean poor industry fit. Turning either into zero builds a data-quality problem into the score.

Unknown is not a polite word for bad. It is an instruction to collect better evidence before pretending the model knows the answer.

I use tiered enrichment. High-intent handraisers, including demo and contact requests, always justify enrichment. Lower-intent records can wait for a scheduled process or a moment when somebody is about to use the field.

The sequence is:

  1. Resolve the company from the business email and submitted identity.
  2. Use known company information where it is reliable.
  3. Call a waterfall provider such as Clay or Apollo when the decision needs structured values or corroboration.
  4. Store the source, observed time, and overwrite policy with each enriched field.
  5. Leave the component unknown when the evidence still does not support a value.

AI can help identify a likely company, industry, and scale from a name and business email. Treat that as inferred evidence until it is corroborated. The evidence and confidence Reference shows how to carry that distinction into a decision output.

For cost and overwrite controls, use moment-of-use enrichment and the enrichment approval-queue pattern.

Make intent decay like the behavior does

Intent gets stale. The decay rule should follow the buying cycle and the strength of the activity instead of copying a generic schedule.

High-intent actions start with more points, so they naturally have more room to decay. Repeated smaller engagements can also build a meaningful signal. One page view should not survive for months. A demo request followed by a meeting should not disappear after a week.

This is a reasonable starting policy for a longer B2B cycle:

Time since meaningful engagementExample treatment
0 to 30 daysKeep high-intent records in active qualification and follow-through
31 to 90 daysDecay activity points and use appropriate nurture or review
91 to 180 daysRequire new evidence before restoring active priority
More than 180 daysReview for disqualification under the company retention and lifecycle policy

I commonly use three months from the last engagement as the main decay horizon and review records with no activity for six months for disqualification. I also run disqualification and deletion audits quarterly. Those dates are operating choices, not legal retention advice. Your buying cycle, consent rules, and data-retention policy decide the final schedule.

Record the activity type and date that contributed each term. A score that decays without exposing its inputs becomes impossible to defend.

Translate the score into operating rules

Once the score has a stable meaning, decide what each band permits.

Score stateQualificationExample downstream treatment
HighQualifiedAE follow-up, shortest SLA, highest work priority
MediumNeeds more qualificationSDR review, enrichment, or AI-assisted first touch
Low but eligibleNot sales-readyMarketing nurture and continued intent collection
Hard disqualifiedDisqualifiedSuppress Sales and Marketing under the documented policy
UnknownIncomplete evidenceEnrich, review, or wait without treating the blank as failure

In organizations where Disqualified means neither Marketing nor Sales will contact the person, use that word carefully. A record that is merely not sales-ready belongs in a different state. The Lead and Contact lifecycle keeps qualification, consent, lifecycle, and current work from collapsing into one field.

An SLA can use a minimum score as its acceptance boundary. Document the start event, the accountable team, and the missed-SLA action separately. The score determines eligibility. The SLA contract determines what timely follow-through means.

Define marketing qualified lead (MQL) versus sales qualified lead (SQL)

An MQL, or marketing qualified lead, should mean that Marketing has applied the agreed qualification contract. An SQL, or sales qualified lead, should mean that Sales has reviewed and accepted the person for active selling. A score can make a record eligible for the handoff; it cannot prove that Sales accepted it.

Record the boundary explicitly:

EventRequired evidenceHistory to preserve
MQL createdScore or hard-gate result meets the published qualification ruleScore, version, component values, time, and qualifying event
Sent to SalesEligible owner and routing policy select the work queueDestination, send time, and routing reason
Sales accepted / SQLSales confirms the record is worth active sellingAcceptor, acceptance time, and current evidence
Sales rejectedSales chooses a governed reasonOriginal score, rejection reason, time, and owner
Returned or recycledThe return policy names the next eligible stateReturn reason, next review date, and prior history

Do not infer SQL from an owner change or task creation. Those events show that a workflow ran, not that Sales accepted the lead. The lifecycle and handoff contract keeps the qualification state, transfer, acceptance, rejection, and recycle path separately reportable.

Copyable lead scorecard template

A useful lead scorecard is the versioned operating contract behind the number, not only a list of positive and negative points. Copy this structure and replace the examples with evidence from your own ICP, personas, buying motion, and historical outcomes.

SignalComponentExample ruleWeightCap or decayMissing-data behaviorOwner
Target industryAccount fitConfirmed industry is in the approved ICP set+15No decayUnknown until enrichedRevOps
Company sizeAccount fitVerified revenue or employee band is eligible+15No decayUnknown; do not award or subtractRevOps
Buying personaPersona fitRole and department match an approved persona+20Re-evaluate after job changeUnknown until matchedMarketing + Sales
Demo requestIntentValid handraiser event is recorded+30Decay after the agreed buying windowReject duplicate or test eventsMarketing Ops
Repeated engagementIntentMultiple qualifying events inside the review window+10Cap repetitions; age by event typeNo event means zero intent, not bad fitMarketing Ops
Personal emailHard gate or reviewApply the company's market-specific policyRe-evaluate if company is identifiedDo not infer company without evidenceRevOps

Then define the score outputs separately:

lead-scorecard-template.yaml
model:
  name: Lead qualification score
  owner: Revenue Operations
  version: 1.0
  components:
    fit_weight: 40
    intent_weight: 60
  hard_gates: []
  bands:
    - name: qualified
      minimum_score: 80
      permitted_action: sales_handoff
    - name: needs_more_qualification
      minimum_score: 55
      permitted_action: sdr_review_or_enrichment
    - name: eligible_not_ready
      minimum_score: 0
      permitted_action: marketing_nurture
  acceptance:
    mql_event: marketing_qualification_completed
    sql_event: sales_acceptance_recorded
    rejection_reason_required: true
  monitoring:
    - sales_acceptance_rate
    - disqualification_rate
    - opportunity_progression
    - override_rate
    - score_distribution

The numbers are placeholders. Simulate them against known records, review the result with stakeholders, and record why each production weight differs from the template.

Calibrate before launch

I start with judgment and historical records. Neither is sufficient alone.

First, inspect won opportunities and deals that reached meaningful stages. Build the initial ICP and persona profile from the attributes those records share. Include lost and disqualified records so the model does not learn only from success.

Then simulate scores at the account and person level. Put the records in front of the people who own the outcome. Ask which high scores excite them, which ones look wrong, and what evidence changed their mind.

The review is part of the build. Stakeholder buy-in is more important than squeezing another point of theoretical accuracy from a formula nobody trusts.

The room test matters: when the team sees the highest-scoring records, the reaction should be recognition rather than surprise. When there is surprise, preserve it. That disagreement is usually where the next useful rule comes from.

A practical launch sequence:

  1. Score historical records without changing workflow behavior.
  2. Review the top and bottom bands with Sales, Marketing, and the relevant customer or success team.
  3. Compare the distribution with accepted, disqualified, progressed, and won records.
  4. Run the score in shadow mode on new records.
  5. Publish the meaning, thresholds, component inputs, and owner.
  6. Turn on one downstream action at a time.

The golden-question testing pattern works well for the edge-case set. Include the high-intent wrong persona, the excellent-fit record with no current intent, the personal-email handraiser, the missing-company record, and the repeated low-value engager.

Use AI lead scoring as an advisory layer

AI is useful around a scoring model when it behaves like an independent evaluator, not an invisible replacement for the methodology.

For a lead score, deterministic fit and intent rules can remain authoritative while AI reads unstructured evidence, checks whether the component values make sense, and points out what the model could not see. For a seller-owned methodology such as MEDDICC, the rep enters the official component scores and AI returns its own advisory assessment. The same pattern applies to account health when a named customer owner remains accountable for the final view.

The advisory evaluation should run even when a person has not completed their score. Leadership still benefits from seeing what the available evidence supports, and the missing human assessment is itself visible. Once both exist, compare them without letting either overwrite the other:

Store separatelyPurpose
Human or deterministic component scoreThe accountable operational assessment
AI suggested component scoreAn independent reading of available evidence
AI confidence by componentHow strongly the accessible evidence supports the suggestion
AI rationale and cited evidenceThe claim a person can challenge
Missing or unavailable evidenceWhy the suggestion may be incomplete
Human contextEvidence outside transcripts, fields, or captured activity
Score and rationale historyWhat each side believed at the time

For a seller-owned assessment such as MEDDICC, rep evidence should win operationally because the rep owns the deal. If the business cannot trust the rep, that is a management problem, not a reason to let a model quietly take authority. An egregious conflict can still be useful: it gives a manager a compass for the next question instead of declaring an automatic winner. In lead scoring, apply the same separation to the deterministic operational score and the AI advisory score.

For example, AI may suggest a lower Economic Buyer score because the available calls show influence but no purchase authority. A leader can then ask the seller to explain why the person is an economic buyer rather than a champion. The rep may have evidence the system never received. They may also realize the score was optimistic. Both are productive outcomes because the rep still has to defend the official score.

Confidence is not authority. Store an overall confidence and component-level confidence beside the AI suggestion. Confidence should generally improve as a deal produces more usable evidence, but a highly confident disagreement still starts a review; it does not grant the model permission to change the score.

Track the delta between the operational and AI scores, then compare low and high scores with opportunity progression. Persistent disagreement can indicate coaching, a missing context rule, incomplete human input, or an ingestion failure. The delta is not only a score-quality metric. It is also an observability signal for the evidence pipeline feeding the evaluator.

The rules-versus-AI architecture defines the authority boundary. The decision explanation and review pattern shows how to preserve the comparison without creating a black box.

Implement HubSpot lead scoring without changing the fundamentals

HubSpot's current lead-scoring tool matches this architecture closely. It supports separate fit and engagement scores plus a combined score. A combined score creates three properties: total, fit, and engagement. The tool also supports score decay, inclusion and exclusion segments, record testing, distribution previews, thresholds, and score history.

Use the combined property for the simple Sales-facing number. Keep the fit and engagement properties on the RevOps view and in calibration reports. HubSpot also produces a fit-by-engagement threshold category such as A1 or C1. That matrix is useful for backend analysis even when Sales works from the combined value.

Before turning the score on:

  1. Choose the Contacts, Companies, or Deals eligible to be scored.
  2. Group property criteria around fit and event criteria around intent.
  3. Cap repeated low-value event groups so they cannot overwhelm the model.
  4. Apply decay to the events whose meaning fades.
  5. Test known records and preview the distribution.
  6. Inspect every workflow, segment, view, and report that uses the score properties.
  7. Turn on downstream actions separately from the score calculation.

HubSpot documents the current behavior in its lead-scoring builder and lead-scoring tool overview. Feature availability depends on subscription.

Implement Salesforce lead scoring without changing the fundamentals

In Salesforce, the same model can live in custom Lead, Contact, and Account fields maintained by Flow or Apex, or use an applicable Einstein scoring feature. Keep the component fields, combined score, qualification result, version, and last-calculated time distinct. A packaged prediction still needs your eligibility, acceptance, override, and downstream workflow contracts.

Salesforce's current Sales Cloud Einstein implementation guide documents native Einstein Lead Scoring setup and considerations. Validate the feature, license, population, and available explanation against your org before making it the source of an operational threshold.

RevOps owns the model

RevOps should own the scoring methodology. Ownership includes getting Marketing, Sales, and Customer Success agreement where the use case touches them.

The initial framework needs stakeholder approval. Routine tuning can belong to RevOps once the principles and decision rights are clear. Every material change still needs an outward explanation: what changed, why the evidence supported it, which records will move, and what the team should expect.

Sales can override qualification. Preserve the original score, final decision, override reason, person, and time. The disagreement is calibration data. Erasing the score destroys the evidence needed to improve it.

RevOps should also meet with the teams using the model. Reports reveal patterns; interviews and surveys reveal why the pattern exists. You need both.

Measure whether the model helps the business

The KPI defines success. Common measures include:

  • Sales acceptance and disqualification rate by score band;
  • opportunity creation and progression;
  • time from qualification to acceptance and meaningful stage movement;
  • conversion, pipeline, and won opportunities;
  • override rate and reason;
  • score distribution by industry, persona, source, and campaign;
  • non-persona contacts who successfully introduce the right buyer.

That last measure catches an easy mistake. A person can be the wrong persona and still have enough influence to reach the right one. The action may be a referral request instead of a normal sales sequence.

Repeated engagement from the wrong ICP also says something about Marketing. If the score keeps finding enthusiastic people Sales cannot serve, targeting or content positioning may need to change. Do not tune the score until it hides a demand-quality problem upstream.

Across implementations at multiple organizations, I have seen tighter persona and fit calibration move Sales acceptance from roughly 30% into the 70% range. Opportunity lifecycle time fell by about 50% to 90% in some segments and products. Those are observed ranges, not a benchmark or promised result. The scoring changes happened alongside tighter marketing targeting and better persona qualification, so the model should not receive all the credit.

The goal is not to manufacture more high scores. It is to make a high score mean something the business recognizes in the records, the conversations, and eventually the revenue. Build the components separately. Give Sales one score. Then keep proving what that score means.

FAQ

Should fit and intent be separate lead scores?
Keep fit and intent separate in the scoring model so RevOps can inspect and tune them. Give Sales one combined score with a plain meaning. Most sales teams will use a number they understand more consistently than a matrix they have to interpret.
Should lead scoring control routing?
Scoring should first answer qualification and prioritization. Routing is a downstream policy. A company may route a high score directly to an AE and send a lower score to SDR review, but those actions should remain configurable without rewriting the score.
How should missing data affect a lead score?
Missing data is unknown, not negative evidence. High-intent requests can trigger enrichment, while lower-intent records can wait for a cheaper scheduled or moment-of-use process. Only score a fit field after you know where its value came from and how fresh it is.
How often should a lead scoring model be reviewed?
RevOps should monitor acceptance, disqualification, progression, conversion, overrides, and score distribution continuously, then run a structured review at least quarterly. A threshold can change sooner when the evidence is clear, as long as RevOps records and communicates the change.
Should AI own a lead score?
Usually no. Let deterministic rules or an accountable person own the operational score. AI can independently evaluate unstructured evidence, suggest component values, explain disagreements, and identify missing context. Store its advisory score separately so a confident model does not silently become the source of truth.
What belongs in a lead scoring scorecard?
A lead scorecard should name each fit and intent signal, its source, eligibility rule, weight, cap, decay, hard gate, missing-data behavior, owner, and evidence for changing it. It should also define score bands, the MQL and SQL acceptance boundaries, and the downstream action each band permits.
What is the difference between an MQL and an SQL?
An MQL is a person Marketing has qualified against the agreed fit and intent contract. An SQL is a person Sales has reviewed and accepted for active selling. The exact labels can vary, but the acceptance event, timestamp, rejection reasons, and return path must be explicit.

Get the next guide

New guides and the occasional note on GTM tooling. Don't worry, I won't drop you into a three-month nurture.