gtmjosh

Operating a CRM context layer

· 6 min read· Salesforce · HubSpot

The part after it works: a maturity model, the rule lifecycle, expiry and review SLAs, what to do when the CRM schema changes underneath you, and the metrics worth watching.

On this page

A lot of implementation guidance stops at "and then it works." This is the part after that: knowing where you are, keeping rules honest as the CRM moves underneath them, and being able to tell whether the thing is healthy.

Where you actually are

You don't need all of this, and the top of the ladder isn't the goal. Find your rung and decide whether the next one is worth it.

StageYou haveYou still get burned by
0. Raw schemaThe agent can see field names and typesField labels can be mistaken for meaning
1. Documented schemaHelp text pulled in as baseline contextStale descriptions, fields with no help text
2. Authority and provenanceWhich fields to trust, who writes them, how freshDefinitions, exclusions, junk records in counts
3. Business rulesFiscal calendar, exclusions, attribution, anti-patternsThe agent reaching things it shouldn't
4. Hard enforcementObject and field allowlists checked in codeRegressions discovered through usage
5. Regression testingGolden questions on both sides of every changeNot knowing why an answer changed
6. ObservabilityRun logs, metrics, attributable failuresOperational blind spots become much smaller

Many teams can get real value at stage 3 and stay there for a while. Stage 4 is the one that changes the risk profile most, because everything below it is guidance a model can fail to follow. If you are going to invest in one additional layer, hard enforcement is often the most consequential.

Stages 5 and 6 earn their keep once more than one person is editing rules, or once the answers start feeding decisions that are not manually checked every time.

The rule lifecycle

A rule isn't a document, it's a small piece of production configuration. Give it states and make the transitions deliberate.

  1. Step 1
    Draft
  2. Step 2
    Approved
  3. Step 3
    Active
  4. Step 4
    Deprecated
The production path is deliberate: write, review, activate, then retire without deleting the history that explains older answers.

A draft can be discarded before activation. An active rule can also be retired when its logic is no longer valid; when another rule replaces it, store the successor relationship explicitly rather than erasing the old record.

draft. Written, not loaded by anything. Where auto-generated candidates land.

approved. Reviewed by somebody other than the author. Still not loaded.

active. Loaded by agents. The state that reaches production.

deprecated. Retired without being deleted, with replacedBy pointing at the successor. Deleting a rule destroys the ability to explain an answer from six months ago.

Each rule should carry who owns it, who reviewed it, when it was created and last reviewed, why it's true, the failure or requirement that caused it, the golden questions covering it, and its successor if any. The schema has a field for each.

originFailure is easy to omit. A year in, it helps distinguish a rule that still protects against a known failure from one that nobody remembers the purpose of.

Expiry and review

Two different clocks, and conflating them is how stale rules survive.

expiresOn is for rules that stop being correct. A fiscal-period definition, a territory rule that only applies to this year's segmentation, or a temporary exclusion during a data migration. Past that date the rule shouldn't load at all.

reviewAfter is a freshness SLA for the rule itself, separate from the freshness of the CRM field it describes. A rule nobody has looked at in two years isn't automatically wrong. It is unverified until somebody reviews it, and that distinction should be visible.

A rule past reviewAfter can still load if that is safer than silently dropping it, but the payload should mark it, and the answer should be able to say the guidance is overdue for review. That's the known_stale resolution state.

The queue this produces is the useful artifact. For example: twelve rules haven't been reviewed in a year, and four of them govern the objects people ask about most.

When the schema moves underneath you

The failure here can be quiet: a field changes while the rule about it keeps loading.

Five drift cases, and they need different handling:

DriftWhat happens without a checkWhat should happen
Field deletedRule references a field that isn't thereRule fails loudly; goes to review queue
Field renamedA lookalike field may be mistaken for the originalDetected by API-name diff; rule flagged
Type changedExisting thresholds or interpretation may stop making senseFlagged; interpretation reviewed
Picklist value addedA rule that enumerates values may omit the new oneFlagged; valueMap reviewed
New field appearsNo approved guidance exists yetEnters the draft queue as unreviewed

The mechanism can be a scheduled diff between the live schema and the fields your rules reference. That's the Properties API anti-pattern: schema review belongs in a maintenance process, not necessarily in every user query.

The last row matters more than it looks. A newly added field starts without approved guidance, and the honest position is that a field with no approved guidance isn't authoritative yet. That's a rule you can write once, globally.

What to watch

If you're already logging guidance version, rule IDs loaded, planned versus actual objects and fields, tool calls, answer or refusal, latency, and cache state, these are useful metrics to derive from that.

MetricReading it
Rule hit rateWhich rules ever load. A rule that never fires may be dead or badly titled
Refusal rateShould be interpreted in context; zero is not automatically healthy
Tool-boundary rejection rateOff-allowlist requests caught in code
Stale-rule rateShare of loaded rules past reviewAfter
Regression pass rateGolden-set results over time, per release
Cache hit rateBelow expectations can indicate volatile data in a stable prefix
Cost per questionTrack over time as payload size changes
Correction rateHow often a human says the answer is wrong

Two deserve special attention.

Tool-boundary rejection rate at a flat zero does not necessarily mean users are perfectly in scope. It can also mean the boundary is not being exercised or logged. Send a known out-of-scope test request and confirm the rejection appears.

Correction rate is a direct signal of user trust in the answers. Everything else mostly says whether the system did what it was configured to do. A rising correction rate with a green golden set is a reason to check whether the test questions still resemble real usage.

Rule hit rate is a quiet diagnostic. A rule that never loads may be genuinely dead or may be titled too generically for retrieval. That's often easy to fix once you can see the pattern.

A reasonable cadence

  • On every rule change: Schema validation, targeted golden questions, deploy rule text separately from retrieval logic.
  • On every release: Broad and adversarial golden tiers, regression diff against the previous guidance version.
  • On a schedule: Schema drift diff into the review queue.
  • Monthly or quarterly: Rules past reviewAfter, rules with zero hits, correction-rate trend.
  • Whenever something goes wrong: Trace it, write the rule, add the golden question, record originFailure.

That last line is the whole loop. Everything else is scaffolding that keeps it working once the library is bigger than one person's memory.

FAQ

How mature does a CRM context layer need to be?
Locate yourself on a ladder rather than aiming for the top: raw schema, documented schema, authority and provenance, business rules, hard enforcement, regression testing, observability. Many teams can get real value at stage 3 or 4 without needing the last two immediately. The stage that changes the risk profile most is hard enforcement, because everything below it is advice the model can ignore.
What happens when a CRM field changes after you've written a rule about it?
The rule can keep loading even after the schema changes, which is the problem. Schema drift needs a scheduled diff between the live schema and the fields your rules reference, producing a review queue: fields that vanished, changed type, gained picklist values, or appeared with no guidance. A rule pointing at a field that no longer exists should fail loudly rather than silently.
Should context rules expire?
Some should. A fiscal-period rule is only correct inside its window, so it gets expiresOn. Other long-lived rules can use reviewAfter, which is a freshness SLA for the rule itself, separate from the freshness of the CRM field it describes. A rule nobody has reviewed in a long time is not automatically wrong, but it is unverified until somebody checks it.
What metrics matter for a CRM AI agent?
Rule hit rate, refusal rate, tool-boundary rejection rate, stale-rule rate, regression pass rate, cache hit rate, cost per question, and correction rate. A flat zero on tool-boundary rejection can mean the boundary is not being exercised or logged, so test an intentionally out-of-scope request rather than assuming zero is healthy.

Get the next guide

New guides and the occasional note on GTM tooling. Don't worry, I won't drop you into a three-month nurture.