Operating a CRM context layer
The part after it works: a maturity model, the rule lifecycle, expiry and review SLAs, what to do when the CRM schema changes underneath you, and the metrics worth watching.
On this page
A lot of implementation guidance stops at "and then it works." This is the part after that: knowing where you are, keeping rules honest as the CRM moves underneath them, and being able to tell whether the thing is healthy.
Where you actually are
You don't need all of this, and the top of the ladder isn't the goal. Find your rung and decide whether the next one is worth it.
| Stage | You have | You still get burned by |
|---|---|---|
| 0. Raw schema | The agent can see field names and types | Field labels can be mistaken for meaning |
| 1. Documented schema | Help text pulled in as baseline context | Stale descriptions, fields with no help text |
| 2. Authority and provenance | Which fields to trust, who writes them, how fresh | Definitions, exclusions, junk records in counts |
| 3. Business rules | Fiscal calendar, exclusions, attribution, anti-patterns | The agent reaching things it shouldn't |
| 4. Hard enforcement | Object and field allowlists checked in code | Regressions discovered through usage |
| 5. Regression testing | Golden questions on both sides of every change | Not knowing why an answer changed |
| 6. Observability | Run logs, metrics, attributable failures | Operational blind spots become much smaller |
Many teams can get real value at stage 3 and stay there for a while. Stage 4 is the one that changes the risk profile most, because everything below it is guidance a model can fail to follow. If you are going to invest in one additional layer, hard enforcement is often the most consequential.
Stages 5 and 6 earn their keep once more than one person is editing rules, or once the answers start feeding decisions that are not manually checked every time.
The rule lifecycle
A rule isn't a document, it's a small piece of production configuration. Give it states and make the transitions deliberate.
- Step 1Draft
- Step 2Approved
- Step 3Active
- Step 4Deprecated
A draft can be discarded before activation. An active rule can also be retired when its logic is no longer valid; when another rule replaces it, store the successor relationship explicitly rather than erasing the old record.
draft. Written, not loaded by anything. Where auto-generated candidates land.
approved. Reviewed by somebody other than the author. Still not loaded.
active. Loaded by agents. The state that reaches production.
deprecated. Retired without being deleted, with replacedBy pointing at the
successor. Deleting a rule destroys the ability to explain an answer from six months
ago.
Each rule should carry who owns it, who reviewed it, when it was created and last reviewed, why it's true, the failure or requirement that caused it, the golden questions covering it, and its successor if any. The schema has a field for each.
originFailure is easy to omit. A year in, it helps distinguish a rule that still
protects against a known failure from one that nobody remembers the purpose of.
Expiry and review
Two different clocks, and conflating them is how stale rules survive.
expiresOn is for rules that stop being correct. A fiscal-period definition, a
territory rule that only applies to this year's segmentation, or a temporary
exclusion during a data migration. Past that date the rule shouldn't load at all.
reviewAfter is a freshness SLA for the rule itself, separate from the freshness
of the CRM field it describes. A rule nobody has looked at in two years isn't
automatically wrong. It is unverified until somebody reviews it, and that distinction
should be visible.
A rule past reviewAfter can still load if that is safer than silently dropping it,
but the payload should mark it, and the answer should be able to say the guidance is
overdue for review. That's the known_stale resolution state.
The queue this produces is the useful artifact. For example: twelve rules haven't been reviewed in a year, and four of them govern the objects people ask about most.
When the schema moves underneath you
The failure here can be quiet: a field changes while the rule about it keeps loading.
Five drift cases, and they need different handling:
| Drift | What happens without a check | What should happen |
|---|---|---|
| Field deleted | Rule references a field that isn't there | Rule fails loudly; goes to review queue |
| Field renamed | A lookalike field may be mistaken for the original | Detected by API-name diff; rule flagged |
| Type changed | Existing thresholds or interpretation may stop making sense | Flagged; interpretation reviewed |
| Picklist value added | A rule that enumerates values may omit the new one | Flagged; valueMap reviewed |
| New field appears | No approved guidance exists yet | Enters the draft queue as unreviewed |
The mechanism can be a scheduled diff between the live schema and the fields your rules reference. That's the Properties API anti-pattern: schema review belongs in a maintenance process, not necessarily in every user query.
The last row matters more than it looks. A newly added field starts without approved guidance, and the honest position is that a field with no approved guidance isn't authoritative yet. That's a rule you can write once, globally.
What to watch
If you're already logging guidance version, rule IDs loaded, planned versus actual objects and fields, tool calls, answer or refusal, latency, and cache state, these are useful metrics to derive from that.
| Metric | Reading it |
|---|---|
| Rule hit rate | Which rules ever load. A rule that never fires may be dead or badly titled |
| Refusal rate | Should be interpreted in context; zero is not automatically healthy |
| Tool-boundary rejection rate | Off-allowlist requests caught in code |
| Stale-rule rate | Share of loaded rules past reviewAfter |
| Regression pass rate | Golden-set results over time, per release |
| Cache hit rate | Below expectations can indicate volatile data in a stable prefix |
| Cost per question | Track over time as payload size changes |
| Correction rate | How often a human says the answer is wrong |
Two deserve special attention.
Tool-boundary rejection rate at a flat zero does not necessarily mean users are perfectly in scope. It can also mean the boundary is not being exercised or logged. Send a known out-of-scope test request and confirm the rejection appears.
Correction rate is a direct signal of user trust in the answers. Everything else mostly says whether the system did what it was configured to do. A rising correction rate with a green golden set is a reason to check whether the test questions still resemble real usage.
Rule hit rate is a quiet diagnostic. A rule that never loads may be genuinely dead or may be titled too generically for retrieval. That's often easy to fix once you can see the pattern.
A reasonable cadence
- On every rule change: Schema validation, targeted golden questions, deploy rule text separately from retrieval logic.
- On every release: Broad and adversarial golden tiers, regression diff against the previous guidance version.
- On a schedule: Schema drift diff into the review queue.
- Monthly or quarterly: Rules past
reviewAfter, rules with zero hits, correction-rate trend. - Whenever something goes wrong: Trace it, write the rule, add the golden question, record
originFailure.
That last line is the whole loop. Everything else is scaffolding that keeps it working once the library is bigger than one person's memory.
- How to build an AI context layer for your CRM: the concept and full build
- The rule schema: every lifecycle and expiry field
- Golden-question testing: the regression half of this
- One CRM question, traced end to end: what a single run log looks like
FAQ
- How mature does a CRM context layer need to be?
- Locate yourself on a ladder rather than aiming for the top: raw schema, documented schema, authority and provenance, business rules, hard enforcement, regression testing, observability. Many teams can get real value at stage 3 or 4 without needing the last two immediately. The stage that changes the risk profile most is hard enforcement, because everything below it is advice the model can ignore.
- What happens when a CRM field changes after you've written a rule about it?
- The rule can keep loading even after the schema changes, which is the problem. Schema drift needs a scheduled diff between the live schema and the fields your rules reference, producing a review queue: fields that vanished, changed type, gained picklist values, or appeared with no guidance. A rule pointing at a field that no longer exists should fail loudly rather than silently.
- Should context rules expire?
- Some should. A fiscal-period rule is only correct inside its window, so it gets expiresOn. Other long-lived rules can use reviewAfter, which is a freshness SLA for the rule itself, separate from the freshness of the CRM field it describes. A rule nobody has reviewed in a long time is not automatically wrong, but it is unverified until somebody checks it.
- What metrics matter for a CRM AI agent?
- Rule hit rate, refusal rate, tool-boundary rejection rate, stale-rule rate, regression pass rate, cache hit rate, cost per question, and correction rate. A flat zero on tool-boundary rejection can mean the boundary is not being exercised or logged, so test an intentionally out-of-scope request rather than assuming zero is healthy.