04 / LAB
Lab Systems research, not a blog. Every experiment has a hypothesis, a build, a measurement, a result and a decision. Where the result is unmeasured, the record says so — unmeasured is a verdict, and nothing changes an operating rule except a result a human decided from an experiment that had a control group.
Experiments … records
all agents curation dashboards data management software workflow
EXP-001 Should an agent ever leave OBSERVE? HYPOTHESIS
An agent that has produced correct drafts for long enough could be allowed to act on the narrowest of them.
BUILD
Three rungs — OBSERVE, DRAFT, ACT — with the rung on the identity card and enforced at the call site rather than in a prompt.
MEASUREMENT
Three recorded evaluations and a deterministic rule are required before any rung change.
DECISION
Every agent stays at OBSERVE. No model is connected, so there is nothing to evaluate.
EXP-002 Computed severity against configured dashboards HYPOTHESIS
A queue derived from the facts of each record will beat a dashboard whose thresholds someone configured once and forgot.
BUILD
One central function computes severity from the record: a person is waiting, a transaction is blocked, a promise is about to be missed.
MEASUREMENT
Items surfaced on open, against the count a badge-driven CRM shows.
RESULT
Five items, ranked, against eighty-four badges.
DECISION
Adopted. The morning brief is capped at five and folds the rest below.
EXP-003 Bypass outcomes HYPOTHESIS
If staff keep skipping a step and the outcomes hold, the step is the problem, not the staff.
BUILD
Every route override requires a reason, and the reason is a field rather than a note. What was skipped is recorded mechanically.
MEASUREMENT
Override frequency against downstream outcome, per step.
DECISION
Hold. This needs live traffic before any step is deleted.
EXP-004 The honest none HYPOTHESIS
Telling a customer that nothing available deserves recommending will cost a sale today and earn the next one.
BUILD
The scorer returns zero when the evidence score sits below the floor, and writes the sentence explaining why.
MEASUREMENT
Return rate and referral rate among customers who received a none.
DECISION
Shipped as the default regardless. Some rules are not run as experiments.
EXP-005 Degradation — does the kernel survive? HYPOTHESIS
The minimum path to a completed transaction should run with every optional layer off.
BUILD
A test that switches each product application off in turn and drives the kernel: vehicle, truth, appointment, numbers, decision, transaction.
MEASUREMENT
Kernel completion with each layer disabled.
RESULT
Passes with every layer off.
DECISION
Adopted as a permanent gate. Nothing above the kernel may become mandatory.
EXP-006 Can any agent reach a human-owned decision? HYPOTHESIS
Written boundaries drift. Executed boundaries do not.
BUILD
Every agent tried against every human-owned decision at every rung.
MEASUREMENT
3,828 attempts. Every one must fail.
DECISION
Adopted. The suite runs on every build.
EXP-007 A fixture model provider HYPOTHESIS
Agent governance can be built and proven before any model is connected.
BUILD
A provider that satisfies the interface and stamps every response as a fixture, so nothing can mistake it for inference.
MEASUREMENT
Whether governance, logging and approval work end to end without a model.
RESULT
They do. Thirty-two drafted actions, each with a recorded decision.
DECISION
Keep until a model is chosen. Connecting one changes no agent's autonomy.
EXP-008 A claims guard in the build HYPOTHESIS
A number that stops being true will otherwise live on a website forever.
BUILD
A guard that fails the build if a retired proof number appears anywhere in the repository.
MEASUREMENT
Builds failed by the guard.
RESULT
In force. The number on this site is the number in the repository.
DECISION
Adopted, and extended to this website.
Rules 6 records
Source → rule → system → measurement. A principle the company wrote down, the operating rule it became, the function that enforces it, and what is counted to prove it holds.
RUL-001 An explorer should not be pursued. PRINCIPLE
Someone browsing on a Tuesday night is not a sales opportunity yet.
OPERATING RULE
Explore intent cannot create a sales lead.
SOFTWARE
The lead-creation function refuses an Explore intent. It is not a step someone follows; it is a record that cannot be created.
MEASUREMENT
Boundary violations are counted. The seed asserts: none.
RUL-002 Every contact needs a reason. PRINCIPLE
Follow-up is not persistence. It is a new reason to continue.
OPERATING RULE
No message leaves the company without a structured reason.
SOFTWARE
One function, through which every message passes. It checks consent, hours, the cadence cap and the reason. There is no force flag.
MEASUREMENT
Refusals are logged with their cause, including the ones that never became a row.
RUL-003 A claim requires evidence. PRINCIPLE
When the seller knows more than the buyer, trust collapses and good stock leaves.
OPERATING RULE
A vehicle file cannot be published without an inspector of record.
SOFTWARE
The controller refuses the state change. An owner's assertion and an inspector's evidence are never quietly merged.
MEASUREMENT
Unpublished files, and the specific evidence each one is missing.
RUL-004 An invented third choice is worse than a short list. PRINCIPLE
Six jams outsell twenty-four. A fabricated sixth is worse than five.
OPERATING RULE
The scorer may recommend three, one, or none.
SOFTWARE
A selection set accepts a reason for none and stays empty. Nothing forces a slot to fill.
MEASUREMENT
The honest-none rate, reported rather than hidden.
RUL-005 A profile is not permission to pursue. PRINCIPLE
Permission is given per channel, and it can be taken back.
OPERATING RULE
Consent is recorded per channel and checked at send time, not at signup.
SOFTWARE
The gate reads the consent record, the person's hours and the cadence cap before the message exists.
MEASUREMENT
Blocked sends, by cause. One blocked message is a working gate, not a bug.
RUL-006 Four funnels, never combined. PRINCIPLE
A person who is not buying yet is not a failed sale.
OPERATING RULE
Lanes are reported separately, always.
SOFTWARE
The function that would merge them raises. Steps the platform does not measure read “not measured”, never zero.
MEASUREMENT
Merge attempts are a test, and they fail.
BENCHMARKS
Not yet. A benchmark needs an anonymised dataset large enough to be honest about. When there is one — lead response time, administrative minutes per transaction, cost per sold unit — it appears here as a record with its sample size.