AI Systems & Workcraft for Entrepreneurs | LachlanCB

Experimentation

The method for deciding whether an idea deserves more resource — and for killing it cleanly when it doesn't.

This works when the experiment is bounded and measurable: a routine change, a content cadence, a tool adoption. It degrades badly for anything with diffuse outcomes, where the scorecard can't hold a real signal and you end up scoring a vibe. The recurring failure mode is running a test with no exit criteria, so it drifts until it's abandoned rather than reaching a verdict — and an abandoned test teaches you nothing while still costing you the month. The structure below exists to force the verdict before the experiment starts rather than after it stalls.

01

Time-boxed Sprint, Forced Verdict

A bounded test against a falsifiable claim, scored on defined metrics, ending in a mandatory binary decision: promote or retire.

Window
Two weeks beats a month, in practice
Hypothesis
Falsifiable claim — not a goal
Scorecard
One primary metric beats a composite
Ending
Mandatory verdict — promote or retire
Record
Both outcomes archived; killed tests are assets

Five components. A named hypothesis — a specific falsifiable claim, not a goal, which is the distinction most people skip and the one that determines whether the test can produce information at all. Defined inputs: the routine, tool or behaviour under test. A scorecard mixing subjective markers with something objectively countable. A reflection loop partway through. And the decision at the end.

That last component is what separates this from an informal trial. The experiment doesn't simply stop — it receives a verdict, and the verdict is recorded whichever way it goes. Killed tests are as much of an asset as promoted ones, and considerably more likely to be forgotten if nobody writes them down. The record is what stops the same idea arriving again in eight months wearing a different name.

In practice the hypothesis and scorecard stages get completed reliably and the mid-point reflection frequently gets skipped. Experiments with a clear daily signal are tracked consistently; ones needing assessment across a longer arc are consistently under-reviewed. The honest read is that this runs better on a two-week window than a thirty-day one, and better against a single primary metric than a composite score. I've adjusted toward both.

Define the exit before the start. A test with no kill criteria doesn't end — it just gets abandoned.

This is the method. What actually got run, and what it cost, lives on the field tests page.