Skip to content
Menu

Method

How Kogsy measures whether an AI process change worked

By Henrik GrünfeldPublished Updated

AI is only useful if it changes an agreed operating result: cycle time, cost per case, rework, conversion, quality, adoption, or whatever else the business already tracks. A model that produces impressive output in a meeting is evidence that the technology works. It is not evidence that the business changed. Only the second claim is worth paying for, and this page sets out how I tell them apart.

Kogsy measures AI-enabled process change against a defined baseline. The method depends on the process: some questions require a census of records, some require work sampling, and some require controlled before-and-after observation. In every case, Kogsy defines the metric, source data, measurement window, uncertainty, client dependencies, and stop condition before judging whether a change worked.

The measurement, end to end
  1. Stage 1

    Define the process boundary

    Where the work starts, where it ends, and who owns it.

  2. Stage 2

    Establish the baseline

    What is true today, in the client’s own operating data.

  3. Stage 3

    Choose the measurement method

    Census, work sampling, or before-and-after observation.

  4. Stage 4

    Observe implementation and adoption

    Whether the system runs, and whether people use it.

  5. Stage 5

    Decide

    Continue, change, pause, or stop.

Where does the process start and end?

Nothing can be measured until it has an edge. Before any baseline work, we write down where the process begins, where it ends, and which adjacent work is deliberately excluded. Skip this and an improvement in one team quietly becomes extra work in the team next door, while the number still looks good.

One named person owns the process. Not a committee. If nobody can be named, that is worth knowing before any money is spent, because it usually predicts what will happen at implementation.

What was true before the change?

The baseline comes from the client's own operating data wherever it exists, not from a survey of opinion and not from my estimate. We state the period it covers, and we record volume, exceptions, rework, time, cost, and quality to the extent each is relevant to the question being asked.

A baseline is an evidence ledger

Separate facts, estimates, assumptions, and gaps before making a business case.

  • Measured

    Observed in named systems or records.

    Traceable source · Defined period

  • Estimated

    Calculated from known inputs.

    Method stated · Range retained

  • Assumed

    A judgement made before evidence exists.

    Owner named · Test planned

  • Missing

    Unavailable or not measurable yet.

    Gap disclosed · No false claims

Rule: no combined headline number until every input is classified.

The four categories stay separate because combining them is how a business case becomes fiction: once an assumption is averaged into a measured figure, nobody can tell which part of the total is real. If the baseline data turns out not to exist, that is a finding, not a setback.

Which measurement method fits the question?

There is no single correct method. There is a correct method for a given question, and choosing it badly produces confident numbers that do not survive scrutiny.

The three methods in detail
Which measurement method to use, when it applies, and what it suits
Census of recordsWork samplingBefore-and-after observation
Use when

Every relevant record can be reviewed, and the records are trustworthy enough to count.

Use when

Logging every activity would distort or interrupt the work being measured.

Use when

A defined intervention can be observed over a period long enough to see the effect.

Best for

Orders, cases, errors, approvals, and transaction history.

Best for

Office time, coordination work, interruptions, and administrative activity.

Best for

Cycle time, rework, adoption, quality, and conversion.

What can the method not establish?

Limitations are written before the results are interpreted, and reported next to the result rather than in a footnote. Four come up almost every time.

  • Data quality

    Records inherit the habits that produced them, so inconsistent systems widen the reported range.

  • Missing data and non-response

    Gaps are rarely random, because people skip logging exactly when they are busiest, and below the agreed response threshold the estimate is reported as indicative only.

  • Observer effect

    Measured work is not identical to unmeasured work, and I do not pretend to correct for the difference.

  • Causation

    Most operational measurement shows that something changed alongside an intervention, not that the intervention caused it, and the report does not claim otherwise.

A result that arrives with its limitations attached is more useful than one that arrives clean, because you can act on the first and only believe the second.

Is the system actually being used?

A technically working system that staff do not use is not a completed transformation. It is a completed build with no operational result, and measuring only the economics will hide that for months.

So the adoption signal is defined before launch: what counts as the system being used, by whom, and how often. A change in the economics without a matching change in adoption is treated as a warning rather than a win. Responsibility for implementation and training is named at the same time, because somebody has to make adoption happen and that person needs the time to do it.

What result would mean stopping?

Before the measurement runs, we agree two thresholds: the result that would justify continuing, and the result that means the initiative should pause or stop. Both are written down while nobody yet knows which way it will go. Setting a stop condition after seeing the data is not a stop condition, it is a negotiation, and the side that wants the project to continue usually wins it.

Doing nothing is a legitimate conclusion. If the honest reading is that the change will not move the numbers, the recommendation is to stop, and the work already produced stays with the client.

What does the client receive?

One document, written so that somebody who was not in the room can check the reasoning and reach their own view.

  • Process boundary
  • Baseline readout
  • Method and source-data record
  • Assumptions and limitations
  • Measurement plan
  • Decision gate
  • Recommended next step

Measurement readout

Illustrative structure, no client data

  1. Process boundaryStart, end, owner, and exclusions.
  2. Baseline and source dataNamed records, period, and data quality.
  3. Metric and measurement windowAgreed before the result is known.
  4. Assumptions and limitationsMeasured, estimated, assumed, or missing.
  5. Adoption signalWhat use looks like.
  6. Decision gateContinue, change, pause, stop.

Every conclusion should trace back to a named source and an agreed rule.

Where can you see this applied in full?

For office-work measurement in US non-medical home care, Kogsy publishes a full work-sampling protocol. It covers randomisation, non-response, confidence intervals, and limitations specific to that setting. That protocol is written for one setting and stays there. This page is the general standard the consulting work is held to.

Applied protocolOffice work sampling protocol, Kogsy GOgokogsy.com →

If you want this applied to one of your processes, the first conversation is thirty minutes.