The lab

AI x
Consulting

Consulting runs on research, synthesis and structure. All three are being rebuilt right now.

I run each change as an experiment with a stated method and a measured result, and publish what works and what does not.

Experiments
Market research with AI, and where it quietly gets things wrong
Question
Can a model produce a defensible market picture without a human checking every source?
Method
Same brief run manually and with AI, then every claim traced back to a primary source.
Result
[TBC]
Verdict
Pending
Competitor benchmarking at speed
Question
How much of a 55-competitor benchmark can be assembled before judgement has to take over?
Method
Rebuild the LivingPaper benchmark with AI collection, then compare coverage and error rate against the version delivered by hand.
Result
[TBC]
Verdict
Pending
Market sizing and issue trees
Question
Does a model build a mutually exclusive tree, or a plausible-looking list?
Method
Ten sizing prompts scored for MECE structure, assumption visibility and sensitivity to a changed input.
Result
[TBC]
Verdict
Pending
Interview and meeting synthesis
Question
Does AI synthesis keep the one sentence that changes the recommendation?
Method
Transcripts from the Promethias and Nikash conversations summarised both ways, then checked against the decisions those conversations actually drove.
Result
[TBC]
Verdict
Pending
Deck logic and slide argumentation
Question
Can a model hold a storyline across twenty slides, or only write good single slides?
Method
Action titles generated for a finished engagement, then tested for whether the titles alone carry the argument.
Result
[TBC]
Verdict
Pending
Where AI made the work worse
Question
Which tasks degraded when a model was inserted into them?
Method
Every experiment above kept a failure log: wrong numbers, invented sources, structure that looked right and was not.
Result
[TBC]
Verdict
Pending
Results marked [TBC] are not yet measured. Nothing is reported here until it is.
Video
Market research with AI, and where it quietly gets things wrong
[TBC] First experiment, walked through end to end with the failure log.
Video IDs [TBC]. The player loads only on click.