Researchers ran a controlled experiment with employees at a fintech firm, splitting them into three groups by relevant expertise and having each conceptualize and draft an article with and without AI. Reported in the March to April 2026 Harvard Business Review, the results split cleanly. With AI, all three groups generated ideas at a similar level. In the writing stage the medium-expertise group nearly matched the experts, while the low-expertise group showed minimal improvement.
The researchers call it knowledge distance and use it to explain the split: the smaller a person's distance from the work, the better they can tell a good draft from a merely plausible one, and refining the output turns out to be where the value sits.
What the experiment measured
The design isolates the variable: same task, same tool, three groups whose jobs sat at different distances from the work (writers, marketing specialists, and developers and data scientists).
At the conceptualization stage, where the job was to generate and shape ideas, all three groups reached roughly the same level with AI. At the writing stage, where the job was to judge what the model produced and refine it, they separated again. The medium-expertise group came close to the experts, and the low-expertise group barely moved. That split is what expertise distance describes: AI closed the gap for people whose jobs sat near the task and hit a wall for people whose jobs sat far from it.
The more interesting study would’ve given all participants access to the same business information, a single source of truth to work off of, and measured how each applied it to their output.
Generating options, and choosing between them
Two things happen when someone puts a task through a model: it produces options, and a person decides which one is right. The experiment separates them. Every group generated comparable ideas with AI. Only the people with some relevant expertise turned what the model wrote into writing close to expert level.
That is why identical software produces such different work across a team. The tool is the constant. The person holding it is the variable, and what varies is whether they can recognize a confident wrong answer.
Why it shows up as output that looks right
If you want to hand work off and you want the standard to hold, you need to build a knowledge database for your firm. That doesn’t just mean materials, vendors, SOPs. Take it a level deeper so the models have access to your project scopes, phases to execute, time per task, etc. to be able to project out longer models of how your business will operate and scale. There are different levels a business can scope out its shared context through project assets and business assets, but it is all about building signals to identify that growth over time.
AI looked like the resolution. Give a less experienced person a model and the model supplies the expertise they haven't accumulated yet.
The experiment suggests that is where it breaks. What comes back reads well, and work that reads well is harder to catch than work that reads badly. You need the same expertise you were trying to route around, and you need it later in the process, with more volume attached.
Expertise distance as a scoping test
You can use expertise distance as a test, but every team member will have different experience; that is how you should build a team. Before you hand a task to a model, ask one question about the person who will receive the output: could they catch a confident wrong answer here? If they can, the tool will make them faster and probably better. If they can't, develop a shared source of truth so the model understands those pitfalls and doesn't make them.
Our Briefings work through what building that looks like in practice, and the manifesto sets out why we treat AI as a collaborator that needs supervision.
Related questions
What is expertise distance?
Expertise distance is the gap between what a person already knows about a task and what the task demands. Research summarized in the March to April 2026 Harvard Business Review uses the same idea to explain why generative AI lifts some workers close to expert output while barely improving others.
Does AI help beginners more than experts?
In the HBR fintech experiment, no. All three expertise groups generated ideas at a similar level with AI, but in the writing stage the medium-expertise group nearly matched experts while the low-expertise group showed minimal improvement. Judging the output takes experience the tool doesn't supply.
Can AI replace the experience of a senior person?
Not on this evidence. Generating options takes far less time now, and choosing correctly among them still depends on what the person already knows. A model can retrieve and draft; someone with the relevant experience still decides whether the result is right.
How do I decide which tasks to automate first?
Ask whether the person receiving the output could catch a confident wrong answer. If yes, automate it now. If no, either keep the task manual or write down the knowledge the judgment depends on before you move it.
How we read this
This is the empirical version of something we hold as a matter of practice: AI is a collaborator that needs supervision. The practical consequence is an order of operations. Get the knowledge your business runs on into one place a person and a system can both reach, then automate against it. Automating first and hoping the model supplies the judgment is what produces volume nobody trusts.
What you can do this week
Take one task you were hoping AI would let you hand off. Write down the three things you check when you review that work, in your own words, at the level of detail you would use to explain it to a new hire in their first week. It takes about fifteen minutes. That short list is the standard that has been living in your head, and it does two jobs at once: it makes the handoff possible for a person, and it is the instruction a system would need later.
In week two you will find the list was incomplete, because what you check changes with the job in front of you. Add to it each time you review something. A list that grows with use is the one that stays true.
Working together
Flow State Found works with a limited number of businesses to make their best work their baseline. Firms usually reach us after handing work over and getting back something that reads well and is wrong. We get the standard out of your head into a form a person and a system can both use, which is what makes the handoff safe.
We take on limited engagements, so it starts with a conversation.
Start a conversationFor Deeper Context
- Harvard Business Review, Gen AI Won't Make Your Employees Experts, March to April 2026 issue
- Vendraminelli, DosSantos DiSorbo, Hildebrandt, McFowland, Karunakaran, Bojinov, The GenAI Wall Effect, HBS working paper, 2025
- HBS Working Knowledge, Gen AI Boosts Productivity, But Can't Turn Novices Into Experts, March 16 2026
Note on scope: the experiment measured writing, at one firm, on one kind of task. Expertise distance is a useful lens for scoping and hasn't been tested across every function a business runs