Testing the productivity claim behind an AI rollout
Developed in my Program Manager role and delivered to executive leadership as an internal analysis with recommendations.
What did the two-hour savings estimate actually measure?
-
A behavioral health organization introduced an AI documentation tool for clinical staff and estimated it would save each clinician about two hours a week. Productivity benchmarks were raised by two hours to absorb the assumed savings.
The tool was introduced through a company-wide live demonstration, with nothing to refer back to afterward. In the period that followed, managers reported strain on their teams, use of the tool was uneven, and turnover was higher than usual across the same stretch.
Leadership was already looking at where the presumed two hours of recovered capacity should go. More of the heaviest client-facing work was not necessarily the answer, so the conversation turned toward where else that time could be used. That made me want to go back to the assumption underneath the conversation: had the two-hour savings actually been measured and realized before expectations were built around them?
-
The estimate had moved from a projected benefit into an expectation people were held to. Before deciding where those presumed two hours should go, I wanted to understand whether the organization had enough evidence to treat them as recovered capacity in the first place.
That also meant looking beyond whether people had access to the tool or were using it at all. I wanted to know what effective use actually looked like across roles and whether the reported savings held up across different ways of working.
-
I designed a short assessment, facilitated it with my teams, and reviewed the responses. Because the group was small, I used the results to identify patterns and better next questions rather than establish a formal baseline.
What emerged was uneven. Some people reported meaningful time savings, while others did not. People were developing their own ways of using AI across the different kinds of documentation their roles required, and the workflows producing those savings were not consistent.
Prior familiarity with AI appeared to be one part of that variation. People who already understood how to work with AI tended to be better equipped to experiment, adjust when the output fell short, and find a method that worked for them. But that was one pattern inside a larger finding: access to the same tool had not produced the same way of working or the same result.
That shifted my attention away from a single estimate and toward the variation underneath it.
-
That started with some of the implementation basics that had not yet been built around the rollout: pulling back on the expanded productivity expectation while the actual savings were assessed, introducing the tool in cohorts, and giving people more than one opportunity to learn how to work it into their roles.
From there, I recommended building role-specific workflows, templates, examples, and guidance into a shared knowledge system people could return to. In a clinical documentation setting, that also meant making expectations for AI-assisted work explicit: what acceptable output looked like, what still required human judgment, and where employees had room to trust the tool without worrying that a different-looking note was automatically a wrong one.
The assessment could then continue alongside implementation, using both the time-savings data and the differences underneath it to see what was actually improving. Once there was a clearer picture of the capacity the tool was reliably creating, expectations could be adjusted from evidence rather than estimates.
My scope ended with the analysis and recommendations.