Selected experience
Global AI assistant: multi-region quality programme at $2M scale
Global consumer AI assistant product
A consumer AI assistant needed consistent quality evaluation across three regions, delivered by a large distributed team with high attrition.
Outcomes
- ~$2M programme owned across APAC, EMEA, and North America
- Programme cost reduced ~15% while output quality improved ~10%
- 42-person multi-country team led; 50% of contractors converted to permanent staff
- Six team members promoted; attrition stabilised
Quality at this scale was a people-system problem wearing a measurement problem's clothing.
Locate → Land → Live, retroactively, describes this engagement.
A global consumer AI assistant product needed its quality evaluation programme run across three regions. The founder led it as Global Quality & Triage Lead, and earlier as EMEA Service Delivery Lead, through a top-five global consultancy — not as A&O.
Locate meant recognising what was actually broken. The programme was framed as a measurement problem: were the evaluations accurate, were the rubrics right. The deeper issue was that a distributed contractor workforce with high churn cannot produce consistent judgment, because consistency is a property of people who have been there long enough to develop it.
Land meant treating the team as the instrument. Roughly half the contractor base was converted to permanent staff, six people were promoted into leadership, and attrition stabilised. A 42-person multi-country team ran evaluation cycles that had previously been undermined by constant re-onboarding.
Live meant holding the improvement while reducing spend. The programme ran at roughly $2M across APAC, EMEA, and North America, and came down about 15% in cost while output quality rose about 10% — an unusual pairing, and one that came from stability rather than from tooling.
This engagement is the origin of a view that runs through our training work: evaluation quality tracks workforce stability far more closely than it tracks rubric design. Most organisations try to fix judgment with better instructions, when the constraint is that nobody stays long enough to build judgment in the first place.
