I once reviewed an SME insolvency model I’d built with an industry data partner. The validation charts were good. Then someone from operations asked what their team should do with an account scored at 0.6, and where an analyst’s disagreement with the score would go. Nobody had a good answer, me included. I’d spent months on the score and almost no time on the decision it fed.
Picture a credit team on a Monday morning. The model has scored four thousand accounts overnight. Two analysts can properly review about thirty a week, so the model’s real output is rows one to thirty.
That changes what a good model is. One with a better overall accuracy score (AUC) can do worse on those thirty rows, because it buys its improvement around row two thousand, where nobody reads. The right choice depends on how many reviews the team can do and what a missed insolvency costs against a wasted afternoon. None of that is in the training data.
The analyst also needs the reason, in their own terms. A score of 0.34 gives them nothing to take to a relationship manager. “Up from 0.08 last month, mostly because days beyond terms have doubled and two county court judgments were registered” is a phone call to make.
Two cheap things decide whether the model gets used. Print the as-at date next to every score, because a weekly committee reading a monthly score is often looking at numbers four weeks old. And record every override with a reason someone can count. If ten overrides in a quarter cite a parent company guarantee the model can’t see, you know what to add next. Without that record, the feedback ends up in an email thread.
The record matters for another reason. If a decision to pull a facility is ever challenged, I doubt anyone will start with the validation charts. I’d expect the questions to be what the lender knew on the day, what the model said and who acted on it.
If I started the insolvency work again, I’d use a baseline model anyone could explain and spend the first weeks on the Monday list: how long it is, what each row says, where the overrides go. A more accurate model can come later, once the analysts can tell me which improvements they’d notice.