10/12/2018
Which borrowers will default, or become delinquent, or prepay? Those are questions that lenders wish they had a crystal ball to answer, and their hopes are being raised by the advent of artificial intelligence.
AI today, in the form of machine learning, is not perfectly predictive, but it is a lot closer than it was. It's in the laboratory – or, more specifically, the accelerator – of Synechron. The New York-based consulting firm announced on October 22 the launch of its AI Data Science Accelerators, and within it a project applying AI to credit risk.
Synechron head of data science Robert Huntsman said that powerful distributed computing software, especially the open-source variety, has enabled significant advances in machine learning and data analytics. AI Data Science Accelerators is the fifth such program launched through the firm's Financial Innovation Labs (FinLabs) innovation hub, operating in 11 locations worldwide.
Robert Huntsman Headshot
Clients are getting “much more insight into their credit portfolios than they have had historically,” says Synechron's Robert Huntsman.
The data science accelerators include AI Data Science, which ingests large volumes of structured and unstructured data, automates the generation of personalized buy- and sell-side research reports, alerts wealth managers to critical events, and identifies factors driving customer complaints. The credit risk component, according to Synechron, “empowers banks to manage their credit portfolios proactively, enabling users to drill down into the factors driving likely credit events and to proactively manage individual risks ranked by probability of incurring a specific credit event.”
Prescriptive Management
Synechron's credit solution cannot be used to underwrite loans. Huntsman said that lenders offering mortgage, credit card and/or auto loans have long used logistic-regression models for underwriting, and regulators have grown comfortable with that approach. The machine learning methods known as random forest, neural networks and clustering have yet to attain that regulatory status, but they can be used to manage credit prescriptively.
“A bank with an existing loan portfolio wants to know who is going to default or become delinquent on loans, and who will prepay them. These are the biggest challenges faced by our clients,” Huntsman said.
The accelerator program is built around massive amounts of data. In the case of mortgage loans, AI Data Science analyzes data on more than 2.1 million loans that Fannie Mae provides on its website. The data is anonymized, but it comprises 55 attributes and the entire payment history for each loan. Synechron's separate models for delinquencies, defaults and prepayments also incorporate macro-economic data, such as state unemployment rates and GDP growth.
High Accuracy
Huntsman said the models predict 95% of prepayments, and 80% to 85% of defaults and delinquencies.
“These numbers are high relative to what we've seen with traditional credit-rating models, where accuracy of 70% is considered very good,” he said. “We've seen that using machine learning provides clients with much more insight into their credit portfolios than they have had historically.”
The credit-related data dates back to 2001 and thus spans different credit cycles including one extreme: the 2008 crisis.
An advantage of neural networks, said Huntsman, is that they don't assume a normal distribution of data. Instead, they can be programmed to update weekly or even daily, thus taking in new data points much more frequently.
Similar to how human brains function, neural networks process new data in layers, with each layer containing multiple nodes. The 55 loan variables pass through the nodes in the first layer, each filtering and transforming input for the next layer, which subsequently processes input from prior nodes until a final layer produces loan predictions.
The last layer decides on whether the loan will default or not. If it decides there's a high probability of default, but in fact that never happens, a loss function proceeds to determine how inaccurate the prediction was and adjusts parameters in each node accordingly. Then the data is fed back through the network of nodes, and the next predictions it generates should be much more accurate.
Getting to Why
In a random forest model, the algorithm knows whether a loan has defaulted or not – the training set – and seeks to determine why by analyzing the 55 loan variables to see which ones made a difference. To do this, the algorithm splits the data, looking at each variable until it comes up with the best mix. It may determine that loans under $100,000 are not likely to default, nor are those whose borrowers have credit scores higher than 700.
“It's machine learning because it has to learn what those splits are, and which splits are most important,” Huntsman explained. “A random forest is nothing more than a group of decision trees, where you take all the loan data, split it X number of ways and create X number of trees.”
The algorithm then searches the different trees to find the data in common, which provides a likely indication for why a loan defaulted or became delinquent.
The random forest model is applied to non-defaulted loans. The 55 attributes are fed into the algorithm, trained to predict defaults, and the decision points for each split applied to non-defaulted loans. The borrower payment history is also incorporated into the model and could provide an important input if, for example, a borrower has a history of making payments on time.
“That serves as an additional tool to try to determine which of those borrowers could potentially default,” Huntsman said.
Responsive to Changes
Those inputs change over time, creating new stresses on borrowers. The high accuracy rates of Synechron's models suggest that they are successfully incorporating those changes.
“A machine learning model can adapt to a changing environment, and it's a big advantage to have adaptive models,” Huntsman said, adding that prior to the financial crisis, such models likely would have seen upticks in key indicators such as delinquency rates.
The term “accelerator” is apt in this case, he said, because the firm has essentially developed all the code necessary to implement the model in clients' existing systems. That amounts to a “git,” or a repository of algorithms that clients can use as a foundation and later customize.
“A client may not want to use the neural network we built, or it may have its own customer data, or different attributes they want to feed into the model,” Huntsman said.
A Human Responsibility
He noted that machine learning is a step toward full-blown AI, in which the computer is capable of autonomous reasoning and decision-making. Machine learning still requires a programmer to write the code and set parameters for how the machine learns over time.
Huntsman said full-blown AI is not in the foreseeable future, because humans ultimately will remain responsible for credit-related decisions. Synechron's tools learn how to analyze huge volumes of data more effectively and dynamically, but their intent is to alert humans to the critical data.
“In banking, there's a strong need for humans, and accordingly there's a strong need to give humans more information,” Huntsman said.