For many financial institutions, 70% internal categorisation accuracy looks passable. It is familiar and keeps the reporting moving. And it’s what our data science experts have seen in the market as an internally acceptable rate for years. But a strong credit risk model leaves little room for ‘mostly right’.
A misread transaction can affect how a lender views income, commitments, affordability, fraud signals, and borrower behaviour. The Prudential Regulation Authority expects banks to manage model risk effectively, including how models are governed, used, and challenged.
So if the data going in is only 70% reliable, firms need to ask a difficult question: how much trust can they place in the risk models decision output?
A false sense of ‘good enough’
A financial risk dashboard that gets 7 out of 10 transactions right can sound reasonable. For broad internal reporting, it may even be workable (though at Moneyhub we’d advise against it). A finance team can still see spending trends. A product team can still spot some customer behaviour. A management dashboard can still tell a loose story. Credit modelling is different.
A credit risk model is not just looking for broad patterns. It is helping a lender decide whether a customer can afford credit, how likely they are to default, whether their behaviour has changed, and whether a case needs further review.
That makes the missing 30% more serious for credit risk management. If rent is mislabelled, affordability can be distorted. If gambling spending is missed, the risk may be understated. If business income is treated as personal income, a borrower may appear more stable than they are.
UK FCA rules require firms to base creditworthiness assessments on sufficient information and consider whether customers can repay sustainably. Those expectations are harder to meet when the transaction layer is treated as ‘good enough’ rather than dependable.
Small errors travel further than expected
Categorisation sits early in the decision chain, which is exactly why it matters. One incorrect category can affect several subsequent checks.
For example, if supplier payments from a self-employed borrower are miscategorised as personal retail spend, the affordability assessment may overstate their day-to-day outgoings.
That same credit analysis error can then distort income assessments, fraud checks, and stress testing, because the model is working from the wrong view of the borrower’s financial behaviour.
By the validation stage, it no longer appears to be a single bad category. It looks like a weaker credit risk assessment built on unreliable inputs:
- A commercial loan repayment, read as retail spend, changes the borrower’s story
- Irregular income placed in the wrong bucket can change the borrower’s risk profile
- A lender may treat one-off payments as stable income, approve a higher credit limit, or miss signs that affordability is tightening
The impact is not just a weaker lending decision. It can increase default risk, create avoidable cash flow losses, and make the model harder to defend during review.
That is the danger with ‘standard’ internal categorisation accuracy at a financial institution. The problem rarely announces itself as one obvious failure. It seeps into the model through thousands of small errors.
Over time, those errors can influence:
- Credit risk assessment
- Affordability calculations
- Fraud checks
- Default risk and probability
- Credit limit decisions
- Stress testing
- Model validation
70% accuracy is not a data tidiness problem but a decision quality problem.
The edge cases are where accuracy earns its keep
Most lenders can spot a clear yes or a clear no.
The harder cases sit in the middle. These are customers with thin credit histories, irregular earnings, shared costs, self-employed income, or spending patterns that do not fit standard credit rating or credit scorecard modeling.
That middle ground is where better transaction insight earns its keep. A borrower with irregular income may still be a strong applicant if their account history shows more money coming in than going out over time. Regular savings behaviour and a clear record of on-time repayments can still signal a strong candidate. Without that context, the applicant may be auto-declined, with the bank missing out on a strong new customer.
A customer under temporary pressure may not pose the same risk as someone experiencing sustained financial stress. Another applicant may look affordable under your standard credit risk definition, until the model picks up commitments that do not appear clearly elsewhere.
This is where alternative data can help spot market risk and support better lending decisions, but only when it is clean enough to use.
Transaction data can give lenders a more current and reliable view of borrower behaviour than a traditional credit file alone. Aggregated account data shows real income, spending, repayments, and commitments across more sources, giving models greater granularity than broader credit reference agency data or demographic estimates.
Poor categorisation weakens that advantage, because the model may still misread the customer’s actual financial behaviour. That’s where Admiral Money turned things on their head – feeding market-leading categorisation and enrichment into their affordability checks for faster approvals, without increasing risk.
By adding richer transaction insights into the underwriting process, Admiral Money was able to approve an additional 18% of loan applications that would previously have been referred or declined at the edge case stage, while reducing underwriting time and strengthening repayment rates and fraud detection.
Manual review should not be the operating model
Low accuracy does not make the work disappear. It pushes the work onto analysts.
When categories cannot be trusted, people have to check them. They correct obvious financial risk management mistakes, question borderline cases, and make judgment calls when the data is unclear. That may be fine for a handful of applications, but it becomes a problem at scale.
Skilled analysts should focus on risk to reduce borrower default rates, not constantly cleaning the inputs. Once manual review becomes the hidden operating model, speed and consistency both suffer.
Fraud teams already know this pattern. Too much noise can create:
- Alert fatigue, where teams are forced to review too many low-value or low-risk cases
- Poor prioritisation, where the wrong cases rise to the top of the queue
- Missed risk signals, because genuine issues become harder to spot
- Operational drag, as teams spend too much time dealing with avoidable friction and false positives
The same problem appears in lending. If the model needs people to keep correcting the data, it is only scalable for as long as the team can absorb the extra work.
This matters more as firms expand the use of Agentic AI across risk, fraud, and customer operations, making the data layer harder to ignore. Faster decision-making only helps if the inputs are clean enough to trust.
Automation can speed up a good process – Moneyhub works through an applied AI methodology to ensure that usage is considered and functional. The wrongful application of AI can also speed up a poor process. We therefore believe that predictive models needs cleaner inputs, not manual rescue work after the fact.
Poor provenance makes decisions harder to defend
Credit decisions need a clear audit trail. A lender needs to show why a customer was accepted, declined, repriced, or flagged. That becomes harder if the underlying categorisation is uncertain.
The FCA’s public censure of Amigo Loans shows why this matters. The regulator found that the company failed to conduct adequate affordability checks, creating a high risk of consumer harm.
This is where 70% accuracy creates a governance issue. If the model uses a category that was really a best guess, the decision built on top of it becomes harder to explain. With Consumer Duty demands also in the picture, firms may find it difficult to evidence good customer outcomes when the financial data underpinning lending decisions is not reliable enough to withstand scrutiny.
Accuracy now affects competitiveness
A more accurate view of income and spending can help lenders identify lower-risk customers. It can also help them spot early signs of pressure before the situation becomes harder to manage.
The opposite, poor categorisation and enrichment, can make a lender too cautious in some cases and too confident in others. That can lead to:
- Good applicants being declined because their affordability looks weaker than it is
- Riskier borrowers passing through because important warning signs are missed
- Analysts spending more time checking transactions than improving credit scoring models
- Lending decisions becoming harder to explain, defend, and improve over time
Fraud detection has the same dependency. Moving beyond batch processing to identify fraudulent transactions in real-time, and based on accurate transaction data, gives fraud teams confidence to act. They can step in swiftly and decisively at the first signs of suspicion, rather than after customer accounts have been drained.
By clarifying the merchant, category, location, and customer behaviour behind a transaction as it happens, firms can spot suspicious spending earlier and trigger faster interventions within the banking app, such as a customer nudge, a step-up check, or a payment block.
Moving beyond 70%
A 70% accurate system can survive because teams learn to work around it.
But workarounds are not free and don’t last for long. They cost time, slow down lending decisions, and make model outputs harder to defend.
For credit risk modelling, categorisation and enrichment influences the quality of the entire decision process.
Categorisation and enrichment shape the quality of:
- Affordability checks, by helping firms understand income, spending, and commitments
- Borrower analysis, by showing how customers actually manage money over time
- Fraud monitoring, by making unusual behaviour easier to spot
- Model validation, by giving teams clearer evidence to test against
- Customer outcome reporting, by helping firms evidence how decisions were reached
The business case for improving accuracy is not about chasing a perfect score. It is about removing avoidable uncertainty from decisions that carry financial, regulatory, and customer impact.
Specialist enrichment engines can push accuracy into the high 90s, giving lenders a clearer view of borrower behaviour and financial resilience.
About Matt Barr
Matt Barr is a Product Director here at Moneyhub. He’s been working either with or for banks since the mid-00s, solving all manner of problems. From ISA transfers to corporate actions, Matt now focuses on transaction categorisation and enrichment. When he’s not solving client problems, you can find Matt buried under his children’s laundry or stomping through the Peak District.
FAQs
It depends on the use case, but 80% accuracy is often too low for financial decisions that affect affordability, pricing, risk assessment, or customer outcomes. If one in five outputs is wrong, firms may decline suitable applicants, approve riskier borrowers, or struggle to evidence that customers received fair and appropriate outcomes
You can automate affordability assessments by using permissioned financial data, accurate categorisation, enrichment, and clear decisioning rules that explain income, spending, commitments, and risk signals.
There is no universal minimum because accuracy depends on the model, use case, and risk level. Still, firms must show that the model is reliable and supported by good-quality data.
share