NG Solution Team
Artificial Intelligence

Alternative Credit Data Is Transforming Credit and Fraud Models

Alternative credit data closes two gaps left by traditional bureau files: it provides information on applicants who have never borrowed and on the device or session submitting the application. These signals now feed both credit-scoring and fraud-detection models, reshaping how lenders assess repayment capacity and detect abusive or automated applications.

What counts as alternative credit data

Alternative credit data means any input used in lending or fraud decisions that does not come from a credit bureau file. In practice it falls into several categories. Cash-flow and transaction records capture deposit regularity, balance volatility and payment behaviour on existing obligations. Utility, telecom and rent payments create a repayment record for applicants with no formal borrowing history. Payroll and employment data provide income verification independent of what the applicant declared. Device and connection attributes cover hardware and browser configuration, virtualisation and remote-access flags, and network and geolocation consistency. Behavioural session signals include typing rhythm, pasting into sensitive fields and navigation that is either too hesitant or too linear. Treating these categories as a single input is a common early-adoption mistake: the first three speak to capacity and willingness to repay, while the last two indicate whether a session looks like one real person or like an emulator, a remote-access tool, or an operator processing applications in volume.

Thin-file borrowers are unmeasured, not necessarily high-risk. Segments include young borrowers, recent migrants, the self-employed and gig workers with irregular deposits. In fast-growing digital-lending markets — where informal income is common and credit penetration low — much of the funnel can have no usable file. Lenders faced with thin files either decline by default and lose the fastest-growing customers, or approve on evidence too thin to price risk correctly. Alternative data restores separation in both cases: these signals are additive, sitting on top of bureau data where coverage is good and carrying more weight where bureau coverage is poor.

Where credit and fraud models meet

Device and session data are the intersection between credit and fraud models, which has pushed alternative data out of isolated credit teams and into fraud operations. Randomisation illustrates the challenge: device and browser attributes can be masked cheaply, so one applicant can present a different technical identity on every attempt. Conventional checks tend not to register this because a randomised device appears new rather than explicitly wrong. Virtualisation flags, geolocation that contradicts the declared address, fields completed by paste rather than typing and sessions with no hesitation behave similarly. None proves fraud alone — legitimate users travel, change handsets and use unfamiliar networks — but correlation across several anomalies in a single session becomes a pattern worth acting on.

This is the logic behind JuicyScore’s approach to fraud prevention: evaluating a wide set of device, connection and behavioural attributes in real time, without relying on direct user identifiers such as name, phone number or email, and returning a score plus a vector of attributes the lender feeds into its own models. The same output serves as fraud stop-markers at the top of the funnel and as additional separation inside the credit model further down.

The supply of usable data is narrowing under two simultaneous pressures. Regulators are the visible constraint: GDPR set the benchmark and comparable regimes now operate across most active lending markets, each with its own consent and data-minimisation requirements. Platform policy is the quieter constraint: app stores have restricted what lending applications may request from a handset, closing off contacts, photos, call logs, precise location and external storage, and browsers block cross-app tracking by default. Data a lending app could routinely collect a few years ago is often now unavailable regardless of local regulatory permission.

Where alternative data models fail, three failure modes recur. First, stacking sources without measuring incremental lift — each new input must be tested for what it adds beyond the model already in production, not only on its standalone performance. Second, overstating the privacy position: device identifiers and IP data can qualify as personal data under several frameworks, so claiming “no personal data” is not a defensible vendor statement on a lender’s behalf. Third, model decay: fraudsters adapt faster than borrowers change spending habits, so fraud-facing signals require revalidation on a shorter cycle than credit-facing ones.

For risk teams, alternative data has moved from the edge of the portfolio to the middle of it. It no longer only addresses applicants the scorecard cannot reach; it helps determine whether the scorecard still works for the applicants it can reach. Two practical implications follow: any new source must prove incremental lift against the model already in production, because adding data that merely correlates with existing inputs increases latency without improving performance; and the durable advantage belongs to lenders that can operate on data carrying no direct identifier, since that is the category both regulators and platform owners are least able to restrict simultaneously.

Related posts

AI momentum in China: dense innovation on show at WAIC 2026

David Jones

What new AI models will Sarvam Epoch launch on July 30?

David Jones

China’s AI showcase: humanoid robots, robot police and a roof collapse

Emily Brown

This website uses cookies to improve your experience. We assume you agree, but you can opt out if you wish. Accept More Info

Privacy & Cookies Policy