Onchain analytics in finance
A public blockchain hands a bank something no other market gives it: the complete transaction record, free, in real time. What it withholds is the only thing the bank needs to use it, which is who each party is. Every number an institution derives from a chain therefore rests on an attribution, and the attribution carries an error rate that the number does not display.
What follows is what is genuinely readable, how clustering works and where it fails, why two providers disagree about the same address, the figures a regulated firm derives, and where the method stops.
What a public chain shows and what it does not
Readable with certainty: every transaction, its amount, its timing, the addresses involved, the current balance of any address, and the complete history of both. No permission is needed and no vendor can withhold it. For a market where a bank would normally pay for position data and still not get it, that is a genuine difference.
Not readable at all: the identity behind an address, whether two addresses belong to one party, whether a transfer was a sale or a move between a holder's own wallets, and whether an address that holds an asset owns it or custodies it for someone else. Everything an institution actually wants to say, exposure to a counterparty, a client's holdings, a reserve's adequacy, lives on that second list and has to be inferred.
Address clustering and the error it carries
Clustering groups addresses that are probably controlled by the same party, and two heuristics do most of the work. Common input ownership says that addresses appearing together as inputs to one transaction are usually controlled by the same party, since the spender needed the keys to all of them. Change address detection identifies the output that returns the remainder to the sender, which links a new address to the same party.
The result is a hypothesis with a confidence level and not a fact, and the two failure directions have different consequences. Over-eager clustering merges unrelated actors, which puts one customer's transactions in another's file; a change address misidentified attaches a stranger's onward activity to a holder. Acting on a false grouping can misdirect an investigation or unfairly flag a customer, which is why a cluster should be treated as something to verify against off-chain evidence before anyone acts on it.
Entity labeling, and why two providers disagree
A label is the next inference on top of a cluster: this group of addresses is an exchange, a bridge contract, a payment processor, a sanctioned entity, a mixer. Labels come from sources that are not the chain. A study measuring one major provider against seized services found its attribution reliable as a lower bound, with accuracy for individual services ranging from about 25 to 95 percent and false positives under half a percent. The label is the provider's conclusion and not a chain fact. Sources include a provider's own test transactions to a service, public reporting, information from investigations, and customer or victim submissions, so the label quality is a function of a provider's research operation and not of its software.
That is why two reputable providers return different answers for the same address, and the difference is not a bug in either. They clustered with different thresholds, labeled from different source material, and updated at different times. For an institution the operational consequences are concrete: a single provider's answer is one opinion, a material decision about a customer deserves a second source, and a figure reported to a supervisor should name which provider and which date produced it.
The figures a regulated firm actually derives
Four, and each inherits the attribution error differently. Counterparty exposure: how much a firm is owed by or has at an entity, which needs the entity's addresses to be complete, since a missed address understates the exposure. Customer risk scoring: whether a client's funds arrived from a service the firm will not accept, which is where a false positive costs a customer relationship and a false negative costs a compliance failure.
Sanctions exposure: whether any holding or counterparty touches a listed address, which is a hard obligation and the one where the firm cannot rely on a provider's judgment alone. And reserve verification: whether an issuer or a platform actually holds what it claims, which only works when the addresses are genuinely the issuer's, and is covered on the proof of reserve page. In each case the honest practice is to report the figure with its method, because a number derived from a hypothesis is not the same kind of number as a custodian's statement.
How this differs from forensic tracing
The methods overlap and the purposes do not, which changes what counts as good enough. Forensic tracing follows specific funds in a specific case, usually after a theft or for law enforcement, and its output has to survive challenge in a proceeding, so an analyst works one path and documents every inference.
Onchain analytics for a financial firm is a population exercise: scoring every client and every counterparty continuously, where the method has to be consistent and automated and no human reviews each conclusion. That difference means a tracing-grade conclusion cannot be produced at analytics scale, and an analytics-grade score should not be presented as a finding about a person. The blockchain forensics page covers the investigative side and crypto asset tracing the recovery work.
Where the method stops: privacy tools and chain hopping
Mixers and CoinJoin break the common input ownership heuristic deliberately, by constructing transactions whose inputs belong to many parties, so the inference that produces most clustering gives a wrong answer by design. Privacy-focused chains hide amounts or parties at the protocol level, which removes the raw material. Coin control and deliberate address hygiene by a sophisticated holder limit clustering without any special tool.
Chain hopping is the practical limit that affects ordinary cases more. Funds moving across a bridge to another chain, or through an exchange that pools customer deposits, break the trail because the thing arriving is a different asset on a different ledger with no cryptographic link to what left. Providers reconnect these hops with heuristics on timing and amount, and that reconnection is the weakest inference in the chain. A firm should know which of its conclusions depend on it.
How accurate is onchain analytics?
It depends entirely on which layer is being asked about, and conflating them is the most common error. The transaction data is exact. Clustering on a transparent chain with naive usage is reliable enough to act on. Entity labels vary by provider and by how well researched that part of the ecosystem is. Reconstructing a path across chains or through a pooling service is the least reliable layer. A firm that treats all four as one number, as a single risk score invites it to, has no way of knowing which of its decisions rests on the weakest inference.
Does a bank need an analytics provider, or can it do this itself?
The clustering is reproducible in-house; the labels are not. A firm with the data engineering can build an index and run the standard heuristics, as the blockchain data indexing page sets out. What it cannot build alone is the label set, because that comes from continuous research into which addresses belong to which services, which is the actual product an analytics vendor sells. Most institutions therefore buy labels and screening and keep their own pipeline for the holdings and exposures they report, where they need the derivation to be their own.
Onchain analytics and Finance Loop
Finance Loop brings the compliance officers who act on these scores together with the analysts who produce them, in its Risk & Compliance track. Finance Loop keeps the subject on the agenda because the question that decides an outcome, how much confidence an attribution deserves, is one neither group answers alone.
Finance Loop is a professional network and has the goal of driving the adoption of emerging technologies in finance, such as AI, tokenization, stablecoins, and DeFi. Finance Loop helps its members build skills and personal networks in these fields: Investment & Digital Assets, Payments & Digital Money, Digital Infrastructure & Sovereignty, and Risk & Compliance.