← All docs

How Confident Is That Answer? Reading Confidence Bands, Data Quality, and Stability

Key takeaways

  • The confidence band (Low / Medium / High) grades how sharply one origin candidate stands out from the field — a measure of separation, never a probability of being correct.
  • A sole candidate always shows 100% share because share is a relative apportionment, not a probability. Four separate ceilings can hold that 100% winner at Medium or Low.
  • Data quality caps the band from above; the stability signal answers a different question — whether the ranking would survive more data.
  • Behind a non-definitive band sits a calibrated probability (IVAP) with an honest error bar. A Deep Scan suggestion only appears when more data could actually change the answer.

Every origin result carries a one-word verdict on how much to trust it: Low, Medium, or High. That band is the most-read and most-misread number on the screen. A wallet can show a single origin candidate with 100% of the share and still come back Low. A result can read Medium while every signal on the page looks strong.

The band is not hedging. It is a precise statement about how decisively the evidence separates one answer from the alternatives, computed after several conservative caps have had their say. Reading it correctly is the difference between citing a finding and overstating it.

This article decodes the caution machinery: what the band measures, why a perfect-looking share can be capped, how data quality and stability feed in, and where the calibrated probability behind the band comes from. It stays on the confidence output specifically — the two-axis attribution model it sits on top of is the subject of Who Controls a TRON Wallet?.

The band is a clarity signal

The band reports how cleanly the origin analysis can pick a top candidate out of the field. It is set only after the agreement model, the window-overlap discount, the data-quality ceiling, and the poisoning and mixer reductions have all run. A High band means the evidence points hard at one answer; the case is not thereby closed.

Confidence is computed independently for each axis. The origin axis — who created and funded the wallet — produces a global band plus a per-candidate band for each contender. The current-control axis — who can sign today — carries its own rule-derived confidence. There is no single merged number spanning both.

The three levels form a strict ladder, and every cap in the system can only move a result down it. Internally each ceiling and floor takes the more conservative of two bands, so the order they apply in does not matter — a result can be lowered many times but never raised back up. This is why stacking caps is safe: the worst applicable verdict wins.

Starting band CREATOR ≥50% → HIGH ELSE AGREEMENT + IVAP STARTS PER EVIDENCE HIGH MEDIUM LOW CAPS · LOWER ONLY POISONING · MIXER DATA QUALITY LOW SPREAD SPARSE HISTORY Each cap takes the more conservative band — it can fall, never rise.
The band starts from the evidence and only ever falls: every cap takes the more conservative level.

One path skips the ladder. When the top candidate sent the account’s AccountCreateContract — the on-chain creation transaction — and holds at least half the share, the band is High immediately. Account creation is proof rather than a calibrated guess, so it does not wait on the probability model. That head start is not immunity, though. The conservative ceilings still run afterward: a creator seen on fewer than three transactions is capped back to Low, and detected poisoning or mixer contact steps a definitive candidate’s band down the same as any other. What creation does skip is the discount that exists only to penalize weak corroboration on a young account — because creation is definitive at any account age.

Why a 100% match can be Low confidence

The share percentage is the single most misleading figure in a result if read as a probability. It is a relative apportionment — a min-max normalization across the candidate field. A wallet with one origin candidate always shows that candidate at 100%, regardless of how thin the underlying evidence is. The absolute, cross-analysis-comparable quantity is the calibrated probability discussed below — the share cannot serve that purpose.

Four independent ceilings can hold a 100%-share winner well below High:

CeilingFires whenCaps at
Insufficient dataFewer than 3 transactions to work fromLow
Limited data3–9 transactions with no creator or delegation anchorMedium
Weak absolute strengthThe winner’s raw score is below 30Medium
Low spreadThe raw margin between the top two candidates is under 10 pointsMedium

The low-spread ceiling is the subtle one. Min-max normalization stretches whatever raw lead exists across the full 0–100 range, so a candidate that beat the runner-up by a hair of raw score can still be shown at 100% share against 0%. The low-spread check looks past the normalized share at the raw merged-score margin: under 10 points apart, the 100-versus-0 split is an artifact of the math rather than a decisive result, and the winner cannot read High. A sole candidate has no runner-up to compare against, so it is never flagged low-spread — its caution comes from the data-quality and absolute-strength ceilings instead.

There is no single “lonely wallet” flag. A sole non-definitive candidate is handled by the stack working together: it earns 100% share, but with nothing to corroborate it the agreement model drops it to Low, the absolute-strength floor holds it there, and on a young account the window-overlap discount keeps it down. The stability signal then spells the situation out in words — “Single candidate with limited supporting signals.”

Data quality sets a ceiling

Before any band is chosen, the engine grades how much material it had to reason from. The grade is a function of three things — transaction count, whether a creator was found, and whether resource delegation was seen — and it resolves to one of four levels:

LevelConditionEffect on the band
InsufficientFewer than 3 transactionsCapped at Low
Limited3–9 transactions, no creator or delegationCapped at Medium
AdequateHas enough to reason from, short of RichNo cap
Rich30+ transactions, or a creator, or a delegationNo cap

Data quality never adds or subtracts points from a candidate’s score. It only sets a ceiling on the confidence band. A definitive creation signal seen on a two-transaction wallet is still capped at Low — the signal is real, but the surrounding record is too thin to trust the reading around it.

The grade is measured against the wallet’s true on-chain size rather than the sample the analysis happened to read. A wallet with thousands of transactions that was analyzed through a bounded sample is still graded Rich by its real footprint, so a busy address is never mislabeled Insufficient because only part of its history was pulled. TronGrid’s account-transaction endpoint returns at most 200 records per page, which is the structural reason a bounded fetch is a sample rather than the whole story for high-volume wallets — and the reason coverage matters to the reading (below).

The 3 / 10 / 30 thresholds are not guesses. They were re-audited against the analysis corpus in mid-2026; the review demoted only three correct results (all from High to Medium, all on wallets with fewer than ten transactions) and surfaced no over-confident wrong calls, so the bins were kept as-is.

Stability: whether the ranking would hold

Confidence and stability answer different questions, and reading them as the same thing is a common error. The band measures how sharply the evidence picks a winner. Stability measures how likely that winner is to change if more data arrived. Stability is reported as HIGH, MODERATE, or LOW, and it is informational — it never moves a score or a band.

The level comes from the gap between the top two candidates, read against how complete the history was:

  • A definitive account creator is always HIGH — creation is on-chain proof and more data cannot overturn it.
  • A commanding lead (a large gap in both raw score and share) is HIGH.
  • A clear lead is HIGH when the whole history was analyzed, but only MODERATE when the history was truncated — because the missing data could still reshuffle the order.
  • A modest lead is MODERATE.
  • A close race is LOW, and a deep scan is recommended.

Whether the history counts as complete is resolved through the data source. When the analysis reads a wallet’s full on-chain history, completeness is known exactly, and a truncation is only recorded when the record genuinely was cut off. When the source cannot prove completeness, the engine falls back to a conservative proxy — treating a fetch that hit its limit as potentially truncated. For stability’s purposes, account creation is the only signal treated as definitive; resource delegation, though strong origin evidence, rides the ordinary margin logic and does not by itself lock a ranking. Delegation’s role in ownership is the subject of Who Pays the Fees?.

The calibrated probability behind the band

For results that are not settled by a definitive creation signal, the band is anchored to a real probability: the model’s calibrated estimate that the top candidate is the true activator. That estimate is produced by an Inductive Venn-Abers predictor — a published calibration method whose probability outputs are provably well-calibrated when the data behaves like the set it was fit on — trained over a 207-case full-history calibration set.

Venn-Abers returns not a single number but a bracket: a lower and an upper probability, whose midpoint is the point estimate. That bracket is the honest error bar. The band then reads the midpoint against two anchored cutoffs — a midpoint at or above 0.80 supports High, at or above 0.50 supports Medium, and below that is Low. In the current production calibration, those cutoffs land at a raw score of roughly 86 for the High crossing and 59 for the Medium crossing.

The calibrated figure is surfaced on the result as its own field, separate from the band. As the FAQ puts it, the probability is an input to the band rather than the band itself — the band is a coarse label, while the probability expresses the model’s certainty more precisely, and the two can differ. The figure is computed for every analysis where calibration is enabled, including definitive creators whose band never needs to consult it.

One guardrail rides this path. The calibration set is drawn from busy, established wallets, so the score-to-probability map is not validated on genuinely thin histories. A non-definitive result on a wallet with 20 or fewer real transactions — or one whose true size cannot be established — is capped at Medium regardless of what the probability suggests. The model does not get to be confident in a regime it was never tested on.

A Low band is not a failure to answer. It is an honest statement that the evidence does not support a confident one.

When a Deep Scan is worth running

A Deep Scan pulls a wider slice of history — up to 200 inbound and 50 outbound transactions, against the standard 30 and 10 — and caches its result separately from the standard read. The interface only suggests one when it could actually help, which means three conditions all hold: there is no definitive creator, the analyzed history was truncated, and stability is not already HIGH.

Each condition is a reason a deeper pull would be wasted otherwise. A definitive creator cannot be overturned by more data. A history already known to be complete has no more data to find. A ranking that is already stable will not move. Strip any one of those away and the suggestion is suppressed.

This is why the suggestion behaves differently once a wallet’s full history is available: a provably complete record is never marked truncated, so the Deep Scan prompt self-suppresses on fully-covered wallets and survives only where the analyzed set is genuinely a sample. When you see the prompt, it is a signal that the current answer rests on partial data, and a reason to pull more before relying on it.

Reading the band as an investigator

Translate the band into claims you can defend:

  • High — the evidence strongly supports the identified candidate. It is the strongest reading the engine offers, and it still is not a closed case.
  • Medium — the most probable candidate given the available data, worth stating with that qualifier attached.
  • Low — honest abstention. Presenting a Low result as a definitive answer overstates what the evidence supports.

A few cross-readings sharpen the picture. A High origin result paired with a current-control finding that points at a different address is a strong hint at a custody transfer or wallet sale — the account was created and funded by one party and is signed by another today. And an unexpectedly Low band with no visible cause is itself a signal: an undetected mixer can leave a result unattributable in fact while the analysis reports Medium or Low without flagging mixer involvement, so a confidence reading that is low for no apparent reason is worth a manual look. Poisoning and mixer contact, when detected, each step the band down directly — mechanics covered in Address Poisoning.

The band is a compression of everything above into one word. Read it as a measure of separation, cross-check it against stability and data quality, and treat the calibrated probability as the more precise figure when the two disagree. That is how the caution machinery is meant to be used — as a guide to how far a finding will carry, not a verdict to quote unread.

Sources