The same facility, arriving under four different names.

A permit filing, a shipping manifest, a market notice and an orbital pass can all describe one asset. Deciding that they do - and being able to show why - is what turns sources into an intelligence layer.

Nothing in the physical record agrees on a name.

One site appears in a permit under its legal owner, in a grid queue under a project code, in a manifest under a shipper abbreviation, and on its operator's own page under a brand. Subsidiaries file separately from parents. Assets are sold and keep trading under the old name for years. None of it is anyone's fault, and none of it joins on a string match.

The consequence is specific rather than cosmetic. Count the suppliers able to deliver a component without resolving them and you get four where there is one, which reads as a healthy market. Resolve them and it is a sole source. The same error runs the other way: two genuinely different sites collapsed into one record invents a capacity that neither of them holds.

Fuzzy name matching is not signal resolution.

String similarity gets the easy cases and is confidently wrong about the ones that matter, because the hard pairs here are similar names belonging to different legal entities, and dissimilar names belonging to the same one. A model that only reads the name has nothing to break the tie with.

Ours reads the evidence around the name: registration identifiers, coordinates and facility footprints, the class of work a permit covers, corporate ownership where it is on the record, orbital observation of the site itself, and what each source claims the asset does. A match is a decision made from several agreeing signals, and a record that only agrees on its name does not clear the bar. Where the evidence genuinely will not decide, the records stay separate - an over-merged asset is a false capacity claim, which is the more expensive of the two errors.

Four stages, continuously.

This is not a one-time cleanup that produces a master list. Sources move underneath the layer by the second, so resolution runs against them continuously and every stage can revise what the last one decided.

  1. Normalize

    Each incoming record is parsed into the same shape: identifiers, names, coordinates, facilities, capacities and claimed activity, with the source and the moment it was observed kept alongside.

  2. Generate candidates

    Cheap signals narrow billions of possible pairs to a reviewable set. This stage is deliberately generous: a pair never generated can never be matched.

  3. Decide

    A model scores each candidate pair on the agreeing and disagreeing evidence, not on the name. Confident matches merge, confident non-matches are recorded so they are not re-litigated, and the ambiguous middle stays separate.

  4. Attach

    The resolved entity inherits every relationship its source records held, which is the moment a pile of documents becomes something a query can traverse.

A merge you cannot inspect is a merge you cannot trade on.

Every resolved entity keeps the records it was built from. Each property points back at the source that asserted it and the moment that source last did, so a reviewer can see that a coordinate came from a filing, a throughput figure came from a licensed dataset, and a construction stage came from an orbital pass - three very different weights of evidence.

That is a requirement, not a feature. Decisions made on this data get audited by risk committees and by regulators, and an answer whose only justification is that a model produced it does not survive the audit. It also gives the resolution itself a way to be corrected: when a merge is wrong, the underlying records are still there to be split, and the correction propagates through the Console rather than being pasted over one screen.

One asset, everything each source knew about it.

The point of resolving is not tidiness. It is that a property discovered in one source becomes usable in a question asked of another.

Identity

Legal owner, operating names, registration identifiers, and the corporate parent where it is on the record.

Location

Coordinates and footprint, because a permit, a queue position and a satellite pass all resolve against a place rather than a letterhead.

Capacity

Rated and observed throughput, separated from what an operator advertises, with the source of each.

Power

Contracted and available supply, queue position, and the load actually drawn against it.

Movement history

Shipments and flows in and out of the asset, tied to the routes and carriers that moved them.

Provenance

For every property above: which source asserted it, and the moment that source last did.

Evidence that keeps pace with the decision.

Evaluate Nuclir against the systems, markets, and decisions that matter to your organization. The result is current intelligence with the context needed to use it responsibly.