Key takeaways
- Use deterministic identifiers before fuzzy similarity.
- Treat over-merging as a higher-risk error than leaving a possible duplicate unresolved.
- Retain match evidence so every link can be reviewed and reversed.
Matching is more than deduplication
Two records can describe the same company while using different names, domains or addresses. They can also look similar while representing separate branches, legal entities or unrelated firms.
Entity resolution turns those ambiguous records into a relationship decision. The goal is not to maximize the number of matches; it is to create links that are defensible for the business use case.
- Define whether the target is legal entity, brand, location or account.
- Separate duplicate detection from parent-child mapping.
- Document the consequences of a false merge.
Start with high-confidence identifiers
Exact matches on a verified domain, registration number or trusted external ID are stronger than a name similarity score. Normalize those fields first, then use them as anchors for later comparisons.
When an identifier is shared by multiple branches, add scope such as country, address or entity type. A key is only useful when its uniqueness assumptions are explicit.
- Normalize domains and legal suffixes.
- Keep country and source context with identifiers.
- Reject conflicting high-confidence keys.
Build evidence tiers instead of one magic score
A match decision should explain which fields agree. A normalized name plus matching domain may be enough for an automatic link, while a similar name alone should remain a candidate.
Use tiers such as confirmed, probable, possible and rejected. Each tier can drive a different action: merge, queue for review, request more evidence or keep records separate.
- Score evidence by reliability, not just similarity.
- Make negative evidence explicit.
- Give every decision a reason code.
Protect against over-merging
Over-merging contaminates every downstream report because activity from two businesses becomes one history. It is usually harder to detect and repair than an unresolved duplicate.
Apply conservative thresholds when the records have conflicting addresses, domains, countries or industries. Preserve the original records and represent the proposed link separately until it is confirmed.
- Prefer unresolved to an unsupported merge.
- Flag contradictory fields for review.
- Keep reversible relationship records.
Design a useful review queue
A review queue should show the fields that drove the suggestion, the fields that disagree and the action that a reviewer can take. Avoid making reviewers open five systems just to understand one match.
Capture the reviewer decision and reason. Those decisions can improve thresholds and reveal source-specific normalization rules without silently changing historical matches.
- Show side-by-side evidence.
- Support confirm, reject and defer actions.
- Measure agreement and turnaround time.
Measure resolution quality
Track precision and recall on a labeled sample, but also monitor business impact. A small number of false merges may be more damaging than many missed duplicates in a sales or compliance workflow.
Recheck a sample after every source change. Entity resolution is a living system because names, domains, ownership and source formatting all change over time.
- Maintain a reviewed golden set.
- Report false-merge and missed-match rates separately.
- Version rules and rerun historical samples.
Frequently asked questions
Should I use fuzzy matching for company names?
Yes, as candidate generation or supporting evidence. Do not let name similarity alone create an irreversible merge when stronger fields conflict.
What is the safest default when two records are ambiguous?
Keep them separate, store the candidate relationship and route the evidence to a review queue.
Can entity resolution link branches to a parent company?
Yes, but model parent-child relationships separately from duplicate records so branch activity is not accidentally collapsed.