Hotel Mapping: Property Entity Resolution Guide
Learn how metasearch systems resolve the same physical hotel across OTAs and providers using canonical property identities and matching signals.
Hotel mapping is the process of deciding whether records from different suppliers represent the same physical property.
One hotel may have several names across channels, so raw supplier IDs and strings cannot be treated as canonical identity.
Why provider IDs are not enough
Every provider owns its own identifier namespace. The same numeric ID can represent completely different properties in different systems. A metasearch platform therefore needs an internal canonical property entity.
Matching signals
Strong signals can include coordinates, address, phone, official website and known chain/property identifiers. Supporting signals include name similarity, city, district, postal code, star category and chain.
Combining multiple signals is safer than trusting one field.
Name normalization
Useful operations include lowercase conversion, diacritic normalization, punctuation cleanup and reducing the weight of generic terms such as hotel, resort or spa. Over-aggressive normalization can create false matches, so it should support rather than replace other signals.
Geographic distance
Coordinates are valuable but not sufficient on their own. Several hotels can exist inside the same resort complex, so distance should be combined with name and address evidence.
Confidence score
A confidence model is usually more useful than a binary decision. For example, high-confidence matches can be automatic, mid-confidence cases can enter a review queue and low-confidence records can remain unmatched.
False merge vs false split
A false merge combines different hotels, mixing prices and content. A false split leaves the same hotel as multiple entities, creating duplicates and incomplete channel coverage.
Both damage metasearch quality.
Manual mapping
Automation still needs an override path for rebrands, relocations, new properties and unusual hotel complexes. Manual mapping should be auditable, not an invisible one-off fix.
Data-model pattern
Keep canonical properties separate from provider mappings. A mapping record can store canonical property ID, provider, provider property ID, confidence, source and review timestamp. This design makes reprocessing and audit much easier.
Why mapping is foundational
If mapping is wrong, prices, availability, reviews, photos and content all attach to the wrong entity. Property identity is therefore one of the most important hidden foundations of hotel metasearch.
Build an evidence model for property mapping
Name similarity is not enough. Store evidence for each candidate:
name similarity
address similarity
geo distance
phone match
website/domain match
brand/chain
postal code
provider cross-reference
manual verificationExample confidence states:
EXACT_VERIFIED
HIGH_CONFIDENCE
AMBIGUOUS
REJECTEDSeparate candidate generation from scoring
First narrow the candidate set, then score it:
provider property
-> geo/name candidate generation
-> feature scoring
-> threshold
-> auto-map
-> manual review
-> rejectThis improves performance and explainability at large catalog scale.
Common failure modes
- merging two hotels with the same name,
- duplicate after rebrand,
- wrong coordinates selecting a nearby property,
- chain name treated as property name,
- provider reusing an ID for another property,
- manual override overwritten by batch mapping.
KPIs
- mapped-catalog coverage,
- auto-map acceptance,
- manual-review queue,
- remap rate,
- duplicate canonical entities,
- high-confidence reversal rate,
- booking/price anomalies after mapping.
Mapping output should carry confidence + evidence + version, not only an ID.
Planning a similar integration?
We can review requirements, feed/API design and the production approach with you.