Hotel Mapping: Property Entity Resolution Guide

Learn how metasearch systems resolve the same physical hotel across OTAs and providers using canonical property identities and matching signals.

Editorial information
Advertisement

Hotel mapping is the process of deciding whether records from different suppliers represent the same physical property.

One hotel may have several names across channels, so raw supplier IDs and strings cannot be treated as canonical identity.

Why provider IDs are not enough

Every provider owns its own identifier namespace. The same numeric ID can represent completely different properties in different systems. A metasearch platform therefore needs an internal canonical property entity.

Matching signals

Strong signals can include coordinates, address, phone, official website and known chain/property identifiers. Supporting signals include name similarity, city, district, postal code, star category and chain.

Combining multiple signals is safer than trusting one field.

Name normalization

Useful operations include lowercase conversion, diacritic normalization, punctuation cleanup and reducing the weight of generic terms such as hotel, resort or spa. Over-aggressive normalization can create false matches, so it should support rather than replace other signals.

Geographic distance

Coordinates are valuable but not sufficient on their own. Several hotels can exist inside the same resort complex, so distance should be combined with name and address evidence.

Confidence score

A confidence model is usually more useful than a binary decision. For example, high-confidence matches can be automatic, mid-confidence cases can enter a review queue and low-confidence records can remain unmatched.

False merge vs false split

A false merge combines different hotels, mixing prices and content. A false split leaves the same hotel as multiple entities, creating duplicates and incomplete channel coverage.

Both damage metasearch quality.

Manual mapping

Automation still needs an override path for rebrands, relocations, new properties and unusual hotel complexes. Manual mapping should be auditable, not an invisible one-off fix.

Data-model pattern

Keep canonical properties separate from provider mappings. A mapping record can store canonical property ID, provider, provider property ID, confidence, source and review timestamp. This design makes reprocessing and audit much easier.

Why mapping is foundational

If mapping is wrong, prices, availability, reviews, photos and content all attach to the wrong entity. Property identity is therefore one of the most important hidden foundations of hotel metasearch.

Build an evidence model for property mapping

Name similarity is not enough. Store evidence for each candidate:

text
name similarity
address similarity
geo distance
phone match
website/domain match
brand/chain
postal code
provider cross-reference
manual verification

Example confidence states:

text
EXACT_VERIFIED
HIGH_CONFIDENCE
AMBIGUOUS
REJECTED

Separate candidate generation from scoring

First narrow the candidate set, then score it:

text
provider property
  -> geo/name candidate generation
  -> feature scoring
  -> threshold
      -> auto-map
      -> manual review
      -> reject

This improves performance and explainability at large catalog scale.

Common failure modes

  • merging two hotels with the same name,
  • duplicate after rebrand,
  • wrong coordinates selecting a nearby property,
  • chain name treated as property name,
  • provider reusing an ID for another property,
  • manual override overwritten by batch mapping.

KPIs

  • mapped-catalog coverage,
  • auto-map acceptance,
  • manual-review queue,
  • remap rate,
  • duplicate canonical entities,
  • high-confidence reversal rate,
  • booking/price anomalies after mapping.

Mapping output should carry confidence + evidence + version, not only an ID.

Technical advisory

Planning a similar integration?

We can review requirements, feed/API design and the production approach with you.

Discuss your project →

Related content