Contents Menu Expand Light mode Dark mode Auto light/dark, in light mode Auto light/dark, in dark mode Skip to content
Documentation for World Historical Gazetteer Latest Release
Logo
Documentation for World Historical Gazetteer Latest Release
  • Introduction
  • Guides & Tutorials
    • 1. Our Indexes
    • 2. Workbench
    • 3. Publishing Data
    • 4. Uploading Data
    • 5. Reconciliation & Accessioning
    • 6. Reviewing accessioning results
    • 7. Collection Groups
  • WHG Staff
    • 1. Gazetteer Configurator
    • 2. User API Profiles
    • 3. Site smoke test
    • 4. Atlas (staff notes)
  • Technical
    • 1. Repositories
    • 2. APIs
    • 3. Issues
  • Development Roadmap
    • 1. v4.0: Atlas, Collaborative Workbench, Phonetics
      • 1.1.1. Draft User Guide
        • 1.1.1.1. Getting Started
        • 1.1.1.2. Atlas: Exploring Places
        • 1.1.1.3. Map your Data
        • 1.1.1.4. Collaborative Collections
        • 1.1.1.5. Suggesting Corrections
        • 1.1.1.6. Sounds-alike Search
        • 1.1.1.7. Identifiers and Citation
        • 1.1.1.8. Data Formats In and Out
      • 1.1.2. Beta Testing Plan
      • 1.1.3. Data Model
        • 1.1.3.1. Introduction
        • 1.1.3.2. Overview
        • 1.1.3.3. Attestations & Relations
        • 1.1.3.4. Vocabularies
        • 1.1.3.5. Special SpatialEntity Patterns
        • 1.1.3.6. Contribution Types & Data Formats
        • 1.1.3.7. RDF Representation
        • 1.1.3.8. Platform Use Cases
        • 1.1.3.9. Summary & Future Directions
      • Toponym Phonetics: Technical Design
        • 1. Overview
        • 2. Elastic Management Guide
        • 3. Infrastructure Summary
        • 4. Components
        • 5. Data Flow
        • 6. Elasticsearch Index Design
        • 7. Training the Phonetic Similarity Model
        • 8. Query Pipeline
        • 9. Advantages of This Architecture
        • 10. Monitoring & Observability
        • 11. Future Extensions
        • 12. Deployment Plan
        • 13. Risk Assessment
        • 14. Success Criteria
        • 15. Summary & References
      • 1.1.4. System Architecture
        • 1.1.4.1. Database Technology Assessment
    • 2. v4.1: Open Educational Resources
    • 3. Archive: the 2025 v4 Design
      • 3.1. User Guide (2025 design)
        • 3.1.1. Quick Start Guide
        • 3.1.2. Understanding WHG Concepts
        • 3.1.3. Place Record Anatomy
        • 3.1.4. Contributing Data Overview
        • 3.1.5. Reconciliation Overview
        • 3.1.6. Tutorial: Creating a Historical Route
        • 3.1.7. Frequently Asked Questions
        • 3.1.8. Glossary
      • 3.2. Implementation in ArangoDB
      • 3.3. Kubernetes Configuration
      • 3.4. SSH Key Setup
      • 3.5. Deploying the Management Pod
      • 3.6. Deploying Services
      • 3.7. Service Configuration
  • License
Back to top
View this page
Edit this page

1.1.1.6. Sounds-alike Search¶

v4.0-beta

Note

Part of the Draft v4 User Guide: under review during beta testing.

Historical place names reach us spelled in many ways and written in many scripts. München, Munich and Мюнхен are one name, heard three ways. A search that only compares letters misses these connections. WHG also compares how names sound.

1.1.1.6.1. What it does¶

When you search, or when Map your Data looks for matches, WHG considers names that sound alike even when they are spelled differently or written in a different script. Results are ranked by how close the sounds are, rather than simply matched or not matched.

It helps most with:

  • historical and variant spellings of the same name;

  • transliterations of one name into different alphabets;

  • cross-script matches, for example Latin and Cyrillic, Greek or Arabic forms.

It does not treat different names for one place as sound-alikes: Deutschland, Germany and Allemagne are different names, and they are connected through the records’ other evidence (location, type, source links), not through sound.

1.1.1.6.2. How it works, in brief¶

Each distinct name is converted into a description of how it is pronounced. Language-specific rules turn written letters into sounds, and a model trained on place names from many languages turns those sounds into a numerical “fingerprint”. Names whose fingerprints are close sound alike. The same method runs on WHG’s servers and inside Map your Data in your browser, so both give the same answers.

For the technical design, see Toponym Phonetics.

1.1.1.6.3. Its limits¶

Sounds-alike search is a discovery aid, not a verdict. It can:

  • miss a variant that looks unlike anything it has learned from;

  • work less well for scripts and languages that are poorly represented in its training data;

  • rest on sound rules that are imperfect for some languages.

Always check a suggested match against the record’s other evidence.

1.1.1.6.4. Help improve it¶

The letter-to-sound rules for each language can be wrong, incomplete, or missing letters the language uses. WHG provides a page where people who know a language can check its rules against real place names from the index, say what is wrong, and propose a correction. Reading it is open to everyone. Contributing needs an account, and contributions are dedicated to the public domain (CC0), so they can be used by anyone, including the open-source project the rules come from.

Copyright ©2017–2026 World Historical Gazetteer
Last updated on 28 September 2026
On this page
  • 1.1.1.6. Sounds-alike Search
    • 1.1.1.6.1. What it does
    • 1.1.1.6.2. How it works, in brief
    • 1.1.1.6.3. Its limits
    • 1.1.1.6.4. Help improve it