Skip to content

Reference

Locales

This reference lists the supported locales and their language analysis features.

A locale is identified by its BCP 47 tag. Locales are used in the following configurations:

en is the default locale when no locale is specified.

The following table lists the supported language tags and their analysis capabilities:

TagLanguageStopwordsStemmingOwn segmentationCompound splitting
arArabicyesyes
bgBulgarianyesyes
bnBengaliyesyes
bsBosnianyesyes
caCatalanyesyes
ckbCentral Kurdish (Sorani)yesyes
csCzechyesyes
daDanishyesyesyes
deGermanyesyesyes
elGreekyesyes
enEnglishyesyes
esSpanishyesyes
etEstonianyesyes
euBasqueyesyes
faPersianyesyes
fiFinnishyesyesyes
frFrenchyesyes
gaIrishyesyes
glGalicianyesyes
guGujaratiyesyes
hiHindiyesyes
hrCroatianyesyes
huHungarianyesyes
hyArmenianyesyes
idIndonesianyesyes
isIcelandicyesyesyes
itItalianyesyes
jaJapaneseyesyesyes
knKannadayesyes
koKoreanyes
ltLithuanianyesyes
lvLatvianyesyes
mlMalayalamyesyes
mrMarathiyesyes
msMalayyesyes
nbNorwegian Bokmålyesyesyes
neNepaliyesyes
nlDutchyesyesyes
nnNorwegian Nynorskyesyesyes
noNorwegianyesyesyes
orOdiayesyes
paPunjabiyesyes
plPolishyesyes
ptPortugueseyesyes
roRomanianyesyes
ruRussianyesyes
skSlovakyesyes
slSlovenianyesyes
srSerbianyesyes
svSwedishyesyesyes
taTamilyesyes
teTeluguyesyes
thThaiyes
trTurkishyesyes
ukUkrainianyesyes
urUrduyesyes
viVietnameseyes
zhChineseyesyesyes
zh-HantChinese (Traditional)yesyesyes
  • Stopwords: The built-in matching analyzer chain applies stopwords for the locale. A custom chain applies them by specifying the locale on a stopwords component. Japanese and Korean drop grammatical parts of speech instead of using a separate stopword list.
  • Stemming: The built-in matching analyzer chain applies stemming for the locale. A custom chain applies stemming by specifying the locale on a stemming component. Stemming behavior varies by language:
    • Japanese reduces elongated final vowels in loanwords.
    • Chinese stems mixed Latin words.
    • Korean, Thai and Vietnamese do not have stemming rules because words do not inflect.
  • Own segmentation: Indicates languages that use a dictionary-based word segmenter instead of Unicode segmentation because words are written without spaces. Thai words are segmented using standard Unicode segmentation. Vietnamese writes a space between syllables rather than between words, so the engine indexes the syllables. The Vietnamese stopword list holds the syllables that are grammar on their own, and leaves out a syllable that is as often part of a content word.
  • Compound splitting: Indicates locales that include decompounding data (for example, searching for jakke matches regnjakke). Decompounding requires decompounding data on the node. For more information, see compound words.
  • Locale data: Icelandic reads its stopwords, stemming, and compound parts from the locale data directory rather than from components built into the engine. A node without the data reports Icelandic as unsupported.
  • Normalization: Applied automatically when a language requires rules beyond Unicode case folding. Normalization covers the following cases:
    • Turkish dotless ı.
    • Greek accents.
    • Elided articles in Catalan, French, Irish, and Italian.
    • Distinct Unicode forms of letters in Arabic, Indic, and Cyrillic scripts.
    • The Arabic and Persian forms of the yeh, kaf and heh in Urdu, folded onto Urdu’s own.
  • Rules of the engine’s own: Gujarati, Kannada, Malayalam, Marathi, Odia, Punjabi, Slovak, Slovenian, Urdu and Vietnamese read stopword lists the engine carries itself, because Lucene ships none for them. All but Vietnamese stem with a light stemmer of the engine’s own, which cuts the case, number and tense endings that are written onto a word, in the way Lucene’s Hindi and Czech stemmers do. Regular inflection is covered. A form that changes the stem itself matches only itself.
  • Shared rules: Standard forms of one language read through the same stopword list and stemmer. Malay (ms) uses the Indonesian rules. Bosnian (bs) and Croatian (hr) use the Serbian rules, which fold both scripts and the letters č, ć, đ, š and ž to plain Latin, and bring the ijekavian mlijeko and the ekavian mleko to one term. A search typed in one of the three spellings finds a value written in another.
  • Script rewriting: zh-Hant rewrites Traditional characters as their Simplified forms before the text is segmented, because the Chinese word model holds the Simplified forms only. A value indexed as zh-Hant produces the same terms as the same sentence written in Simplified and indexed as zh. Character positions are unchanged, so highlights point at the text as it was sent.

nb (Norwegian Bokmål) and nn (Norwegian Nynorsk) both resolve to no (Norwegian) when a field does not specify the narrower tag. A field configured with no matches a search for either nb or nn.

Chinese is split by script rather than by written form. zh-TW, zh-HK and zh-MO resolve to zh-Hant, because those regions write Traditional without stating it in the tag. A field that holds only zh still answers a search for them, because there is no closer variant to read. To index both scripts separately, configure the field with both zh and zh-Hant.

Other language tags match by dropping subtags that the available locales do not distinguish. For example, a search request specifying sv-SE matches a field configured with sv.

Case carries no meaning in a tag. zh-hant, ZH-HANT and zh-Hant all name the same locale, and a definition stores whichever spelling you send as the canonical one, zh-Hant. Reading the definition back returns the canonical spelling.

For information about fields configured with multiple locales, see Localize fields.

A search in user mode reads a number typed next to a comparative word of the search locale as a filter. For more information, see reading numbers and units.

Every locale reads a number written with a unit, but only the locales in the following table have comparative words:

LocaleBelowAt mostAboveAt leastRange
da Danishunder, mindre end, billigere endmax, maks, højst, op tilover, mere end, dyrere endmin, mindst, framellem … og, fra … til
de Germanunter, weniger als, billiger alsmax, maximal, höchstens, bis, bis zuüber, mehr als, teurer alsmin, mindestens, abzwischen … und, von … bis
en Englishunder, below, less than, cheaper thanmax, maximum, at most, up toover, above, more thanmin, minimum, at least, frombetween … and, from … to, … to …
es Spanishmenos de, por debajo demax, máximo, hasta, como máximomás de, por encima demin, mínimo, al menos, desde, a partir deentre … y, de … a, desde … a
fi Finnishalle, vähemmän kuinmax, enintään, korkeintaanyli, enemmän kuinmin, vähintäännone
fr Frenchmoins de, sousmax, maximum, au plus, jusqu’àplus de, au-dessus demin, minimum, au moins, à partir deentre … et, de … à
it Italianmeno di, sottomax, massimo, al massimo, fino apiù di, oltre, sopramin, minimo, almeno, datra … e, fra … e, da … a
nb, nn, no Norwegianunder, mindre enn, billigere ennmax, maks, høyst, opp tilover, mer enn, dyrere ennmin, minst, framellom … og, fra … til
nl Dutchonder, minder dan, goedkoper danmax, maximaal, hoogstens, totboven, meer dan, duurder danmin, minimaal, minstens, vanaftussen … en, van … tot
pt Portuguesemenos de, abaixo demax, máximo, no máximo, atémais de, acima demin, mínimo, pelo menos, desde, a partir deentre … e, de … a, desde … a
sv Swedishunder, mindre än, billigare änmax, högst, upp tillöver, mer än, dyrare änmin, minst, frånmellan … och, från … till

Collation uses International Components for Unicode (ICU), which supports all locales. A sort definition with "collation": "locale" sorts values according to the rules of the locale assigned to the value.

The supported languages table lists locales supported for text analysis during indexing. A locale is included only when the engine provides specific analysis rules for that language. Icelandic is supported only on a node that has its locale data installed.

The engine rejects unsupported locale tags. The following table lists the error codes returned when a tag is unsupported:

ErrorCondition
index:field:locales:locale_unsupportedA field definition specifies an unsupported locale tag in locales or defaultLocale.
index:field:analyzer:locale_unsupportedAn analysis chain specifies an unsupported locale tag.
index:locale_fallback:locale_unsupportedA fallback chain specifies an unsupported locale tag.
search:locale_unsupportedA search query specifies an unsupported locale tag.

The following table lists the error codes returned for index-level locale declarations:

ErrorCondition
index:locales:default_locale_requiredAn index definition specifies locales without defaultLocale.
index:field:locales:locale_unknownA field definition specifies a locale in only or defaultLocale that the index does not declare.
index:field:locales:default_not_in_onlyA field definition specifies an only list that does not contain the default locale of the field.
index:field:locales:list_with_declarationA field definition specifies a locales array on an index that declares locales.
index:field:locales:only_without_declarationA field definition specifies only on an index that does not declare locales.

Each definition records its required locales in its features as locale.<tag>. A node built without a locale rejects an index that uses that locale instead of indexing the text as English.

Exofind is built by Level Four AB and is available under the Apache License 2.0.