Accents are not properly configured for the french language, I updated the related config files:

config/optional/search_api_solr.solr_field_type.text_fr_6_0_0.yml
config/optional/search_api_solr.solr_field_type.text_fr_7_0_0.yml

Comments

B2F created an issue. See original summary.

b2f’s picture

Updated supporting organization

sylvainm’s picture

+1

mkalkbrenner’s picture

Status: Needs review » Needs work
+++ b/config/optional/search_api_solr.solr_field_type.text_fr_7_0_0.yml
@@ -405,166 +405,166 @@ text_files:
     # é => e
-    #"\u00E9" => "e"
+    "\u00E9" => "e"
     # ê => e

This patch will remove all special characters for French!

We wanted exactly the opposite to distinguish between the accents.

Could you describe the issue you have?

b2f’s picture

Indeed it could be useful that accents are removed so that when a user type the unaccented character in a fulltext input, it will give the same results regardless.

For instance on a website we are working on, if I'm searching "endometriose", I should find "endométriose" in the results. It seem to be a common need.

Cheers.

b2f’s picture

Actually the client request is especially true for the opposite, when we type the "é" character in the fulltext search, we want to find the unaccented "e" in the results.

mkalkbrenner’s picture

Category: Bug report » Support request
Status: Needs work » Fixed

It's impossible to define a configuration that fits all needs.
But the good thing is, that you can simply leverage drupal's config management to adjust the settings to your customer's requirement without a patch!
I suggest that you create your own "domain" and adjust the accents there. Have a look at the documentation:
https://www.drupal.org/docs/8/modules/search-api-solr/search-api-solr-ho...

b2f’s picture

Thanks for the tip, I didn't know about that feature.

DeFr’s picture

Title: Update text_fr accents in solr_field_type » Improve french solr_field_type
Category: Support request » Feature request
Status: Fixed » Needs work

Sorry for re-opening this, but please read the following. While I'm pretty sure everyone agree that providing a single config that will work for all French users isn't possible, I still think that there's a few stuff that can be tweaked to better fit the needs of the majority of the french users out there.

Note : everything below is specific to French ; it's the only language for which I've actually investigated what the various stemmers algorithm provided by Solr do, and it's the only one I can claim I've implemented on a lot of sites targeted toward different audience. Some might apply to other language too, but I can't guarantee it.

First of all, I'm kinda surprised about the statement in #4 , "We wanted exactly the opposite to distinguish between the accents." . Do you have some links to public discussion about that ? Because (spoiler alert), that's not what's going on with the current configuration, more on that later.

As is, the configuration shipped for the french language feelds weird :

  • First one is easy enough : accents_fr as is isn't consistent : some french characters ( for example œ, æ ) are transliterated ; some non french characters ( ì, ý ) aren't.
  • Unless you're on a site that mixes language, transliterating characters from other scripts language isn't the correct thing to do. Virtually no French users will expect to find documents with words starting with ass for the search aß*, or to find words with th for a search for þ. In most case for French users, those would be better left alone.
  • Applying a French stemmer algorithm to the transliterated version of foreign characters is pretty much guaranteed to produce unexpected results
  • Comment #4 states that the goal is to distinguish between the accent, but the config use the SnowballPorter stemmer (algorithm here) . The step 6 of that stemmer is to explicitely remove the accents that might be at the end of the words, because they're all from the same stem. Once that final accent is removed and taking into account the stopwords provided by default, there's pretty much no word left in the french language that would vary only by accent, so I'm a bit confused ; you generally want to remove all the remaining accent because French speakers are rather bad at correctly entering accented letter when they should, so this is seen as another kind of stemming step. That being said, the mis-understanding might have been that Snowball Porter stemmer for French does expect to see correctly accentuated letter in the provided tokens to correctly remove the suffix, so you're using that stemmer you shouldn't remove them with a CharFilter, but instead use another filter after the tokenizer. To avoid that and to slightly alleviate what some people see as a flaw in snowball porter (because suffix are correctly removed only when end users enter correct accents in your search form too), you can use the dedicated FrenchLightStemFilterFactory (original algorithm here, lucene solr implementation ) : that one will remove all accents and not only the last one, and does it slightly earlier in the algorithm

All of this being said, I'm not completely sure what the best way forward is. A few options I can think of

- If we assume that end users pretty much never input accents correctly, then @B2F patch above makes sense : it means that there's a few stemming case in Snowball Porter that won't be it, but it'll be consistent between the query and indexing, and thus make end users find what they're looking for, correctly written or not
- Another approach would be to switch to FrenchLightStemFilterFactory

At the very least, I think accents_fr should be made consistent.

mkalkbrenner’s picture

Obviously languages are different ;-)

In German it is a big mistake to transliterate an Umlaut. Examples:

Küchen => kitchens
Kuchen => cake

Ich fahre mit der Fähre. => I "take" the ferry.

I don't know if there're similar examples for French or not.

The discussions about the default configurations for several languages happened at different Drupal Cons with people from different countries.
And I always encourage people to contribute their configs.

So I'll accept a patch to change the French default field type. But the patch has to be complete including the upgrade path.

mkalkbrenner’s picture

I can at least help with the update.

b2f’s picture

Hello Markus,

How should I go about making a better patch ? Thanks.

mkalkbrenner’s picture

For existing installations the new configs don't get magically applied.
You need to implement an update hook. Have a look at search_api_solr_update_8311().

b2f’s picture

Assigned: Unassigned » b2f
gonssal’s picture

I just wanted to chim in to say that in both Spanish and Catalan, by default it should also be expected to transliterate all the accents. Except in some very rare cases, there's no change of meaning for the same word with accent (tilde) and without it and, even in those cases, it would generally be fine to get all the results.

I'm not a french speaker but I think in the provided patch, the Ç => C and ç => c conversions should be left commented. For example plaçage is not the same as placage. This is indeed complex stuff.

Also I don't think the tweaked configs should be automatically applied on existing installations, just changed when a new config.zip is generated and the tweaks documented in the release notes, so people don't get sudden unexpected behaviour changes. In other words, I don't think an upgrade path should be included.

mkalkbrenner’s picture

Gondel, please Open a dedicated issue.
Fun fact, the current config has been developed in Barcelona at DrupalCon. But maybe some didn’t understand the “nature” of search and followed the German reference too strictly.

gonssal’s picture

@mkalkbrenner I will open a new issue, but will it be ok if there's no upgrade path as I suggested? I don't think automatically upgrading configs without user action is a good idea.

b2f’s picture

Status: Needs work » Needs review

I think the update is indeed questionable, the provided example above in search_api_solr_update_8311 is adding new values not replacing existing ones.

mkalkbrenner’s picture

Status: Needs review » Needs work

OK, I agree on your suggestion to skip the upgrade path. But the patch should be extended by comments added to the yml files why these characters are normalized. maybe including a link to this issue.

mkalkbrenner’s picture

Status: Needs work » Reviewed & tested by the community

  • mkalkbrenner committed 076e7c9 on 8.x-3.x authored by B2F
    Issue #3096210 by B2F, mkalkbrenner, gonssal, SylvainM, DeFr: Improve...
mkalkbrenner’s picture

Assigned: b2f » Unassigned
Status: Reviewed & tested by the community » Fixed
b2f’s picture

Status: Fixed » Closed (fixed)

Automatically closed - issue fixed for 2 weeks with no activity.