Trying to use the ASCIIFoldingFilterFactory to eliminate diacritical marks in the analysis chain and search query with no results. Drupal field used is body with type fulltext. The ASCIIFoldingFilterFactory is even added to every field type in the schema file with no results.

CommentFileSizeAuthor
#18 3003524.patch7.58 KBmkalkbrenner
#9 3.png86.12 KBleoperezpulido
#9 2.png86.87 KBleoperezpulido
#9 1.png90.4 KBleoperezpulido

Comments

leoperezpulido created an issue. See original summary.

mkalkbrenner’s picture

Issue tags: -search api solr, -solr, -schema, -filter

Can you provide some more details about your setup?
Solr version, cloud or single core, ...
Are you sure that Solr uses your schema?

leoperezpulido’s picture

Yes. It is a local setup. Solr version 7.4.0, single core, default managed-schema. PHP version 7.2.10-0ubuntu0.18.04.1. Working with Search API 8.x.1-10, and Search API Pages 8.x.1.0-alpha12, along with Search API Autocomplete 8.x.1.0.
When query is introduced inside search box, the autocomplete shows the title of the node where the term is located, but when results are returned from the query nothing is rendered on the results page. No fragment with highlighted term with diacritical mark, only the title where the term is located.
Solr can't start up without a schema file and I have one single managed-schema file inside conf folder.

leoperezpulido’s picture

Yes. It is a local setup. Solr version 7.4.0, single core, default schema.xml. PHP version 7.2.10-0ubuntu0.18.04.1. Working with Search API 8.x.1-10, and Search API Pages 8.x.1.0-alpha12, along with Search API Autocomplete 8.x.1.0.
When query is introduced inside search box, the autocomplete shows the title of the node where the term is located. But when results are returned, only the title of the node where the term is located is rendered on the results page, no fragment with highlighted term with diacritical mark.
Solr can't start up without a schema file and I have one single schema.xml file inside conf folder.

mkalkbrenner’s picture

What do you mean with "default schema"? search_api_solr requires a drupal schema.

search_api_solr 8.x-1.x doesn't support Solr 7.x. You need to use search_api_solr 2.x and follow the instructions in README.md.

I can't recommend search_api_pages. It's not up to date. You should use Views to display your search results.

leoperezpulido’s picture

By default schema I mean the schema file that comes when you download config.zip file from Search API > Server.

Let me upgrade to search_api_solr 8.x-2.2 and use Views to see if that solves the issue.

leoperezpulido’s picture

Brand new installation of Drupal 8.6.1 with installed search_api_solr 2.x. Try to render search results with the ASCIIFoldingFilterFactory applied to every field in the schema file (just in case). Fields to display: title and excerpt.

- Tested with a custom search View, get title field displayed but no fragment with keyword highlighted.
- Tested with Solr Search Defaults, get title field displayed but no fragment with keyword highlighted.

In every case the behavior is the same: to show the title of the node where the term is located in search results but not the fragment with the keyword highlighted. So that if, for example, I have only two documents indexed with the same content but one of them have diacritical marks, a query for a term on this documents will display the title fields of both documents but only one of them, the one without the diacritical marks, will be displayed as highlighted fragment.

mkalkbrenner’s picture

Version: 8.x-1.0 » 8.x-2.x-dev

OK, let me recap:
You have a working setup and you see highlighted search results if they only contain ASCII characters?
Is that correct?

So what's your expected behavior? Can you describe an example, what are your search keys, what do you expect to find and what should be highlighted?

Are you targeting a specific language and did you try the multilingual Solr backend?

leoperezpulido’s picture

StatusFileSize
new90.4 KB
new86.87 KB
new86.12 KB

I have two documents with two fields each: title and body. Content for the first is: Drupal is the leading open-source CMS. Content for the second is: Drûpal is the leading open-source CMS. Notice the diacritic in the second document. If you enter he query term: leading, everything works as expected [1]. Both documents return with the token highlighted.

Now I want to search for the term: drupal. Assuming that I have applied the ASCII filter, both documents would have to return with the tokens highlighted (as is the default behavior in Solr) but instead I get only one and a half because the document with the diacritic doesn't return the excerpt [2], only the title. To search for a token with a diacritic I need to type in the term with the accent, in this case: drûpal [3]. Please notice that also in this image I get both documents returned but this time the one without diacritics doesn't return the excerpt. That is not the expected behavior. Search best practices recommend to free users to enter query terms with diacritics and that is what the ASCII filter is for.

mkalkbrenner’s picture

leoperezpulido’s picture

Thank you anyway.

mkalkbrenner’s picture

Does the patch solve your issue?

leoperezpulido’s picture

No. Same behavior as before.

mkalkbrenner’s picture

Maybe we should move the discussion to https://drupalchat.me/channel/search
That would be easier.

leoperezpulido’s picture

Write to you there, then.

leoperezpulido’s picture

Now, the patch did solved the issue. The whole process boils down to:

  1. We don't need to apply any ASCIIFoldingFilterFactory into the schema.xml.
  2. We apply the patch.
  3. We need to add (/admin/config/regional/language) the language(s) we are going to work with before we do any other Multilingual Solr operation.
  4. Then, we can create our Solr Server with a Multilingual Solr backend, and remember to check (in Advanced options) "Retrieve Result Data From Solr" and "Highlight Retrieved Data".
  5. After the Solr Server is created, the config.zip must be downloaded from /admin/config/search/search-api/server/solr_multilingual/solr_field_type. This config.zip file have a schema.xml file with all the necessary field types for our multilingual search created automatically. Nice touch!
  6. After the Solr core is created, the highlighting processor (/admin/config/search/search-api/index/test_index/processors) must be enabled.
  7. The content we create must be made with the right language selected (/admin/config/regional/content-language). Remember to check the option "Show language selector on create and edit pages" to be able to do it.
  8. We can now index content and search.
  9. It all works!
mkalkbrenner’s picture

Status: Active » Fixed

Thanks for the feedback. It would be great if you confirm that the patch is working at #3001030: The excerpt should leverage the backend highlighter with stemmer & accents, too!

mkalkbrenner’s picture

Title: ASCIIFoldingFilterFactory not working with Drupal » Excerpts don't work with stemmed keys or diacritical marks
Component: Miscellaneous » Code
Category: Support request » Bug report
Status: Fixed » Needs review
Related issues: +#2973763: Highlight excerpts in view results don't work with Arabic anymore.
StatusFileSize
new7.58 KB

  • mkalkbrenner committed 894265d on 8.x-2.x
    Issue #3003524 by mkalkbrenner: Excerpts don't work with stemmed keys or...
  • mkalkbrenner committed f6e8612 on 8.x-2.x
    Issue #3003524 by mkalkbrenner: Excerpts don't work with stemmed keys or...
mkalkbrenner’s picture

mkalkbrenner’s picture

Status: Needs review » Fixed

Status: Fixed » Closed (fixed)

Automatically closed - issue fixed for 2 weeks with no activity.

ncg777’s picture

It would be really nice to include a fix for this issue in the next release for 8.x-3.x.