I'm trying to get the search API running with SOLR backend, czech stemmer and ASCIIFoldingFilterFactory.

If I search for 'zakon', backend will correctly match it to 'zákon' and will return the expected results.

While highlighted fields will get generated correctly, excerpt seems to completely ignore the 'highlighted_fields' extraData and tries to build it by itself - and fails as it can't match zakon to 'zákon'. I've tried fixing it inside the excerpt code and it's fine, until you start talking about stemming. Then there's not much I can do in the code any more.

Would it not make sense to use the highlighted_data for excerpt generation too?

If you need any more info, let me know.

Comments

mvantuch created an issue. See original summary.

mvantuch’s picture

StatusFileSize
new2.87 KB

There's a patch I use just to fix the accents.

drunken monkey’s picture

Project: Search API » Search API Solr
Component: General code » Code

As you correctly say, we don’t really have a chance to take all possible filters in Solr into account when creating the excerpt. We just use a comparatively simple algorithm to get something.
Probably the best approach to improve this, at least for Solr, is to take the highlighted fields data Solr provides and create the excerpt from that. The problem with that, though, is that we can’t know what markup the Solr backend is using for highlighting.
So, probably, this should be implemented in the Solr backend, but I’m not sure. I’ll just move the issue for now and we’ll see what Markus, the Solr module maintainer, says.

mkalkbrenner’s picture

If you use search_api_solr 8.x-2.x it should work well without any patch required!

Just enable the Highlight processor and excerpt AND check 'Retrieve result data from Solr' and 'Highlight retrieved data' on the server edit page.

Does it work now?

mkalkbrenner’s picture

Title: Excerpt with stemmer & accents » Document highlighting and excerpt with stemmer & accents
Category: Bug report » Task
Status: Active » Needs review
StatusFileSize
new872 bytes

  • mkalkbrenner committed cc4974a on 8.x-2.x
    Issue #3001030 by mkalkbrenner: Document highlighting and excerpt with...
mvantuch’s picture

Hi Markus,

Well... not really. It kind of works when the searched phrase is directly the same as the content returned from the server, but not if they are different in the accents. I'll give you an example:

If I search for 'noveho', SOLR correctly matches it to 'nového' and I even get the correct highlight data. The SearchApiSolrBackend::getHighlighting works well and correctly assigns the data there (something like "Potřebný například v případě [HIGHLIGHT]nového[/HIGHLIGHT] připojení k elektřině, změny jističe nebo jeho zaplombování.")

Up to this point it all looks good. But when going through the code, I can't find anywhere it'd use it and more importantly, it doesn't get picked up by the views - the `search_api_excerpt` is empty.

I might be just as well missing something and am sorry if I am.

Just to add info about my setup, in the fields in schema.xml I use:

Probably the first one will be sufficient to replicate this.

mkalkbrenner’s picture

Status: Needs review » Needs work

OK, I confirm that field highlighting works, but excerpt doesn't.
The highlighting processor uses the already highlighted fields provided by the backend.
But the excerpt builder doesn't respect these fields and applies it's pure PHP logic.

mvantuch’s picture

Yes, that sounds right. Sorry I wasn't more specific with this issue but I'm still trying to find my way around all this in code.

I've tried simply using the highlighted fields which are already present in the data, it works fine untill there is a fulltext field, which obviously breaks the excerpt by being "too long".

mkalkbrenner’s picture

Title: Document highlighting and excerpt with stemmer & accents » The excerpt should leverage the backend highlighter with stemmer & accents
Project: Search API Solr » Search API
Component: Code » General code
Category: Task » Feature request
Status: Needs work » Needs review
StatusFileSize
new1.92 KB

Here's a different patch that uses the highlighted fields provided by the backend if available.
It's hackish and I wonder if Thomas will come up with a better implementation.
Compared to the patch in #2 it solves the issue for stemmed keys, too.

mvantuch’s picture

Just tested it and seems to work well enough.

drunken monkey’s picture

Hm, that’s an interesting solution. Just using the highlighted data to retrieve the matched keys and then do highlighting like normal is not a solution I would have thought about.
However, it does suffer the same problem as just using the highlighted data directly: we don’t actually know the tags the backend will use for highlighting. Using the same as configured for the Highlight processor is really just guesswork.

The solution I’d have thought about would be to use the highlighted field data directly in the Solr backend to piece together an excerpt manually. I.e., get all highlighted data, look for the matches in there ([HIGHLIGHT] tags) and then use the same algorithm as in the Highlight processor to extract snippets based on that, with some context before and after.
The downside here would of course be (apart from bloating up the backend code further) that this would have to be re-implemented for other backends with highlighting support. We might be able to provide a static helper method for that on the Highlight processor, though – maybe even refactor it to use the same internally, too.
(One change that would have to happen in the Highlight processor in any case is that it seems we currently don’t check for existing excerpts, so we’d just override those with our own if the user enabled excerpts in the Highlight processor, too.)

To make this work completely in the Highlight processor, we’d have to somehow get the information from the backend about the tags used for highlighting. Then we could just have the code discussed above in the Highlight processor directly.
The downside here would be, though, that that way we couldn’t get the raw data as retrieved from Solr, just the data already preprocessed and escaped. Not sure if that would be a problem, though.

Either way, this would be quite a bit of code, so we’d need someone to work on that. As it’s a feature request, I don’t think I’ll have time to work on it any time soon – sorry.

mkalkbrenner’s picture

Thomas, what about an intermediate solution and use my hackish patch as it is and to open a follow-up.
I assume that the Solr backend is the only one at the moment that provides highlighted snippets AND uses the Search API Highlighting processor.

We can add open a follow up and add a todo to the code.

On the other hand it would be very easy for my to capture the matched tags in the backend. If you suggest an API to set them on the result, I can do it.

mkalkbrenner’s picture

As the multilingual backend is now part of search_api_solr 8.x-2.x itself, I expect more and more bug reports.

drunken monkey’s picture

Thomas, what about an intermediate solution and use my hackish patch as it is and to open a follow-up.
I assume that the Solr backend is the only one at the moment that provides highlighted snippets AND uses the Search API Highlighting processor.

Usually, when we do that and the problem is thereby hidden, no-one ever works on the proper solution. If you’ll still work on a proper solution afterwards, however, sure, we can commit this as a quick workaround for now. (Although one or two more people saying it works at least “well enough” wouldn’t hurt, either.)

On the other hand it would be very easy for my to capture the matched tags in the backend. If you suggest an API to set them on the result, I can do it.

How about this:

$results->setExtraData('highlighted_fields_metadata', [
  'highlight_prefix' => $highlight_config['prefix'],
  'highlight_suffix' => $highlight_config['suffix'],
]);

However, the tricky part would then of course be to actually use that information to extract the matched fragments in the processor.
The upside would be that this could also be made to work with the processor’s own field highlighting – which would save a bit of performance in cases where you want both highlighted fields and an excerpt.

mvantuch’s picture

(Although one or two more people saying it works at least “well enough” wouldn’t hurt, either.)

Yep, this works well enough.

Using this patch, when searching for 'nove' I do get the expected words highlighted: 'nového' and 'nové' - so yeah, much better!

mkalkbrenner’s picture

StatusFileSize
new774 bytes

How about this:

$results->setExtraData('highlighted_fields_metadata', [
  'highlight_prefix' => $highlight_config['prefix'],
  'highlight_suffix' => $highlight_config['suffix'],
]);

However, the tricky part would then of course be to actually use that information to extract the matched fragments in the processor.
The upside would be that this could also be made to work with the processor’s own field highlighting – which would save a bit of performance in cases where you want both highlighted fields and an excerpt.

I thought about an API to pass the highlighted "keys" to Search API:

$results->setExtraData('highlighted_keys', [
  'nového',
  'nové',
]);

The patch would for Search API would be very simple in that case.

mkalkbrenner’s picture

StatusFileSize
new1.15 KB

Sorry, this one looks better. Note, it's just a proposal and I'll implement the counterpart once you agree on it.

The last submitted patch, 17: 3001030_highlighted_keys.patch, failed testing. View results

drunken monkey’s picture

Component: General code » Plugins
StatusFileSize
new4.04 KB
new4.17 KB

Ah, OK. That should of course also work, sure. And if it’s easier to implement, then let’s go with that one, yes.

Just a few things were missing: some inline comment explaining what’s going on, documentation of the key in itemInterface::getExtraData() and a test.
All added in the attached patch, please test/review!

leoperezpulido’s picture

Hi, I come from a related issue. I have been looking for a solution to eliminate the diacritical marks on tokens from content. Due to my Solr background, my first attempt was to apply the ASCIIFoldingFilterFactory editing the schema.xml file. Then I came across that Drupal have its own code layer on top of Solr.

I have tested the patch 3001030.patch with Spanish and French languages and it do what I have been looking for, it eliminates the diacritics from words and let me do search without typing accents.

mkalkbrenner’s picture

Status: Needs review » Reviewed & tested by the community

Thanks, Thomas!

I tested your latest patch manually and now I'm preparing a patch and tests for search_api_solr.

drunken monkey’s picture

Status: Reviewed & tested by the community » Fixed

OK, great to hear!
Committed.
Thanks again, everyone!

Status: Fixed » Closed (fixed)

Automatically closed - issue fixed for 2 weeks with no activity.