I'm using Search API with a database backend to index content which is mostly the common names of species of wildlife. I'm struggling to get the configuration right for species names with hyphens in, e.g. Lesser-Spotted Woodpecker. Ideally, this should be indexed as 3 words and the hyphen treated as a word break, so the indexing would capture "Lesser", "Spotted" and "Woodpecker". However, the only options for handling the hyphen are:
* Ignore characters processor can remove the hyphen
* Tokenizer processor had a white space regex which should be usable to treat the hyphen as white space. But, on inspecting the code I found the following in the simplifyText method which removes the hyphen completely before the regex gets a chance to convert it to a space:

    // The dot, underscore and dash are simply removed. This allows meaningful
    // search behavior with acronyms and URLs. See Unicode note directly above.
    $text = preg_replace('/[._-]+/', '', $text);

In most cases in my index, I want the hyphen to be treated like whitespace but it doesn't seem to be possible because of the above. Has anyone any tips for this scenario please?

Comments

johnvb created an issue. See original summary.

danharper’s picture

Hi,

I have a similar issue with SKU codes for example I have a product JLR-ASS-0501 and I want the product to return in the search when the user searches JLR and JLR-ASS but I can't seem to make it work.

Cheers Dan

drunken monkey’s picture

Title: Handling hyphenated text » Add Tokenizer option to handle dashes as whitespace
Version: 8.x-1.17 » 8.x-1.x-dev
Component: General code » Plugins
Category: Support request » Feature request

Urgh, we should never have split out “Ignore characters” into its own processor …
You’re right, it’s definitely not ideal that this can’t be handled properly. (Also, it seems really strange to me that underscores would be ignored. They seem pretty whitespace-y to me. Maybe we had a good reason? (But why shouldn’t multiple underscores, at least, count as whitespace?))

If one of you would be willing to provide a patch, I’m up for making this configurable. We’d have a second option for the “ignored” regex, just as for spaces, and maybe even a checkbox for counting multiple consecutive “ignorables” as spaces. (Also a few data sets added to TokenizerTest to make sure this works.)

b_sharpe’s picture

Status: Active » Needs review
StatusFileSize
new539 bytes
new3.98 KB

Here's a patch/test to allow specifying the ignored characters.

One note is that you can't really have the best of both worlds as far as I can see. By which I mean if your word is "month-end" you can use this to allow both "month-end" and "month end", but not "monthend" since the processor will run either before/after the ignore characters processor, so it's kinda one or the other, which in my opinion is fine as it's either a dash is a space or it's ignored...

The last submitted patch, 4: 3145955-ignore-characters-test-only.patch, failed testing. View results

drunken monkey’s picture

StatusFileSize
new5.5 KB
new6.01 KB

Thanks a lot, looks great already!
I just saw that we already have an ignorable key in the config schema for the processor – seems we forgot that when removing the setting. Let’s overwrite it now (but ignored probably still makes more sense, you’re right).
Also, as said, I think the characters treated as word boundaries when repeated should be the same as the ignored characters. Also adding test for that.

Anyways, thanks again, great work! Please test/review my attached revision!

b_sharpe’s picture

Going to test this shortly, but was looking at the diff and this confuses me:

+        'foobar',
+        [Utility::createTextToken('bar')],
+        ['ignored' => 'o'],

shouldn't that return "fbar"? why would it split the "f"?

EDIT: NVM, missed that part about multiple characters in the interdiff.

b_sharpe’s picture

Status: Needs review » Reviewed & tested by the community

Looks good!

drunken monkey’s picture

Status: Reviewed & tested by the community » Fixed

Good to hear, thanks a lot for reviewing!
Committed.
Thanks again!

Status: Fixed » Closed (fixed)

Automatically closed - issue fixed for 2 weeks with no activity.