I'm using Search API with a database backend to index content which is mostly the common names of species of wildlife. I'm struggling to get the configuration right for species names with hyphens in, e.g. Lesser-Spotted Woodpecker. Ideally, this should be indexed as 3 words and the hyphen treated as a word break, so the indexing would capture "Lesser", "Spotted" and "Woodpecker". However, the only options for handling the hyphen are:
* Ignore characters processor can remove the hyphen
* Tokenizer processor had a white space regex which should be usable to treat the hyphen as white space. But, on inspecting the code I found the following in the simplifyText method which removes the hyphen completely before the regex gets a chance to convert it to a space:
// The dot, underscore and dash are simply removed. This allows meaningful
// search behavior with acronyms and URLs. See Unicode note directly above.
$text = preg_replace('/[._-]+/', '', $text);
In most cases in my index, I want the hyphen to be treated like whitespace but it doesn't seem to be possible because of the above. Has anyone any tips for this scenario please?
| Comment | File | Size | Author |
|---|---|---|---|
| #6 | 3145955-6--tokenizer_ignored_characters.patch | 6.01 KB | drunken monkey |
Comments
Comment #2
danharper commentedHi,
I have a similar issue with SKU codes for example I have a product JLR-ASS-0501 and I want the product to return in the search when the user searches JLR and JLR-ASS but I can't seem to make it work.
Cheers Dan
Comment #3
drunken monkeyUrgh, we should never have split out “Ignore characters” into its own processor …
You’re right, it’s definitely not ideal that this can’t be handled properly. (Also, it seems really strange to me that underscores would be ignored. They seem pretty whitespace-y to me. Maybe we had a good reason? (But why shouldn’t multiple underscores, at least, count as whitespace?))
If one of you would be willing to provide a patch, I’m up for making this configurable. We’d have a second option for the “ignored” regex, just as for spaces, and maybe even a checkbox for counting multiple consecutive “ignorables” as spaces. (Also a few data sets added to
TokenizerTestto make sure this works.)Comment #4
b_sharpe commentedHere's a patch/test to allow specifying the ignored characters.
One note is that you can't really have the best of both worlds as far as I can see. By which I mean if your word is "month-end" you can use this to allow both "month-end" and "month end", but not "monthend" since the processor will run either before/after the ignore characters processor, so it's kinda one or the other, which in my opinion is fine as it's either a dash is a space or it's ignored...
Comment #6
drunken monkeyThanks a lot, looks great already!
I just saw that we already have an
ignorablekey in the config schema for the processor – seems we forgot that when removing the setting. Let’s overwrite it now (butignoredprobably still makes more sense, you’re right).Also, as said, I think the characters treated as word boundaries when repeated should be the same as the ignored characters. Also adding test for that.
Anyways, thanks again, great work! Please test/review my attached revision!
Comment #7
b_sharpe commentedGoing to test this shortly, but was looking at the diff and this confuses me:shouldn't that return "fbar"? why would it split the "f"?EDIT: NVM, missed that part about multiple characters in the interdiff.
Comment #8
b_sharpe commentedLooks good!
Comment #10
drunken monkeyGood to hear, thanks a lot for reviewing!
Committed.
Thanks again!