Problem/Motivation
Under index >> Processors I have enabled the "Ignore characters" Processor to strip out the characters ['¿¡!?,.:;] (the default) under "Strip by regular expression" and it works really well.
But there is an edge case. I am using Search API Autocomplete and I'd like to actually not strip the dot from URLs such as "finland.fi" in the suggestions.
Currently, I get the suggestion "finlandfi" where I'd like to get "finland.fi" (or any other domain example.com, example.org etc.) suggested, and subsequent do a search for "finland.fi".
Is this possible?
My processors
Preprocess index
- Ignore case
- HTML filter
- Stopwords
- Transliteration
- Tokenizer
- Ignore characters
Preprocess query
- Content access
- Ignore case
- HTML filter
- Stopwords
- Transliteration
- Tokenizer
- Ignore characters
Steps to reproduce
- Index text containing domain names such as example.org
- Use Search API Autocomplete and see that exampleorg is suggested (the dot is removed)
- Try to update the "Ignore characters" > "Strip by regular expression" to something else to only strip dots immediately followed by space, and see that it doesn't make any difference
The default value in the "Strip by regular expression" field:
['¿¡!?,.:;]
Suggested alternatives, which don't seem to only strip dots followed by a space character, while retaining dots immediately followed by another character:
['¿¡!?,:;]|\.\s['¿¡!?,:;]|\.(=? |$)
Proposed resolution
Remaining tasks
| Comment | File | Size | Author |
|---|---|---|---|
| #9 | 3359908-9--ignore_characters_add_test.patch | 598 bytes | drunken monkey |
Comments
Comment #2
ressaComment #3
ressaComment #4
drunken monkeyNo, I’m afraid it is not, at least as far as I can see.
I think your best bet for this would be to create a custom processor plugin that does this. Maybe, after tokenizing, it would just remove dots from the beginning and end of each token, but not from the middle?
Anyways, I don’t think this is possible without custom code (or a backend like Solr) currently, sorry.
As a support request, I’d say this is fixed, unless you have further questions on this.
However, also feel free to re-open it as a feature request. In that case, I’d leave it open and wait whether more people are interested in this – but that would surely take some time. In any case, it shouldn’t be too hard to implement, but I’m very wary of feature creep, so try not to add functionality that will be used by too few people.
Comment #5
ressaThanks for a fast reply, I really appreciate it. And your answer makes total sense, to avoid creating custom solutions as much as possible ...
I guess my hope was that since all dots I want removed will always be followed by a space i.e "the system. But then again" and never "the system.But then again".
So I hoped a rule like this would work, but it doesn't:
['¿¡!?,:;]|\.\sAnyway, as you write, Solr might be able to do this, so I will keep my fingers crossed that Solr Cloud integration for Lando will arrive some day :)
Comment #6
drunken monkeyOh, that is a good point about the regular expression. Didn’t occur to me to use it like that.
Yes, something like this should work:
['¿¡!?,:;]|\.(=? |$).Thanks for mentioning it.
Comment #7
ressaYou're welcome, and thanks for your suggestion. But the thing is, it doesn't seem to work ... Changing status, maybe it turns out we found a bug? :)
Comment #8
ressaComment #9
drunken monkeyNo, I just produced the regular expression from memory and flipped the syntax for a lookahead assertion. Otherwise it works as expected – adding a test to confirm this.
Please try this instead:
['¿¡!?,:;]|\.(?= |$)Comment #11
drunken monkeyCommitted the additional test.
Please report back when you tried out the corrected regular expression so we can, hopefully, mark this ticket “Fixed”.
Comment #12
ressaI am not able to make it work, just with the bare minimum, and all filters removed, to make it really simple ... are you able to suggest the steps? With the steps below, I only get
examplesuggested, notexample.com:examplesuggested, notexample.comIs it Search API Autocomplete which strips dots?
Comment #13
drunken monkeyIt won’t work without any processors, as the DB backend will treat all non-alphanumeric characters as whitespace if text is not already tokenized. Please keep at least “Ignore characters” and “Tokenizer” enabled, and configure them accordingly: the former to ignore
['¿¡!?,:;]|\.(?= |$)(or something like that), the latter to not ignore periods or treat them as whitespace.