Problem/Motivation

Under index >> Processors I have enabled the "Ignore characters" Processor to strip out the characters ['¿¡!?,.:;] (the default) under "Strip by regular expression" and it works really well.

But there is an edge case. I am using Search API Autocomplete and I'd like to actually not strip the dot from URLs such as "finland.fi" in the suggestions.

Currently, I get the suggestion "finlandfi" where I'd like to get "finland.fi" (or any other domain example.com, example.org etc.) suggested, and subsequent do a search for "finland.fi".

Is this possible?

My processors

Preprocess index

  • Ignore case
  • HTML filter
  • Stopwords
  • Transliteration
  • Tokenizer
  • Ignore characters

Preprocess query

  • Content access
  • Ignore case
  • HTML filter
  • Stopwords
  • Transliteration
  • Tokenizer
  • Ignore characters

Steps to reproduce

  1. Index text containing domain names such as example.org
  2. Use Search API Autocomplete and see that exampleorg is suggested (the dot is removed)
  3. Try to update the "Ignore characters" > "Strip by regular expression" to something else to only strip dots immediately followed by space, and see that it doesn't make any difference

The default value in the "Strip by regular expression" field:
['¿¡!?,.:;]

Suggested alternatives, which don't seem to only strip dots followed by a space character, while retaining dots immediately followed by another character:

  • ['¿¡!?,:;]|\.\s
  • ['¿¡!?,:;]|\.(=? |$)

Proposed resolution

Remaining tasks

Comments

ressa created an issue. See original summary.

ressa’s picture

Issue summary: View changes
ressa’s picture

Issue summary: View changes
drunken monkey’s picture

Status: Active » Fixed

No, I’m afraid it is not, at least as far as I can see.
I think your best bet for this would be to create a custom processor plugin that does this. Maybe, after tokenizing, it would just remove dots from the beginning and end of each token, but not from the middle?
Anyways, I don’t think this is possible without custom code (or a backend like Solr) currently, sorry.

As a support request, I’d say this is fixed, unless you have further questions on this.
However, also feel free to re-open it as a feature request. In that case, I’d leave it open and wait whether more people are interested in this – but that would surely take some time. In any case, it shouldn’t be too hard to implement, but I’m very wary of feature creep, so try not to add functionality that will be used by too few people.

ressa’s picture

Thanks for a fast reply, I really appreciate it. And your answer makes total sense, to avoid creating custom solutions as much as possible ...

I guess my hope was that since all dots I want removed will always be followed by a space i.e "the system. But then again" and never "the system.But then again".

So I hoped a rule like this would work, but it doesn't: ['¿¡!?,:;]|\.\s

Anyway, as you write, Solr might be able to do this, so I will keep my fingers crossed that Solr Cloud integration for Lando will arrive some day :)

drunken monkey’s picture

Oh, that is a good point about the regular expression. Didn’t occur to me to use it like that.
Yes, something like this should work: ['¿¡!?,:;]|\.(=? |$).
Thanks for mentioning it.

ressa’s picture

Category: Support request » Bug report
Issue summary: View changes

You're welcome, and thanks for your suggestion. But the thing is, it doesn't seem to work ... Changing status, maybe it turns out we found a bug? :)

ressa’s picture

Issue summary: View changes
drunken monkey’s picture

Component: General code » Plugins
Category: Bug report » Support request
Status: Fixed » Needs review
StatusFileSize
new598 bytes

No, I just produced the regular expression from memory and flipped the syntax for a lookahead assertion. Otherwise it works as expected – adding a test to confirm this.
Please try this instead: ['¿¡!?,:;]|\.(?= |$)

  • drunken monkey committed a7e96e22 on 8.x-1.x
    Issue #3359908 by drunken monkey: Added another test case for the "...
drunken monkey’s picture

Status: Needs review » Active

Committed the additional test.
Please report back when you tried out the corrected regular expression so we can, hopefully, mark this ticket “Fixed”.

ressa’s picture

I am not able to make it work, just with the bare minimum, and all filters removed, to make it really simple ... are you able to suggest the steps? With the steps below, I only get example suggested, not example.com:

  1. Download drupal/search_api drupal/search_api_autocomplete
  2. Enable search_api search_api_db search_api_db_defaults search_api_autocomplete
  3. Remove all filters
  4. Enable Search API autocomplete
  5. Enable Retrieve from server
  6. Add the string example.com in a node and index it
  7. Type "exam" in the field, and get example suggested, not example.com

Is it Search API Autocomplete which strips dots?

drunken monkey’s picture

It won’t work without any processors, as the DB backend will treat all non-alphanumeric characters as whitespace if text is not already tokenized. Please keep at least “Ignore characters” and “Tokenizer” enabled, and configure them accordingly: the former to ignore ['¿¡!?,:;]|\.(?= |$) (or something like that), the latter to not ignore periods or treat them as whitespace.