Closed (fixed)
Project:
Search API
Version:
8.x-1.30
Component:
Database backend
Priority:
Normal
Category:
Bug report
Assigned:
Unassigned
Reporter:
Created:
26 Oct 2023 at 19:23 UTC
Updated:
25 Nov 2023 at 14:24 UTC
Jump to comment: Most recent, Most recent file
Comments
Comment #2
gaddman commentedFrom what I can tell the indexed entries are almost unique, but the trailing space is causing problems. So there are two entries:
1. 'abcdefjklmnopqrstuvwxyzabcdefjklmnopqrstuvw': the 49-char word
2. 'abcdefjklmnopqrstuvwxyzabcdefjklmnopqrstuvw ': the 49-char word plus a space (it would have been a phrase consisting of the 49-char word plus a space plus the following word but is trimmed to 50 chars).
These are being treated as identical in the MariaDB database:
MariaDB version: 10.4.17-MariaDB
From my read of the docs on varchar and collation this is expected behaviour with the configured
utf8mb4_bincollation.Comment #3
gaddman commentedComment #4
gaddman commentedIssue #3199355 introduced a fix to strip trailing spaces to avoid these errors but it excludes type=='text'. See /modules/search_api_db/src/DatabaseCompatibility/MySql.php. Not sure why it excludes text?
Comment #5
drunken monkeyThanks a lot for reporting this problem!
I could easily replicate and fix it. And while the problem is only present in MySQL, I guess adding a token that is just a word plus a single trailing space doesn’t make sense for any of the other DBMSs, either. Still, applying the
rtrim()also to text fields probably makes sense, too, if MySQL is incapable of treating those correctly.Which is all you need to know about MySQL. Seriously, wtf?
Comment #6
gaddman commentedAwesomely quick patch! Tested and works fine, thanks a heap.
This is a bit of a nitpick+tangent, but your comment here got me thinking:
The code is checking the length of the prev_word, not the bigram itself. No big deal, but that got me looking into what happens if the bigram itself is too long, and I see it gets truncated further down, eg the text abcdefjhijklmnopqrstuvwxyzabcdefjhijklmnopqrs tuvwxyz is stored as the phrase
abcdefjhijklmnopqrstuvwxyzabcdefjhijklmnopqrs tuvw. It's a different scenario and may not cause any problems, but is there any value in storing a truncated bigram or should they just be discarded? I did some brief testing and don't get any results when searching for abcdefjhijklmnopqrstuvwxyzabcdefjhijklmnopqrs tuvw but I do get results for abcdefjhijklmnopqrstuvwxyzabcdefjhijklmnopqrs tuvwxyz, with and without quotes, so I guess it still works somehow?Comment #8
drunken monkeyThanks for reporting back, good to hear it works for you.
Merged. Thanks again!
And yes, overlong bigrams are already handled “properly” (i.e., as well as possible) in
\Drupal\search_api_db\Plugin\search_api\backend\Database::splitKeys().