Problem/Motivation
Algolia has a maximum size limit for a single record. search_api_algolia currently provides a truncate option which truncates strings to 10000 characters in an effort to avoid hitting the limit. This is not ideal because it results in data not being indexed.
Proposed resolution
In an effort to solve this problem I created an Algolia Item Splitter processor. The processor allows setting a record size limit, which should be 10k in most cases. A base record is created with all non-splitted attributes. Splitted attributes will keep being added until the record hits the limit. in that case, the attribute will be split and a new splitted object will be created.
In order to avoid duplicate records in Algolia results, we have to set an attributeForDistinct in the config for the index. I'm personally using url but you can use whatever you want, as long as it's unique per record. Algolia then combines all the records with the same url into one (as far as the end user can see).
How to test patch:
- Apply patch.
- Go to your Algolia search_api server and make sure that truncation is disabled.
- Go to an Algolia search_api index. Click the Processors tab.
- Enable the Algolia Item Splitter processor.
- Enable the processor only on fields that are likely to be long, such as the body field or rendered HTML.
- If you enable on rendered HTML, you should also enable the HTML trimmer processor on the rendered HTML field.
- Set the character limit. For testing purposes, set it low enough that it is guaranteed to cause splits to be created.
- Identify a node that will be indexed that has a field value that will be split based on the character limit you set.
- Clear and reindex.
- Go to your Algolia dashboard.
- In the Browse tab, search for the node you identified. You should see that there are multiple records for the same node. The value of the field you enabled the processor for should be split amongst the records.
- In the Algolia dashboard, go to Configuration > Deduplication and Grouping. Set Distinct to true. Set the Attribute for Distinct to a field guaranteed to have a unique value, such as the url or nid.
- Go back to the Browse tab and search for the node again. You should only see one record. The field that was split will not show the entire field value. This is apparently another quirk of Algolia. I do not know what happens when you try to render the split field on the frontend.
- In the Browse tab, search for something that appears in the split field value that does not appear on the one record you can see. You should find that even though you can't see the entire field value, the record is still returned in search results, meaning that the entire value of the field is being searched.
Remaining tasks
None.
Known issues
None.
User interface changes
Creates an Algolia Item Splitter processor available under Processors tab of index.
| Comment | File | Size | Author |
|---|---|---|---|
| #34 | search_api_algolia_3256840.patch | 29.07 KB | josh.stewart |
Issue fork search_api_algolia-3256840
Show commands
Start within a Git clone of the project using the version control instructions.
Or, if you do not have SSH keys set up on git.drupalcode.org:
Comments
Comment #2
maskedjellybeanApologies for the formatting issues of the description. I wish this site understood markdown.
Comment #3
maskedjellybeanComment #4
bgilhome commentedThis is great! Thanks for this @maskedjellybean. I currently have an approach which uses hook_search_api_algolia_objects_alter() to add the split items - but your process plugin is a much better approach IMO.
Re: deletion of split items on deletion of the original item (e.g. on 'Clear all indexed data') - I use a custom search_api backend plugin which extends SearchApiAlgoliaBackend and overrides the deleteItems() method to add a \Algolia\AlgoliaSearch\SearchIndex::deleteBy() query using filters to select any items with the same value for the distinct attribute.
Re: avoiding splits in the middle of words etc - I have a 100 char overlap between split fragments (i.e. last 100 chars of split 1 is the same as first 100 chars of split 2), the idea being to return hits for a search for a phrase that would otherwise be split across the two (the deduplication will avoid returning both). Is that sensible? I can't decide if it's just hacky or not :)
Comment #5
cesarg commentedThanks for publishing this, I am about to add it to my dev environment and test it out.
Comment #6
cesarg commentedI'm getting an error on all records that exceed the character limit:
Notice: Trying to access array offset on value of type null in Drupal\search_api_algolia\Plugin\search_api\processor\ItemSplitter->getSplitsForItem() (line 68 of /app/web/modules/contrib/search_api_algolia/src/Plugin/search_api/processor/ItemSplitter.php)This is on Drupal 9.4.5 running on PHP Version 7.4.3
I'll keep debugging but if you fellas have an idea, I would greatly appreciate the help.
Comment #7
cesarg commentedI think ItemSplitter.php line 68 should change from this:
if (!empty($this->splits[$item_id] && !empty($this->splits[$item_id][$field_name]))) {to this:
if (!empty($this->splits[$item_id]) && !empty($this->splits[$item_id][$field_name])) {Unfortunately, I still don't see multiple records on the algolia side tho.
Comment #8
cesarg commentedGot it to work! My issue... aside from the code change above; was that I was using "Fulltext" as field type. As noted in the instructions, I switched it to "String", and I now see my split records in Algolia.
Comment #9
cesarg commentedOne more note.... the instructions say "should also enable HTML trimmer" it "must" also be enabled or else it didn't work for me.Comment #10
jonloh commentedTried the patch, but unfortunately this does not work well in Multi-lingual setup.Sorry just saw above that it has some issue with multilingual code.
Comment #11
akhil babuThanks for the patch. I have created a new patch with few changes.
Comment #12
akhil babuComment #13
maskedjellybeanThank you for carrying this forward! Sadly I no longer have an Algolia project to work with so I can't test the new patch out.
Comment #14
nikunjkotechaThis is good. I am not convinced though that we should index huge objects in Algolia, can we have some real use case to help understand the need for this?
Comment #15
maskedjellybeanThe use case is if you want to index more than 10000 characters in one record. :-)
Algolia offers the ability to split records in order to get around their character count limitation, so it would be great if search_api_algolia leveraged this ability.
Potentially site builders/developers may not realize their records are being truncated. When it is truncated search does not search the entire record because only part of it is indexed. This means worse search results without any indication why.
Comment #16
kevinb623 commentedThis patch is working wonderfully to properly index and discover lengthy pages on a content rich website we manage.
Only suggestion is to update ItemSplitter.php line 68 to use isset() to reduce PHP warnings related to unknown and null array keys.
Very nice work!
Comment #17
reecemarsland commentedOur use case is indexing PDF files attached to content and we need the PDF content to be searchable.
Comment #18
den tweed commentedSame as in #17 our use case is making attached documents searchable
I've worked further on patch #11 and changed following:
Comment #19
dieterholvoet commentedWe hit this limit regularly on projects, when e.g. indexing long text fields or paragraphs for search. This is a very valid use case.
Comment #21
dieterholvoet commentedI started a MR based on the latest patch. I'm sometimes still getting the following error, even with the patch applied:
I'll do some debugging.
Comment #22
dieterholvoet commentedComment #23
dieterholvoet commentedI can't figure out the problem. That project might have been using an outdated patch, I updated it and I'll wait and see if the issue happens again.
Comment #25
dieterholvoet commentedThe existing splitter doesn't work consistently for me. Splitting up all enabled fields on a fixed amount of characters works quite well if you only have one very big body field. If you have multiple fields with a lot of content, splitting on a fixed amount of characters still has the risk of creating records that are too big, unless you set the amount of characters to a low value.
That's why I decided to rewrite everything and to come up with a smarter splitter. Instead of splitting all enabled fields on a fixed amount of characters, my splitter fills up records until the limit which is dictated by Algolia (usually 10K bytes), before it starts splitting text into multiple records. This makes it practically impossible to create records that are too big + it's a lot more efficient, it will only create as much records as necessary.
Most of the logic was moved from the field processor to the search backend code, right before the record is sent to Algolia, in order to be able to calculate the record sizes as efficiently as possible with all base fields included.
Comment #26
dieterholvoet commentedI also improved documentation of the processor, warning users to set up things correctly on the Algolia side. I also changed it so the truncate option is automatically disabled for an index when splitting is enabled.
Comment #27
dieterholvoet commentedI cleaned up the issue description. About the known issues previously listed:
This is not true. When clicking that button,
deleteAllIndexItems()is triggered, which clears the whole index instead of specific objects. I'll remove this from the known issues.This is not the case in my implementation. Removing from known issues.
I would also say this works as expected. The splitted items are an implementation detail of this specific backend and are not necessary to be listed in the UI. When you display the index on dashboard.algolia.com, splitted items are also not listed or counted separately. Removing from known issues.
This is not the case in my implementation since splits are indexed together with regular objects. Removing from known issues.
This is not true. Algolia doesn't merge the contents of splitted items. When searching an Algolia index and multiple splitted items match the query, the splitted item that matches the most will be returned to the user. This means that all non-splitted attributes need to be present on all splitted objects. Removing from known issues.
This is not true. It splits on spaces, so it shouldn't break words. The current code looks plenty smart to me. Removing from known issues.
Comment #30
josh.stewart commentedWe ended up using the following patch after we ran into some issues. Might not be 100% perfect but we hit some snags with Korean characters and the splitting actually malforming them because of their bytes length. Hopefully this helps someone. Previous work on the MR worked great as far as I can tell other than that issue.
Comment #31
dieterholvoet commented@josh.stewart I added the changes from your patch to the MR. What do you mean by 'Might not be 100% perfect', anything specific to look out for?
Comment #32
josh.stewart commented@dieterholvoet I was able to index 70k records without issue after the updates so the comment about not 100% perfect was maybe just a lack of confidence around testing edge cases. But it's working really well for us at the moment.
Comment #33
dieterholvoet commentedGood to hear, thanks for the help!
Comment #34
josh.stewart commentedRan into a warning message with the Annotation parser dealing with the link in the description.
[error] Doctrine\Common\Annotations\AnnotationException while computing Views data for index Main: [Syntax Error] Expected Doctrine\Common\Annotations\DocLexer::T_CLOSE_PARENTHESIS, got 'https' at position 234 in class Drupal\search_api_algolia\Plugin\search_api\processor\ItemSplitter. in Doctrine\Common\Annotations\AnnotationException::syntaxError() (line 28 of /app/vendor/doctrine/annotations/lib/Doctrine/Common/Annotations/AnnotationException.php).So this file is to update it to remove the html in there.
Comment #35
dieterholvoet commented@josh.stewart could you please add those changes to the MR? Patches aren't being used anymore on Drupal.org. Thanks!
Comment #37
jordan.caldwell commentedI ran into the Annotation parser issue as well. I've pushed an update to the MR to resolve it.
Comment #39
nikunjkotechaThanks everyone for hard work, based on last few comments I can say it is already tested in multiple projects so merging it.