Algolia limits record size to 10kb.

Sometimes there is more text in the body field than that and algolia will return an error when trying to index a record.

Maybe it would be possible to somehow truncate the body field value to for example 8 or 9kb before sending it to algolia?

Comments

lennart created an issue. See original summary.

mlbrgl’s picture

Status: Active » Closed (won't fix)

Hi, I see your point but adding a rule for trimming the body when the whole record approaches 8 - 9 KB is only going to work in some cases, where most of the content is in the body. On more complex content types, the data might be spread out evenly in more fields, or the body might just be absent, etc ... so hardcoding "trim body if record > 9 KB" will not have any effect.

The way I see this being solved at the moment is on a per site basis, through search api processors defined in a custom module.

More information about processors:
- https://www.drupal.org/node/2004270
- https://www.drupal.org/node/1254452

And an implementation example:
https://www.previousnext.com.au/blog/writing-custom-drupal-search-api-pr...

If and when you have a working example, through which people can choose what field to trim, how many words are being kept, etc..., please post it back here and I will consider providing it as part of this module.

As a side note, I actually talked to Algolia about this as some of my records were also going over the limit. They indicated that one of the reason for small records was keeping the relevance of the hits high. You can read more about how and why they split long articles into small records on their blog : https://blog.algolia.com/how-to-build-a-helpful-search-for-technical-doc.... I also went on and created a plugin for Kirby to somehow replicate this behavior: https://github.com/mlbrgl/kirby-algolia. You might find some pointers there too if you want to go down that route.

Matthieu

lennart’s picture

Thanks for the answer and links,

Splitting / fragment indexing like suggested by Algolia and in your Kirby plugin seems to be the better way forward since all data will be available in the index and at the same time relevance could be improved.

The only drawback is that splitting seems to require a more structured text with headings dividing the paragraphs into smuller chunks in the text though I suppose it could be done even in cases were there are just a lot of (long) paragraphs and no headings in a longer text.

Regs,
Lennart

lennart’s picture

The following might also be of interest:

https://github.com/algolia/php-dom-parser