Hi -

I'm not sure if this belongs in the Search API issue queue or the Search API Solr issue queue, so I'll start here. I'm running the latest dev versions of everything - I'm using Search API with the Solr backend. The search index I'm trying to build is based on nodes. Whenever I try to index something, I get an error message, which stems from the Solr backend.

What's happening is that I'm trying to index node titles as Strings, not Fulltext, so that I can use them for sorting (I also have the Search API Sorts module). Even though I thought the Tokenizer processor was only supposed to run on Fulltext fields, it's running on the node titles (Strings) as well. This causes an array of values to be sent to Apache Solr. Since Apache Solr recognizes the node title field as "ss_title," it chokes on the array (expecting a single value instead), and I get messages like this in the dblog:

HTTP ERROR: 400
ERROR: [default_node_index-115] multiple values encountered for non multiValued field ss_title

I can help with this further and look into making a patch, but I'm not sure where in the pipeline this problem should be solved. Should the tokenizer not run on String fields? Should the Solr backend implode arrays when it receives them for single-valued fields? Let me know what you think - or if you think this is a special case stemming from something weird I did on my site.

Thanks - and also, thank you for this incredible module and all the time you put into this issue queue. The Drupal community is lucky to have people like you.

Comments

nkschaefer’s picture

Status: Active » Closed (works as designed)

I'm an idiot. While I spent a lot of time reading code, I never even noticed the checkbox that lets you exclude fields from the tokenizer. It works; sorry for the post here.

drunken monkey’s picture

Title: Tokenizer breaks indexing node titles » Tokenizer should only run on fulltext fields
Status: Closed (works as designed) » Active

You're not an idiot – the Tokenizer processor should not allow users to enable it for non-fulltext fields instead of relying on users to guess this. Other processors (e.g., Stopwords) likewise.

Will have to look into that, thanks for reporting!

BarisW’s picture

Same here, it seems that the search keeps indexing (without loggin an error) when there are String fields assigned to the Tokenizer.

finex’s picture

This problem is reproducible with the latest stable search api release: 7.x-1.5.

drunken monkey’s picture

Status: Active » Needs review
StatusFileSize
new1.13 KB

Thanks for bumping, FiNeX! Attached is a patch that should fix that, only presenting fulltext fields as options for the Tokenizer. Users with an existing misconfiguration would have to re-save the workflow form, though, as this is almost impossible to do in an update hook.
Does anyone know of other processors that should only run on fulltext fields? After a quick scan I think all others should be safe for string fields as well.

drunken monkey’s picture

Status: Needs review » Fixed

Committed.

finex’s picture

Thanks.

Status: Fixed » Closed (fixed)

Automatically closed -- issue fixed for 2 weeks with no activity.