It would be great to be able to modify the item boost on a per-item basis. I.e., we should offer a processor that lets you pick a numeric field for each datasource/bundle and at indexing time multiplies the item boost with that field’s value. This way, you could easily configure individual items as being especially important – or, conversely, as not being very relevant as a search result.
If you would be interested in such a feature, please speak up here and, ideally, provide a patch. (A copy of \Drupal\search_api\Plugin\search_api\processor\TypeBoost would be a good place to start.)
Comments
Comment #2
mkalkbrennerIn general I'm not a friend of elevate.xml and prefer LTR (learn to rank).
In your case, why not just adding a keywords field to the entities and give them a high boost?
Comment #3
mpp commentedI like the idea of adding a keywords field with a high boost. Only issue here is that there's weight to distinguish between the keywords.
Comment #4
mkalkbrennerusing LTR it's possible to weight terms individually.
Another approach are payloads as we already use for boosting entity types and HTML tags.
If you use (e)dismax you can also use "bq" and "bf".
And for sure you can use "eleavate".
From a module maintainers perspective, LTR is the way to go and the only approach worth implementing a UI for.
(Elevate is simply too limited and not worth the effort.)
Comment #5
mpp commentedComment #6
mpp commented@mkalkbrenner, let me know if you want me to open another feature request to implementing a UI for LTR to boost individual keywords.
Comment #7
mkalkbrennerFrom a discussion on Slack (I hate it):
To recap: Someone wants to keep the entire boost calculation that exists already and wants to add another criteria based on a numeric field added to the entity.
I see two reliable ways to achieve such a feature
1. LTR
you can create a simple "model" that leverages the calculated relevance as Solr would normally so according to our Search API configuration as one "feature". Then you add a second "feature" based on your field's value.
LTR means learn to rank
2. payloads
Have a look at flattenKeysToPayloadScore().
You can add a numeric value as payload to your custom field when indexing. If I assume that your custom field is an integer or float you can index the same value again in it's payload. These payload values could be "added" to the relevancy calculation.
That's how I reimplemented the HTML tag based boosts.
Comment #8
mkalkbrennerI think that it's not too hard to create a search_api_solr sub-module to offer such an entity field.
Or a Search API processor that acts on a numeric field of choice.
Comment #9
mkalkbrennerMeanwhile there's another contrib module for LTR.
If someone is interested we should implement this simple approach:
Extend the TypeBoost Processor or derive from it. Offer a configuration form to select a numeric field for each datasource / bundle. This field's value has to be set as or add to the calculation of
$item->setBoost(...)inpreprocessIndexItems().Comment #10
drunken monkeyIf a few people would be interested in such a feature, sure, we could add such a processor. Would seem to be rather a niche functionality to me, though.
Also, in any case, I unfortunately don’t have the capacity to work on such feature requests in my spare time, so would wait for someone to provide a patch first. (Would provide details if someone is interested.)
Comment #11
dwwWould this be the right issue for adding a processor to search_api to boost results based on a date field?
https://www.drupal.org/project/issues/search_api?text=date+boost is turning up empty. ;)
I'm trying to improve the relevancy of search results on a site I'm building that involves events. The site owners would like upcoming events to be much more relevant than expired events.
I can't find any existing solution that lets me define a boost factor based on a single (date) field. I can boost all events relative to other node types, but I can't boost upcoming events and unboost (or whatever you call it when you reduce relevancy) to past events.
https://www.drupal.org/docs/8/modules/search-api-solr/search-api-solr-ho... is all about Solr, and sadly, I can't use Solr for this site due to hosting limitations.
https://www.drupal.org/project/search_api_boost_priority doesn't provide a processor like this.
Seems like a not-so-uncommon scenario. Folks also seem to love to be able to boost based on node creation date.
Is this the right issue to discuss/develop said feature, or should I open a new one?
Thanks!
-Derek
Comment #12
laura.gatesI would love a feature like this. Google Search Appliance used to have something like what was described in the description. With GSA hitting end of life a few years ago, UI solutions for things like this have been harder to find in search products. Modifying things like Best Bets in the UI would give clients a bit more freedom in their search results.
Comment #13
drunken monkeyNo, I think that’s a different feature, as it’s more backend-specific. (I.e., I don’t think this is currently even possible on the framework level with pure Search API.) This issue is about adding just a flat numeric boost on a per-item basis.
I see now, though, that the issue title and description are completely out of sync at this point, so I’d better bring them in line again …
For your suggestion, if there is really no issue for it yet, please open a new one. (Do check the Search API Solr issue queue first, though – as said, that feature would probably have to be backend-specific.)
@ Laura: With “a feature like this” do you mean the one from the (former) issue description or the one discussed in #9, and in the new issue description? I.e., what exactly do you want to be able to do?
Comment #14
laura.gates@drunken monkey, I was referring to being able to edit the elevate.xml file in a Search API UI solution that is in the ticket description. I don't have easy access to the elevate.xml file settings with one of my projects.
Comment #15
drunken monkeyAs that is Solr-specific, you should create an issue in the Search API Solr issue queue for that (if there isn’t one already).
Comment #16
mkalkbrennerComment #17
mkalkbrennerhere's a first draft.
Comment #18
mkalkbrennerComment #20
mkalkbrennerComment #22
drunken monkeyThanks a lot for the patch, great work!
I’m not sure that much configuration is really needed, though. I would have just used a select box for a single field to use (always using
1.0andmaxfor boost/aggregation), or maybe checkboxes if we really want to allow multiple fields. It seems to me, though, that the use case would be a separate field, added specifically for this purpose, so there would always be just a single field with a single value. Or do you have a use case for actually re-using some other field(s), where this wouldn’t apply?Anyways, if you do see a use case, I’m not that opposed to leaving the configuration in, but I usually aim to not add unnecessary configs, as the configuration for this module is far too complicated already anyways. (Partly because I haven’t always done a great job of this in the past.)
Speaking of this use case: I think in most cases, people will not actually want to index that boost field, just use its value for the item boost. So it would seem better to leave the user to choose any of the available (first-level) properties, not just the indexed fields.
However, I do see that that is a bit harder to implement – and to fit into the UI. So maybe just doing it like this is still the better solution – the overhead in the index should be pretty small anyways. (Or maybe we could even add an option that will remove the field from indexing after using its value? Would that make sense and be worthwhile, or just too complicated or too confusing?)
Furthermore, as discussed on Slack, item-level boosts for the DB backend unfortunately don’t work without keywords present. So you’ll either need to add a common keyword (like “node”) to all searches, or change the test to only check for correct computation of the boost values (like
TypeBoostTestdoes).Actually, I think the former makes the most sense, especially if we just move the query creation to a helper method. That way, you could also easily override that for the Solr tests. Implemented in the attached patch revision, along with some other changes (mostly nit-picks).
Some further remarks (plus one question):
Haven’t seen this anywhere yet, but good to know it exists!
However, why is it needed here? Is it in case the list is empty? (Don’t think that made problems for us anywhere else before.)
I don’t think this needs to be a property. If we make this a variable in the form builder method, we can also just use a label for the “Don’t use” value, instead of a magic “0.0” value. (I used “Ignore” for the label, but open to input.)
Can’t use that yet, as we still want to stay compatible to Drupal 8 and, thus, PHP 7.0 for a few more months.
Are you sure this worked for you? You’re using
'boost'as the field ID here, but'field_boost'everywhere else. Had to change that to make the tests pass for me.Comment #23
mkalkbrennerScoring is always a difficult task with Solr. It might be that you perfectly adjusted your scoring but later you need to add a different scoring property like location or recent date and notice that the dimensions / relations aren't working. In this case it's easier to adjust the boost "factor" here instead of changing the field value of thousands of entities. That's why I introduced it.
The "first value" is always used for sorting. So this option should be available for boosting, too. But I agree that max is the most common option.
I would suggest to keep the field in the index. But if you want to offer the optionally remove it, you can do it. I saw both use-cases.
In some cases the requirement was to boost on existing fields that are needed and present in the index anyway, for example stock of a product. Not the absolute value but an integer indicating "no stock", "low stock", "full stock".
In another case there's a special dedicated field for boosting which could be removed.
I general see your concerns and agree to keep things simple. But I already went through multiple iterations with this patch and already considered my proposal as the minimum set of features. So I agree with your patch in #22.
Thanks for your help!
Comment #25
drunken monkeyOK, all makes sense. Thanks for your detailed reply!
Committed. Thanks again!