Problem/Motivation

setupVdbServerWithDefaults hardcodes two values in this recipe's recipe.yml that don't hold up across sites:

  • backend_config.chat_model is openai__gpt-3.5-turbo. A site with no OpenAI provider configured self-hosted embeddings only, or a different chat provider entirely, ends up with a server pointing at a model it doesn't have. This field only affects chunk-size token-count accuracy, not correctness, so a blank default is safe and doesn't assume any particular provider is present.
  • backend_config.embedding_strategy_configuration.chunk_size is 1000 tokens. Embeddings models with a smaller context window than that e5-large's ~512 tokens, for example return HTTP 500s once contextual fields are appended to each chunk, since the combined payload exceeds what the model accepts. Reproduced directly against e5-large, not theoretical.

Steps to reproduce

  1. Apply this recipe on a site with no OpenAI provider configured, or whose default embeddings provider has a ~512 token (or smaller) context window.
  2. Check search_api.server.content_vector's backend_config.chat_model it's openai__gpt-3.5-turbo regardless of what the site has available.
  3. Index any content long enough to need contextual chunking, this returns an HTTP 500 from the embeddings provider once chunk size overflows its context window.

Proposed resolution

  • chat_model: leave it blank by default, or resolve it from the site's default chat provider the same way embeddings_engine is already resolved dynamically in the same action.
  • chunk_size: either lower the default to something safer across common models (300 leaves real margin), or validate it against the resolved embeddings model's actual context window at setup time and clamp it down automatically.

Remaining tasks

  • Agree on blank-default vs. dynamic resolution for chat_model.
  • Agree on fixed-lower-default vs. dynamic clamping for chunk_size.
  • Patch SetupVdbServer::apply() (PHP code change(in the ai module) only needed for the dynamic-resolution options; the simpler fixes are a one-line recipe.yml edit).
Command icon Show commands

Start within a Git clone of the project using the version control instructions.

Or, if you do not have SSH keys set up on git.drupalcode.org:

Comments

abhisekmazumdar created an issue. See original summary.

abhisekmazumdar’s picture

Issue summary: View changes
a.dmitriiev’s picture

Chat model in the settings is not about the real model of Open AI provider. The chat model setting is only there to evaluate the size of the chunk in tokens with tiktoken library (composer dependency of AI Core). The setting chat model is just a parameter in tiktoken that controls how tiktoken counts tokens. This will be anyway an estimation value, as a real embeddings model will count tokens on its own anyway. So it is safe to leave it like that.

Regarding the chunk size - yes, let's lower it to 300

abhisekmazumdar’s picture

Assigned: Unassigned » abhisekmazumdar
arianraeesi’s picture

abhisekmazumdar’s picture

Assigned: abhisekmazumdar » Unassigned
Status: Active » Needs review

Open MR for lower it to 300

Thank You for inputs

a.dmitriiev’s picture

Status: Needs review » Fixed

Merged

Now that this issue is closed, review the contribution record.

As a contributor, attribute any organization that helped you, or if you volunteered your own time.

Maintainers, credit people who helped resolve this issue.

Status: Fixed » Closed (fixed)

Automatically closed - issue fixed for 2 weeks with no activity.