When a Redis connection is reset mid-request (e.g. Connection reset by peer / RELAY_ERR_IO), the following methods throw an unhandled exception that propagates as an HTTP 500:
RedisBackend::getMultiple() — hgetall() pipeline
RedisBackend::setMultiple() — hMset()/expire() pipeline
RedisBackend::invalidateMultiple() — hGet()/hSet() loop
RedisCacheTagsChecksum::doInvalidateTags() — incr() pipeline
RedisCacheTagsChecksum::getTagInvalidationCounts() — get()/mget()
Drupal core's DatabaseBackend swallows connection errors gracefully (the cache simply misses). The Redis backend should do the same — a transient Redis connection issue should degrade to a cache miss, not crash the request.
Additionally, RelayFactory::getClient() calls $relay->connect($host, $port) without forwarding the timeout and read_timeout values from $settings, making those config keys silently ineffective when using the Relay interface.
Steps to reproduce:
Configure Drupal to use drupal/redis with the Relay interface.
Trigger a TCP connection reset on the Redis socket mid-request (e.g. via a rolling Redis pod restart in Kubernetes, or by simulating RST with tc).
Observe unhandled Relay\Exception: Connection reset by peer propagating as a 500.
Expected behavior:
getMultiple() returns [] (cache miss) on connection error, logs via \Drupal::logger('redis').
setMultiple() / invalidateMultiple() / doInvalidateTags() log the error and return without throwing.
getTagInvalidationCounts() returns [] on connection error.
RelayFactory forwards timeout and read_timeout from settings to connect()/pconnect().
Patch attached (against 2.0.0-alpha2): covers all five methods and the RelayFactory timeout fix.
Comments
Comment #2
berdirLogging for this is tricky. First, this might happen in early bootstrap where there is no container yet. That can be accounted for. Second, if this happens, it possibly happens on all future calls in the same requests, which might easily result in hundreds of log messages. A static counter that checks failures might be an option, that could then log once or twice and after a few more errors, completely shut down and not even try to connect again in that request? But even with a limit per request, this could quickly flood the logs if redis is permanently down.
There are existing issues about error handling and fallbacks to other bins, but I think that's quite challenging.