Skip to content

Cluster topology is not updated before configCheckSeconds on RESP3 and requires @admin or @dangerous command #3254

Description

@seppo498573908457

Hi. I've been testing the newest 3.3.1.2706 version and found a behaviour that causes downtime up to configCheckSeconds (60 s default) when a cluster changes topology or a node goes down. This is a problem, because it doesn't allow 0 downtime when updating a Redis cluster to newer software version. A similar behavior was observed on earlier versions 2.7.27 and 3.1.31, but the following analysis was made with the newest.

My Redis (Redis Open Source) is a cluster of 6 nodes (3 master, 3 replicas). I tested different connection strings, but the most interesting one is:
<six-dns-names-to-cluster-nodes-and-default-port>,protocol=resp3,ssl=true,user=testuser,password=pass,channelPrefix=testuser-,maintNotifications=auto
I'm not sure if maintNotifications=auto should have an effect here, but it was on in hopes it helped.
And the users ACL is like this:
user testuser on sanitize-payload #hash resetchannels ~testuser-* ~__Booksleeve_TieBreak &__Booksleeve_MasterChanged +@all
Ie. all commands, username prefixed keys, booksleeve channel and key.

Flow

I bound some loggers to multiplexer event handlers and the order of events goes like this:

  1. Start round-trip key-value test. (write, read, delete, using random guid as key and value)(this is me clicking a tester app)
  2. Connection Multiplexer is created.
  3. Test success.

Now a simulated server maintenance and topology change by using direct Redis command "CLUSTER FAILOVER" on all replica nodes. This causes the replicas to become masters in a graceful manner and old masters become automatically replicas.

  1. Start round-trip key-value test.
  2. (Using cached Multiplexer from here on)
  3. ERROR in write:

StackExchange.Redis.RedisConnectionException: InternalFailure on [0]:SETEX testuser-f7eea6da-cbe4-41a0-8e82-7bebd92b54f2 (BooleanProcessor)
---> StackExchange.Redis.RedisCommandException: Command cannot be issued to a replica: SETEX testuser-f7eea6da-cbe4-41a0-8e82-7bebd92b54f2
at StackExchange.Redis.PhysicalBridge.WriteMessageToServerInsideWriteLock(PhysicalConnection connection, Message message) in /_/src/StackExchange.Redis/PhysicalBridge.cs:line 1654
--- End of inner exception stack trace ---

  1. Event: Redis Hash Slot Moved. HashSlot: 2901, OldEndPoint: Unspecified/server1-dns-name:6379, NewEndPoint: 10.10.10.15:6379.
    Note that the connection string defines servers with dns names, but the new endpoint server5 is seen as an IP address. There are sometimes more than one of these events. The "Unspecified" refers to dotnet EndPoint AddressFamily.
  2. Event: Redis Configuration Changed. EndPoint: 10.10.10.15:6379.
    This is server5. Note that this is ConfigurationChanged event, not ConfigurationChangedBroadcast.
  3. Start round-trip key-value test.
  4. ERROR in write. The same as last time.
  5. Start round-trip key-value test.
  6. ERROR in write. The same as last time.
    This keeps going until the next event:
  7. Event: Redis Configuration Changed. EndPoint: Unspecified/server6-dns-name.fi:6379.
    This is the configCheckSeconds triggered check up. Sometimes there are two of these events with different endpoints.
  8. Start round-trip key-value test.
  9. Test success. Subsequent tests pass, although a few errors might still occur.

Sender of all events is: "SERVER(SE.Redis-v3.3.1.2706)" where the SERVER is the one hosting the tester app that I was clicking.
Redis ACL LOG is empty, so user privileges were not limiting in this case.
There is no indication of activity on either booksleeve channel or key, however there still could be some.
Docs state that:

An endpoint that refuses every connection is evidence that what the client believes about the deployment may be wrong, so after three consecutive failed connection attempts to the same endpoint, the client re-reads the topology - the same refresh a MOVED or a configuration announcement would have caused, jittered and coalesced in the same way.

But this doesn't seem to apply when the endpoint throws "Command cannot be issued to a replica".

The Issue(s)

So the issues are:

  • I'm committed to high availability in my environments so the downtime of up to 60 seconds is a problem.
  • Seems that the endpoint will not respond with "MOVED" when it's the replica of the master of the correct slot, for whatever reason (see summary below for further analysis). This cuts one of the change detection methods.
  • It still uses "INFO" command (@slow, @dangerous) to resolve a changed cluster.
  • There seems to be no way to issue controlled downtime to cluster nodes for maintenance. At least the documentation isn't clearly stating it.

FWIW Opus 5.5 suggested to manually send a message to __Booksleeve_MasterChanged channel (redis-cli PUBLISH __Booksleeve_MasterChanged "*") after issuing CLUSTER FAILOVER so that the library would take notice when using RESP2. This doesn't work with RESP3, which it claims to be a bug in SE.Redis.

Workaround attempt: broadcast on the config channel

So I tried to make the library react immediately by publishing to the config channel after CLUSTER FAILOVER. Findings:

RESP3: the config channel is never subscribed.

  • PUBSUB NUMSUB __Booksleeve_MasterChanged returns "0" on every node, with and without the channelPrefix.
  • PUBLISH to either name returns "(integer) 0".
  • CLIENT LIST shows "sub=0" on testuser, with "flags=r" or "flags=N".

RESP2: the config channel is subscribed, but under the channelPrefix name, testuser-__Booksleeve_MasterChanged.

  • The unprefixed channel has no subscribers.
  • Before allowing &testuser-__Booksleeve_MasterChanged in the ACL, I get ACL LOG entries that state the user trying to access that channel.
  • After allowing &testuser-__Booksleeve_MasterChanged, PUBLISH testuser-__Booksleeve_MasterChanged "*" returns "(integer) 1".

I believe applying channelPrefix to the config channel is a bug. The cluster topology is shared by every client, whatever its prefix. With a prefix applied, a message would need to be sent out to all clientses MasterChanged channel (currently dozens), which almost completely defeats the purpose of that channel. It also means the ACL has to grant the prefixed name, which is not documented anywhere.

AI-Generated, Human Edited Summary

After a CLUSTER FAILOVER, writes fail with Command cannot be issued to a replica until the next configCheckSeconds tick (up to 60 s by default).

  • The client already knows the old primary is now a replica, but its slot map still routes writes to it. The error is thrown client-side in PhysicalBridge.WriteMessageToServerInsideWriteLock, so the command never reaches the server. No MOVED comes back, and the MOVED-triggered topology refresh never fires.
  • Hitting this condition does not trigger a topology refresh either. I would expect "slot owner is a known replica" to trigger an immediate refresh, just like MOVED does.
  • Under RESP3, the config channel is never subscribed, so the __Booksleeve_MasterChanged broadcast can't be used as a workaround. ServerEndPoint.HandshakeAsync subscribes it only when connType == ConnectionType.Subscription, and ConnectionMultiplexer.ActivateServer never activates the subscription bridge when RESP3 is known or assumed. Observed: sub=0 on all connections, and PUBLISH returns 0 with and without the channelPrefix.
  • Under RESP2, the config channel is subscribed with channelPrefix applied (testuser-__Booksleeve_MasterChanged). The handshake uses RedisChannel.Literal(...) without IgnoreChannelPrefix, so MessageWriter.Write(RedisChannel) prefixes it. Since cluster topology is shared by all clients, the config channel should ignore channelPrefix. At minimum, this should be documented, including the required ACL channel permission.
  • Possibly related: the endpoints are configured as DNS names, but the cluster reports IPs (Unspecified/server1-dns-name:6379 vs 10.10.10.15:6379). The same physical node may therefore be tracked as two separate endpoints.

Expected: the client should recover on the first rejected write (or within a few seconds), not after configCheckSeconds.

PS. When manually changing node roles in the cluster, should I also always manually send message to "__Booksleeve_MasterChanged" to minimize downtime?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions