A hold waits on one anchor row per scope that bounds it, and takes them all with a single SELECT ... FOR UPDATE. Two holds naming overlapping sets have to take the rows they share the same way round, or each holds one and waits for the other. The statement said ORDER BY id, and that never did it.

ORDER BY sorts the rows that come back. The locks are taken while the rows are being read, so the order is the order the plan reaches them in, and the plan changes with the shape of the condition: a hold naming two scopes and a hold naming three can walk the same rows differently. Sorting afterwards is too late.

Measured on a load harness against MySQL 8: a house of eight thousand places in sixteen blocks with a tariff quota over all of it, forty bookers claiming at once for ninety seconds from an empty hall. 1891 and 2047 claims refused out of about six thousand, every one an HTTP 500 from SQLSTATE[40001] Deadlock found when trying to get lock raised in LockAnchors::lockRows(). It needs a few thousand lines against the slot before it shows, which is why nothing had caught it.

The fix

The one order every plan agrees on is the order the rows are stored in. So what a row is about becomes its primary key, (slot, scope, ref), and the serial id goes: a caller already holds the scopes it wants, so it addresses the rows directly and the clustered index is walked in that order whichever way the plan reaches it. Nothing referred to an anchor by id, the unique key this replaces was the same three columns, and dropping a slot anchors is now a prefix of the key rather than an index of its own.

It costs nothing. One statement, as before, so no query budget moves: holding one line is ten statements and appending is eleven, exactly as they are today. There is a version of this that keeps the serial id, looks the rows up first and locks them by id in a second statement. It works equally well on the harness and costs one statement per hold, and it is not what this does, because the identity is already in the caller hands.

The same run now leaves 50 refusals of 3752 against about 1950 of 6000, and the log contains no deadlock at all. Those fifty are a transaction stack error rather than a lock cycle, they are not diagnosed, and they are not this.

What is left

Deciding whether a line fits reads a grouped SUM over every consuming booking of the slot, taken FOR UPDATE, which holds a next-key lock over that range while other holds insert into it. Removing that read is #3617101, and the arm that has it was the only one to reach zero refusals.

Upgrading

The table changes shape, and this is before 1.0 so there is no update hook: a site that reinstalls gets it, and a site that does not needs the primary key moved by hand. The rows carry no information, so they can also simply be dropped and rewritten.

Tests

A kernel test cannot see any of this: Drupal SQLite driver drops FOR UPDATE silently, so there are no row locks to collide. What is asserted instead is the shape that makes the order deterministic, which is one statement, addressing the key, with no ordering clause pretending to be the guarantee. LockAnchorScopeTest recognised that statement by its ORDER BY and recognises it by the columns it asks for now. The load harness is the evidence for the behavior.

AI-Generated: Yes (Claude Code found this while measuring #3617101, wrote the fix and its tests, and took the measurements. The measurements are from a local load harness against MySQL 8.)

Issue fork yoyaku-3618513

Command icon Show commands

Start within a Git clone of the project using the version control instructions.

Or, if you do not have SSH keys set up on git.drupalcode.org:

Comments

mably created an issue. See original summary.

mably’s picture

Issue summary: View changes
mably’s picture

Status: Active » Needs review

  • mably committed 717a867a on 1.x
    fix: #3618513 Anchor locks are taken in the order the plan reaches the...
mably’s picture

Status: Needs review » Fixed

Now that this issue is closed, review the contribution record.

As a contributor, attribute any organization that helped you, or if you volunteered your own time.

Maintainers, credit people who helped resolve this issue.

Status: Fixed » Closed (fixed)

Automatically closed - issue fixed for 2 weeks with no activity.