Query counts are pinned by PlaceMapPerformanceTest and SlotBookingPerformanceTest, and they answer how much work one request does. Nothing answers the question an operator actually asks before an on-sale: how many people can book at once before the site stops coping, and what gives way first.
A functional test issues one request at a time, so it cannot answer that even in principle. This needs a load driver against a containerised site, and a documented baseline that can be re-run per release.
The scenario has to be the real journey
Hammering one endpoint measures that endpoint. What matters is the booking journey as it is actually walked, because the expensive parts are not evenly spread:
- Arriving on the calendar or the place map, which is a render.
- The first click, which opens the basket: it fetches a token in a request of its own and writes the order, and is measurably dearer than the ones after it.
- Two or three more places, which is the recurring click and the one most visitors repeat.
- Checkout and payment, which locks the order.
- Abandonment, which is common and leaves the sweep to clean up.
With a mix, not a uniform march: most arrivals browse, fewer book, some abandon mid-basket. The ratio changes which bottleneck arrives first.
Refusals are the majority case, not the edge
Capacity is finite, so a load run books the hall out and then every further hold is refused. That is not the test breaking down, it is the test reaching the interesting part: during a real on-sale most people meet a refusal, not a success. The refusal path has to stay cheap and correct under load, and measuring only successful holds measures the minority.
It does mean the harness has to be explicit about capacity: seed a hall large enough that the run does not exhaust it, or reset between arms, or deliberately exhaust it and measure the refusal path on purpose. All three are legitimate and they answer different questions.
What to record
- Throughput: successful holds per second, and the arrival rate at which it stops rising.
- Latency percentiles per step, p50, p95 and p99 separately for the render, the opening click and the recurring click. A mean hides the thing visitors complain about.
- Outcome mix, separating a legitimate refusal, a place already taken, from a failure: a timeout, a deadlock, a 500. These look alike in a status code count and mean opposite things.
- What saturates first. This is the real deliverable. PHP workers busy and queueing, MySQL
Threads_runningand connection limit,Innodb_row_lock_waitsandInnodb_row_lock_time_avg, memcache hit rate and evictions. The answer is a named resource, not a number.
Lock contention, as one output among these
Two specific questions fall out of the same instrumentation, and were the original reason for this issue.
- The slot row lock. Taking a place does
SELECT ... FOR UPDATEon the session row, serialising every hold on that session by design. That is the module's central promise, that a rush cannot overbook, and nothing demonstrates it under actual concurrency. - The booking list cache tag. Every hold invalidates
yoyaku_booking_list, two queries againstcachetags. #3615587: Reduce the query cost of taking a place records that as a known cost precisely because it is unmeasured, and deciding to stop emitting it means breaking a core-wide contract, which is worth doing only against evidence.
A prediction to falsify, reached by reading the code rather than by measuring: holds on the same session already serialise on the slot row lock, and tag invalidation is deferred to transaction commit, so it happens inside that already serialised section. If that holds, cachetags contention cannot appear in a single-session rush at all, and could only appear across concurrent holds on different sessions, which do not serialise against each other yet all write the same row. Running one arm on a single session and one across many would settle it. performance_schema.data_lock_waits names the row being waited on, which is what turns "the run was slow" into "the waiting was here".
Correctness under load, not only cost
The harness should assert the invariant while it is at it: with capacity N and more than N concurrent holds, exactly N succeed and the rest are refused, with no place held twice and no order left holding what it did not get. That is worth having on its own merits, independently of any optimisation it informs, and it is the only thing that would demonstrate the no-overbooking guarantee under real concurrency rather than by argument.
Caveats to record with every figure
One container on one machine gives a number for that machine. What transfers is the shape: which resource saturates first, whether a change moves the ceiling up or down, and where the latency knee is. What does not transfer is the absolute count of simultaneous bookers. Say measured or inferred, and never carry a figure from this harness into a claim about a live site.
Issue fork yoyaku-3615593
Show commands
Start within a Git clone of the project using the version control instructions.
Or, if you do not have SSH keys set up on git.drupalcode.org:
Comments
Comment #2
mably commentedComment #3
mably commentedOne measurement from the pre-alpha3 audit #3615981: Pre-alpha3 audit: refuse operator markup and scripted artwork on the place map, take the summary off its per-line reads, and put back the vocabulary, translations and docs the last 92 commits moved, for whoever writes the load test:
ConstraintPolicyResolver::resourceIdsSharingLimit()loads every resource type and every resource on each evaluation of across_resource_limitpolicy, with no memo and no cache, and the policy is evaluated at the hold and again at the checkout gate.It reads like an N+1 and is not one. Measured with
Database::startLog()aroundvalidateOrder()at 2, 10 and 40 resources, the query count stays flat at 5 to 6, because aloadMultiple()with no ids is a single read. What grows linearly is the number of entities hydrated, so the cost is memory and PHP time rather than queries, and it will only show up under the concurrency this issue is about.Left alone deliberately rather than pre-optimised: the right shape of a fix (a per-request memo, or a cache keyed on the shared name and tagged on the two list tags) depends on which of the two evaluations actually hurts, and that is what the load test is for.
Comment #5
mably commentedComment #6
mably commentedMR !239 adds the harness and the baseline it produced. The driver walks the journey rather than one endpoint, with a mix of visitors who browse, book or wander off, each with a cookie jar of their own; a sampler names what is saturating, reading the row a hold waits on out of performance_schema; and the site's half of it seeds the hall and checks afterwards that nothing was booked twice. There is a kernel test for the check itself, because a check that cannot fail asserts nothing: mutating it to count places the way PlaceAvailability does, as a set, makes it fail.
Measured on one container of Apache with mod_php against MySQL 8 and memcache, a hall of 8000 places on four sessions, visitors mixed 50 browsing, 40 booking, 10 abandoning, sixty second windows after a thirty second warmup, the hall emptied between rungs.
How many at once
The knee is between 40 and 80. Every doubling up to 40 bought nearly double; the next two bought 35 and 34 per cent. It is a bend rather than a wall, and the opening click stays about twice the recurring one until the knee, after which both medians converge and both tails run away.
What gives way first
The slot row lock, and nothing else is close. At 80 bookers the sampler caught 48 transactions waiting at once, with 13871 of 13877 wait edges on yoyaku_slot row 1; at 160, every one of 38392 was. Meanwhile the web server used 54 of its 150 workers at the median, memcache evicted nothing at any rung, and the machine never dropped below 7.8 GB free. Adding workers would not move a single booking.
The two arms, and the prediction
Holds only, same total concurrency, one session against four: 51.3 holds a second against 84.6, so spreading is worth 65 per cent. That is the price of the guarantee, quantified.
The summary predicted that cachetags contention could not appear in a single-session rush at all, because tag invalidation is deferred to commit and therefore happens inside the already serialized section. Half right. Across sessions it is the majority of what remains once the slot row is out of the way, 169 wait edges against 42 on the slot. On a single session it is not zero, it is 2 against 2582: deferred invalidation runs after the transaction commits, which is outside the lock, so two holds can reach that row together. It is small either way and nothing here argues for breaking the invalidation contract. The cost of that tag is now a measurement rather than a suspicion, and the arm that would show a regression exists.
Correctness under load
60 bookers against a hall of 200 places, no warmup so every hold falls inside the window: exactly 200 holds succeeded, 4281 refusals, zero failures, no place held twice, and the driver's count agreed with the hall's to the unit. The refusal path is also the cheap one, 16 ms at the median against 16 and 46 for a click that got its place, which matters because during an on-sale most people meet a refusal.
Two things worth recording about the method
A ten second warmup was not enough. Emptying the hall means rebuilding the caches, and a cold start under concurrency puts requests on the semaphore rows core uses to stop two of them rebuilding the same cache collector; that stampede sat inside the measured window and put a fourteen second reading in the map page's p99. At thirty seconds it is 225 ms.
The first attempt at 160 died with the kernel killing the database, and the explanation first written for it was wrong. Summing per-process RSS over the workers counts the shared binary and the opcode cache once per process, which made a worker look like 48 MB and the worker cap look unaffordable; the container's own accounting says about 7 MB each. The machine was simply busy with other work. Run again with it free, 160 completed. The sampler now records available memory per tick, because a site reaching its limit and a box reaching its limit look identical from the outside.
Follow-up
#3616896: Seat a party under the hold's own locks, so a rush is served rather than refused came out of this: the anchor is chosen from the lines a hold names when it should be chosen from the rules that apply to them. It also corrects a documentation claim this work turned up, that a tariff row is locked where no overall cap is set, which stopped being true when one way of taking capacity replaced two. Two tariffs may price the same grade, so a tariff anchor would let two holds pass the check for one seat.
Absolute figures are one machine's and do not transfer. What transfers is the shape: the queue forms at the session row, and the lever is spreading load across sessions rather than adding workers.
AI-Generated: Yes (Claude Code was used to write the harness, run it and draft this comment. I reviewed the code and the figures; every number above came out of a run rather than a reading of the code, and where something is inferred instead of measured the documentation says so.)
Comment #8
mably commentedComment #11
mably commented