Problem/Motivation
https://dri.es/what-it-costs-to-run-drupal-infrastructure
I’ve been contemplating the issue of DA budget and infrastructure costs.
In a Slack thread in #drupal-infrastructure discussing that post, @drumm said:
GitLab CI is the real cost, no one else does 1/10th of the amount we provide for free
I’ve been on a kick to reduce waste via the gitlab_templates world, but that’s only for contrib. Unsurprisingly, the vast bulk of CI usage is from core. Here are the top projects circa January 2025 (the most recent data I have access to):

Yes, there is still a lot of room for incremental improvement: skipping known flakey tests to avoid having to re-run so many pipelines, converting Functional to Kernel or Unit, etc. But a radical idea came to me that I think we should explore:
What if we completely disable running any automatic pipelines on each push to a core branch?
Allow me to explain. Every commit we might push comes from an issue that already ran numerous pipelines, and the last pipeline must pass before we even consider merging it. Why re-run a whole other pipeline on the push? We’d still have daily jobs that would alert us if a missing rebase allowed 1 change to “break HEAD” due to interaction with another push that just happened. Worst case, we’d know about this within 24 hours. Meanwhile, our core developer base seems active enough that Slack usually lights up very quickly if HEAD is failing tests for any reason. Finally, I know we’d be running every possible pipeline configuration before we actually tag another release.
So I assert that the chances of shipping a core release with a regression due to not running the pipeline after every push is 0%. Meanwhile, it seems like this would approximately cut in half the total CI usage for "native core pipelines" with a single easy change. No other optimization is going to get anywhere close to that. However, since I don't have access to the underlying numbers, I have no idea if this "approximately cut in half" is at all true, nor if cutting that value in half is even meaningful improvement compared to all the issue forks.
I asked in the core subsystem maintainers' Slack channel about this idea to "read the temperature" and gauge interest. Generally, folks thought it was worth exploring, but we lack the info needed to make an informed decision. So I'm opening this public issue to explore it.
Steps to reproduce
Proposed resolution
Modify what branches run automatic pipelines "on-commit". Rely on nightly pipelines (and the crowd-sourced core contributor community) to alert us if we "break HEAD" due to interactions of different commits in the same day breaking each other due to missing rebases or whatever.
- Skipping all of the frequent random test failures from #3600653: Temporarily skip failing tests, round 2
- Release branches - 1) no daily pipeline, 2) run all environments on push
- Development branches - 1) No on push pipelines, 2) daily pipeline using all environments at ~0400 UTC 3) extra daily pipeline on weekdays UTC, 1200 UTC, using the current on-push matrix.
Issue to fixed the skipped tests in this round: #3600655: [meta] Fix and re-enable tests skipped for random failures, round 2
Remaining tasks
- Get real info on how much CI time we're using on:
- Scheduled pipelines
- on-commit pipelines when pushing changes to core branches (what this issue aims to eliminate)
- Total CI usage from all core issue forks
- Assess if eliminating on-commit pipelines would be the massive improvement I think it might be, or if everything is dwarfed by issue forks.
- Done. Risks from release management perspective. See #6
- On commit test runs usually find the following types of errors
- something that breaks a specific database driver
- commit of an MR that has a very stale test run that was green at the time but would have failed if run again
- same thing but two up to date MRs that cause each other to fail without merge conflicts
- cherry picks to another branch that didn't get a test run.
- If this change is made, then on release day release managers would have to remember to manually trigger a branch run if something for the release was just committed.
- On commit test runs usually find the following types of errors
- If this is likely to result in significant savings, open an MR against
mainto alter the.gitlab-ciworld to make it so. Done: MR !15152 - Decide how far to backport the change (if we get this far)
Estimates on the number of pipelines over a 6 month period: Comments #25 - #30
User interface changes
Introduced terminology
API changes
Data model changes
Release notes snippet
| Comment | File | Size | Author |
|---|---|---|---|
| 2025-07-31-Drupal-GitLab-CI-usage-from-Jan-2025.png | 172.1 KB | dww |
Issue fork drupal-3580398
Show commands
Start within a Git clone of the project using the version control instructions.
Or, if you do not have SSH keys set up on git.drupalcode.org:
Comments
Comment #2
dwwPS we could move the entire “on_push” configuration to “on_tag” to ensure full pipelines run whenever release managers push a tag, but before they create a release node.
Comment #3
dwwComment #4
mstrelan commentedMost of the issue fork pipelines run against main. I think we would still need on-commit pipelines for other branches.
Comment #5
dwwDefinitely.
Not so sure. 😅 Between nightly builds and
on_tagpipelines, I still think we could skipon_pushon non-mainbranches, too.But let's see what the numbers reveal. Maybe this is a drop in the bucket compared to issue forks, and we need to stay focused on the sorts of efforts I linked in related issues.
Comment #6
catchSo 99% of the on-commit runs don't find any issues. The ones that do tend to be:
- something that breaks a specific database driver
- commit of an MR that has a very stale test run that was green at the time but would have failed if run again
- same thing but two up to date MRs that cause each other to fail without merge conflicts
- cherry picks to 11.x that didn't get a test run. I did this yesterday by cherry picking a unit test change that triggered a deprecation on phpunit 10.
For everything except the database driver/11.x backports, MR runs in issue forks tend to find out that head is broken very quickly. Because of the disruption I tend to get pinged (if I'm the culprit or if the culprit isn't online) pretty quickly.
The main advantage of the branch runs then is to confirm that it is indeed head that is broken.
For database drivers and 11.x it can take longer to spot that head is broken anyway. Mostly because of the frequency of random test failures meaning that email alerts are unreliable.
So short version is, I think we could trial this.
The one bigger downside would be on release days but we could manually trigger a branch run if we've just committed something for the release.
Comment #7
dwwFantastic, thanks for that update! Really helpful and encouraging to hear. I'm crossing off "Explore other risks from release management perspective..." from remaining tasks.
Indeed. And/or what I proposed at #2:
Comment #8
catchThe tag runs fail every time because we have tests that don't like running against a tag. That is potentially fixable one by one but very hard to enforce over time.
Comment #9
quietone commentedI've summarized catch's response to the issue summary. It is a nice list of what on-commit finds and the impact on release managers.
Comment #11
dwwGuess it's inherently impossible to test this via an MR. 😉 But I left the pipeline running to make sure I didn't break anything else. Obviously needs careful review before we try it.
Comment #12
dwwComment #13
catchThinking about it more and briefly discussing with @quietone in slack, I think them main thing I'd want to go along with this is skipping all of the frequent random test failures from #2829040: [meta] Known intermittent, random, and environment-specific test failures. People are understandably reluctant to skip tests, but if we want to know when we've actually newly broken something, it's more important that a pipeline failure is a surprise rather than immediately triggering a hunt for which job to re-run. Especially with re-runs being broken for everyone in MAINTAINERS.txt at the moment.
e.g. #3285193: Temporarily skip random test failures that hide real test failures, part 4 where we did this previously, and #3267247: [meta] Fix and re-enable tests skipped for random failures where we unskipped them again. I don't think there's a current issue/MR to do that for the current batch but could have missed one.
That would independently be a quality of life improvement for most contributors, and we can skip the tests just at the point before the random test failure so we don't completely lose the test coverage that actually passes.
--
Also one possible way to roughly gauge the impact of this without the DA having to provide it:
A branch run looks like this: https://git.drupalcode.org/project/drupal/-/pipelines/774226/ that took 20 minutes wall time. I don't know how to convert that to CI/compute minutes though, and the total time report in https://git.drupalcode.org/project/drupal/-/pipelines/774226/test_report looks broken to me.
This month we had:
215 commits to main.
159 commits to 11.x
36 commits to 11.3.x
12 commits to 10.6.x
We could count the total requested CPUs in gitlab-ci.yml for jobs, then multiply that by the minutes. If we remove the lint + unit test jobs from the count, that will compensate for the fact they don't run on the different database pipelines. Then we have a CPU request * walltime calculation per branch run. Even that won't be accurate because jobs finish at different times, but most individual core test jobs run 2-4 minutes these days so it wouldn't be completely inaccurate, probably better than what we get from the gitlab UI by the looks of things.
This also suggests to me that if we can identify release branches and run the on-commit tests on those, that would require a fraction of the CI time compared to main/11.x branch runs, but also strike out the release day problem from the issue summary.
Comment #14
catchOne thing that has changed significantly in the past five years compared to the five previous years, is when issue pipelines used to take 55 minutes to run, it was a very time consuming process to re-run tests prior to commit. I don't always remember to do that when the last run is very old, but it's now very easy to type
/rebasein the comment box on the MR and get a fresh run in less time than it usually takes to do a last pass over the MR and figure out issue credit. Also pretty common for people to manually rebase MRs if they notice they're 150 commits behind HEAD or similar. We might end up doing a bit more than we do now if we go ahead here, but I don't think it would happen dramatically more than it already does.Comment #15
catchI'm +1 to trialling this on the non-release branches, tagging for RM review so the other release managers get a chance to chime in here.
Comment #16
nicxvan commentedI am all for this.
I also think just skipping the flaky tests will have a huge impact, the number of tests I rerun just for flaky tests is very high.
Comment #17
godotislateRe: #15, I'm +1 as well for trying out the non-release branches.
Comment #18
kentr commentedIt was suggested in #2829040-235: [meta] Known intermittent, random, and environment-specific test failures to put flakey tests in their own group (job?).
Could we run those first and use Gitlab workflow rules to abort if they fail?
Comment #19
catch@kentr I think that would make them even more disruptive than they are now - e.g. you wouldn't be able to see any more results until you've passed the gauntlet of the flakiest tests all passing at the same time.
For on-commit and scheduled tests, what we really need is that when it fails, we know it's because we committed something bad recently, not clicking through and seeing it's blockfilteruitest failing for the 2,000th time in three years.
Comment #20
quietone commentedI've updated the issue summary to include current proposal of
How long would the trial last?
Comment #21
catchI think we could try it for six months with the option to revert earlier if it causes noticeable problems before then. Noticeable problems would be main or 11.x broken for more than a day, or more often (we already break them occasionally but find out very quickly usually).
If it doesn't cause noticeable problems in six months then that's probably enough to stick with it - can always reverse that decision later.
Comment #22
catchOnly a theory but we regularly get gitlab queues backing up, delaying branch creation, MR pipelines and everything else, and this seems to coincide with commits to core.
Today we committed four issues to main between 11:03 and 11:28, the were all backport to 11.4.x, so that's 12 on-commit pipelines in the space of 25 minutes.
If we trialled this, only the 11.4 commit would have triggered on-commit pipelines, so a 1/3rd reduction. Commits to 11.3 and 10.6 would also trigger pipelines but that's less frequent and also desired anyway.
Comment #23
godotislateI also have the impression that gitlab queues get backed up at times coinciding with a series of core commits. Still +1 for a trial.
Comment #24
catchSummarising some notes from https://drupal.slack.com/archives/CGKLP028K/p1781003925872199 where we are trying to diagnose the gitlab sidekiq (probably) queue getting backed up for 30-60 minutes at a time.
Looking at an on commit pipeline. This one is 36 jobs https://git.drupalcode.org/project/drupal/-/pipelines/842881 for the main pipeline (couple of these are manual so slightly less actually run), + 21 jobs for each environment. We run three non-default environments for on-commit jobs. So that's 36 + 63 = > 96 jobs per commit, per branch.
So when we committed and backported four issues to main, 11.x, 11.4.x in 30 minutes. That's 150 jobs * 3 * 4 = ~1152 jobs in half an hour.
Given those branches were main, 11.x and 11.4.x, this change would mean that instead of 1152 jobs in half an hour, we'd do closer to 400 for four commits, so 1/3rd the workload when we backport to a release branch, and no jobs at all when we don't.
For comparison, project_analysis does about 50 jobs to analyze 7,000 modules, then another 20-50 to post issues. So about 100 jobs max in total per run.
edit: counted the jobs wrong in the gitlab UI in the first attempt at this comment, fixed since.
Comment #25
catchDiscussed this a bit with @longwave in slack based on the above and also looked at commit rate against the various active branches.
Background is:
On-commit pipelines we do default + 3 extra environments. Daily pipelines we do default + 6 extra environments.
So four environments for on commit and 7 environments for daily. Let's call them 'environment pipelines', there is probably a proper gitlab term which I can't remember.
On the 10.6.x branch, we've made exactly 100 commits to that branch in the past six months:
183 daily+ 100 on-commit
(183 * 7=1281) + (100 * 4= 400) = 1681 environment pipelines.
On the 11.3.x branch, we've made 249 commits in the past six months
183 daily + 249 on-commit.
(183 * 7=1281) + (249 * 4) = 2227 environment pipelines.
For those release branches, we can compare to if we just ran every pipeline on commit and dropped the daily job altogether:
10.6.x
100 commits * 7 = 700 (compared to 1681)
11.3.x * 7 - 1743 (compared to 2227).
So if we had run all environments for all commits for the two release branches, we would have run 2443 environment pipelines instead of 3908 and got better feedback.
~
1465 less environment pipelines and got more instant feedback - something like a 1/3rd saving in CI time for better results.
On the other hand.
For main, we have made 924 commits in the past six months.
(183 * 7=1281) + (924 * 4 = 3696) = 4977
For 11.x, we've made 763 commits in the past six months.
(183 * 7 = 1281) + (763*4=3052) = 4333
Add them together and you get 9310.
@longwave suggested two different ideas in slack:
1. Using a delay + cancellation for on-commit jobs to try to run on-commit pipelines every x hours.
2. A more simple version of that: drop the on-commit jobs but make the daily jobs twice-daily.
If we had twice-daily pipelines for main and 11.x and no on-commit pipelines, that's
183 * 7 * 2 * 2 in six months, so 5124 environment pipelines, compared to 9310.
For the approximately 1/3rd of commits that make it back to release branches, we'd still have an on-commit pipeline for those anyway with the above plan. For the other 2/3rds, if we don't find out before, we'd know we broke head in < 12 hours.
Comment #26
catchRefinement of #25.
Most of our commits happen on weekdays, so we could try something like this for development branches (11.x and main):
Keep the daily scheduled pipeline more or less as is, at say 4am UTC.
Add a 'twice daily weekday pipeline' that runs the current on-push matrix at 12pm and 8pm UTC.
That would be 7 * 7 + 10 * 4 environment pipelines per week. Which works out as 2314 * 2 = 5096 - less CI time than if we run the daily job twice per day every day, but more of a safety net during the week.
Comment #27
quietone commentedIf I understand correctly, the last proposal is to make changes for 11.x and main only. The changes are 1) stop on_push pipelines 2) add a pipeline that runs twice a day on weekdays UTC using the current on-push matrix.
If that is right, it is worth trying. Thanks for working through the numbers.
Comment #28
catch@quietone there would be a separate change to expand the matrix of tests for on-push pipelines for release branches and get rid of the daily jobs entirely, it's bit lost in #25 nearer the beginning.
I think we can split this into two separate issues - changing the on-commit matrix for release branches (might need a RELEASE_BRANCH=1 in gitlab.yml which we update when branching a new release branch) the continuing here for the development branches.
Comment #29
quietone commentedLost idea.
Is this the intention?
Release branches - 1) no daily pipeline, 2) run all environments on push
Development branches - 1) No on push pipelines, 2) daily pipeline using all environments at ~0400 UTC 3) Twice daily pipeline on weekdays UTC, 1200 and 2000 UTC, using the current on-push matrix.
For weekdays, instead of 3 pipeline runs on weekdays can it be reduced to 2, 1 on all environments and 1 on the on-push matrix? They could be at the times suggested for the twice daily, 1200 and 2000. Just that one would be full.
Comment #30
catch@quietone
Yes that sounds right and I think twice daily would be fine. Once we've made all the code changes to support this we'll be able to vary the schedule easily via gitlab without core changes. We could also open a follow-up for @longwave's rate limiting idea which will be harder to do but give us more flexibility than the scheduling.
Comment #31
quietone commentedJust updating the issue summary with the latest decision
Comment #32
quietone commentedComment #33
quietone commentedCreated an issue to skip tests and a followup to fix the ones skipped in this round.
#3600653: Temporarily skip failing tests, round 2
#3600655: [meta] Fix and re-enable tests skipped for random failures, round 2
Most of the release managers have commented in agreement. And this is a trail, so I am removing the tag.
Comment #34
catchI think we need to change the MR so that it checks a RELEASE_BRANCH variable from gitlab pipeline settings. When RELEASE_BRANCH is set, we'll run the on-commit pipelines, then it's not, we don't.
That way, we can have the same gitlab-ci.yml for main, 11.x and 11.4, otherwise we'll need to revert this commit when we branch 11.5.x off main.
Also to keep changes reviewable, let's commit that first in this issue. We can then schedule an extra daily pipeline on main and 11.x which gives us the 'two daily runs per day'.
Then in a follow-up, we can tweak the matrix for daily to support 'full' and 'light' workloads.
Comment #36
catchJust pushed a commit for #34 then realised it won't work as-is because there's nowhere to put variables for on-push jobs in the UI unlike scheduled jobs.
So... maybe we can just set a variable higher up in the YAML, then when we branch a new release branch, we'll need to flip from 0 to 1, this could be added to the branching script if we want or at minimum it's an easy change.
Comment #37
fjgarlin commentedDoes that variable need to be documented somewhere?
For the variables to appear in the UI, they need to have "value" and "description", so they need to be defined with long format, instead of just "NAME_OF_THE_VAR: value".
Comment #38
catch@fjgarlin I don't think so particularly - it'll only be used in the
scheduled jobsone place in the pipeline, not even on the scheduled jobs or not for the moment, we don't really have anywhere to document any of this except maybe an inline comment.I don't think I've ever seen a gitlab variable show up in the UI, is that on the pipeline schedule page (where we currently set variables like that via the text fields)?
Comment #39
bbralaYou can manually start a schedules job (i do that a lot), which will surface the vars.
Comment #40
cmlaraSince the team always uses the git CLi to push they can set variables while pushing as a push option.
https://docs.gitlab.com/topics/git/commit/#push-options
Notably a predefined options already exists for this
ci.skipandci.no_pipelines.I would suggest use a CI variable ( Settings > CI/CD > Variables), set it to whatever is your current release branch and use a comparison in the job rules rather than having to flick 0/1 on multiple branches. Yes the value should be updated using the GitLab API as part of your branching script.
Comment #41
catchDon't think this helps - would add an extra thing to forget when pushing.
We have up to three release branches. But we can probably do RELEASE_BRANCH_NEXT || RELEASE_BRANCH_CURRENT || RELEASE_BRANCH_LTS
Comment #42
cmlaraIf your willing to expand the definition a bit and say "anything that isn't N.x or main is a release branch" you could also use a simple regex match.
The grey area would be when 11.5.x splits from 11.x, however IIRC that is usually only around what a month or two ahead of release isn't it? At that point arguably testing on every push becomes more critical anyways as your working to an actual release. similar holds true for when splitting 12.0.x from Main. You still save on Main and N.x not being tested in those cases.
The security branches would see every commits tested compared to a scheduled job however the security branches also need to be ready to release a security critical fix with zero notice so you likely want to be testing those on commit anyways.
Comment #43
catchYes we'd want on-commit against those (and very infrequent scheduled pipelines, if at all), which makes me realise we can probably hard-code main and 11.x, and be fine here for a very long time. We may never open a 12.x branch, 11.x was a hack to get around not having main as an option on d.o
Comment #44
catchPushed a commit for #43.
Comment #45
catchRe-titling.
Comment #46
catchComment #47
smustgrave commentedIs this something we should just try for a week or two to see if it causes issues?
Comment #48
mondrake#47 let’s try! It’s a one line change, easy to revert if something breaks horribly.
Comment #51
godotislateCommitted and pushed 5c22a54 to main and 8799478 to 11.x. Thanks!
Comment #53
godotislateI'm not sure the change is working as expected because it seems like on commit jobs are still running for main https://git.drupalcode.org/project/drupal/-/commit/5c22a54bcdd8a23991186... and 11.x https://git.drupalcode.org/project/drupal/-/commit/87994782d9693ecb72c79....
Comment #54
longwaveThere are a bunch of other rules sections that I think we need to add this to:
We could block this at the top level workflow.rules - but I think we shouldn't, because we do need the "Lint cache warming" jobs to run on 11.x/main. These are used to update caches that are used in merge request pipelines.
Comment #55
longwaveWhat didn't run on main/11.x is the jobs tagged
run-on-commit, so this did partially work.I think instead of
we should try something like
Then the
run-on-commitrule is more reusable in other workflows in the same wayComment #56
catchPut an MR up for #55.
Comment #58
smustgrave commented2nd times the charm?
Comment #59
longwaveThat solves the spirit of #55 but not the actual problem, in that there are other instances of
elsewhere in the file, that also need changing here.
Comment #60
catchAdded to more places.
After doing that I realised we possibly only needed to make this change in the
workflow:section, because it's that which determines whether the pipeline will run at all, and if the pipeline doesn't run, then none of the other jobs will anyway. However probably doesn't hurt to have the logic in the macro and double up anyway.Comment #61
longwave@catch adding it to the top level workflow will also stop the lint cache jobs from running, which are needed for downstream MRs.
Comment #62
nicxvan commentedCan we add a separate job for priming the cache that does run on commit? So we can still block it top level and not worry about all additional jobs?
Comment #63
catch@longwave if there's no pipeline at all then I think the lint cache will come from the previous pipeline which should be OK. Then the daily job will refresh the artifacts. Could be 100% wrong about that but this was my guess - e.g. we only need the lint cache job when a pipeline actually runs (but otherwise would wipe them).
Comment #64
smustgrave commentedThis one good to least try for a trial period? Since we just did a core release hope we have a window before the next one.
Comment #65
catchComment #66
catchThe YAML is invalid in the current MR.
Comment #68
longwaveI think this is correct now...
Comment #69
catchLet's try this and see what it looks like.
I still think it ought to be safe to skip the lint jobs if the entire pipeline doesn't run, but they are cheap to run, and it also means the cache will be more up-to-date for MR runs after a commit, especially if it's something like a phpstan update which completely invalidates the cache, so we might as well keep those anyway.
edit: also phpstan is sometimes enough to spot problems with cross-commits that introduce issues, so having that run gives us a bit of early feedback if something silly happened.
Comment #70
smustgrave commentedPer last comment
Comment #72
catchCommitted/pushed to main, thanks!
This will need a backport MR for 11.x
If we break HEAD and realise that on-commit tests would have notified us quicker, we might need to revisit this, but let's see how it goes for a while. I think keeping the on-commit pipelines for release branches means the risk here is a lot lower than it otherwise might be.
edit: looks like this might completely skip the pipeline after all, maybe I misread the YAML, but also think that's fine to try for a while.
Comment #74
longwaveBackported to 11.x with the help of GPT 5.6, because I don't remember how 11.x pipelines work really. I wish we had backported the recursive pipeline there as it makes more sense to me now...
Comment #75
smustgrave commentedLooks good.
Comment #77
catchCommitted/pushed to 11.x, thanks! Let's see how this goes for a bit.
Comment #78
catchCredited @xjm for discussion in slack - this is partly what led to skipping only on main/11.x and not release branches, so we get immediate feedback if we break something on those that otherwise doesn't get caught.