Two hundred thousand objects, and the collector never ran

Two hundred thousand objects created and dropped, and zero collections. What puts an object in the root buffer, and why the collector's cost is set by what you keep alive.

Alden Pike12 min read

Two hundred thousand objects created and destroyed in one loop, and the cycle collector ran zero times. Not because it was disabled — because nothing in that loop was ever a candidate.

<?php

declare(strict_types=1);

final class Row
{
    public function __construct(public int $id, public string $title) {}
}

for ($i = 0; $i < 200_000; $i++) {
    $row = new Row($i, 'row');
    unset($row);
}

$status = gc_status();
printf("runs %d  roots %d  collected %d\n", $status['runs'], $status['roots'], $status['collected']);
runs 0  roots 0  collected 0

Now give the object one reference back to its owner. Same count, same loop shape, and the collector runs thirty-nine times.

<?php

declare(strict_types=1);

final class Comment
{
    public ?Post $post = null;
}

final class Post
{
    /** @var Comment[] */
    public array $comments = [];

    public function addComment(Comment $comment): void
    {
        $comment->post = $this;
        $this->comments[] = $comment;
    }
}

for ($i = 0; $i < 200_000; $i++) {
    $post = new Post();
    $post->addComment(new Comment());
    unset($post);
}

$status = gc_status();
printf("runs %d  roots %d  collected %d\n", $status['runs'], $status['roots'], $status['collected']);
runs 39  roots 10000  collected 585000

That loop took 23.354 ms against the first one’s 8.040, and the second program allocates more than the first, so not all of that gap is the collector. The part that is, is 13.553 ms — 58% of its own runtime, against zero for the loop above it. One assignment is what moved it there.

How these numbers were taken

PHP 8.5.10, NTS, arm64, Homebrew build, on an Apple M4 Pro laptop with 24 GB and 12 logical cores, running macOS and not quiesced. OPcache and the JIT are off, and memory_limit is 4G so the disabled-collector runs fit.

Timings are the median of 15 interleaved runs with three warmup rounds discarded, taken with hrtime() inside the process, and the range is given with each one. Process totals are wall-clock around the whole interpreter, taken outside it. Counts and memory figures from gc_status() and memory_get_peak_usage() were byte-identical across all 15 rounds of every configuration below, so they are printed as single values. Read the timings relative to each other rather than as absolutes for your hardware, for the reasons the method post sets out.

One measurement compares 8.5.10 against 8.4.23, which is a Homebrew build against a Herd build on the same machine. That section says so again where the numbers are.

What puts an object in the root buffer

Most PHP memory is not collected. It is released the moment the reference count on the value reaches zero, by the same code that decremented it, with no collector involved. The first loop above allocated 200,000 objects and freed 200,000 objects that way.

The collector exists for the case that mechanism cannot see. If two objects hold each other, dropping every outside reference leaves both counts at one, and neither will ever reach zero. The memory is unreachable and still accounted for.

So the engine watches for the only event that can create such a cycle: a refcount that is decremented and does not reach zero, on a type that could contain a reference back. That object is written into the root buffer as a possible root. Not garbage — possible.

<?php

declare(strict_types=1);

final class Node
{
    public ?Node $peer = null;
}

$baseline = memory_get_usage();
$alone = new Node();
unset($alone);
printf("plain object, unset       %+d bytes  roots %d\n", memory_get_usage() - $baseline, gc_status()['roots']);

$baseline = memory_get_usage();
$left = new Node();
$right = new Node();
$left->peer = $right;
$right->peer = $left;
unset($left, $right);
printf("cycle, unset             %+d bytes  roots %d\n", memory_get_usage() - $baseline, gc_status()['roots']);

$collected = gc_collect_cycles();
printf("after gc_collect_cycles()  %+d bytes  collected %d\n", memory_get_usage() - $baseline, $collected);
plain object, unset       +0 bytes  roots 0
cycle, unset             +112 bytes  roots 2
after gc_collect_cycles()  +0 bytes  collected 2

The first object cost nothing to release and was never buffered. The pair cost 112 bytes that unset() did not recover, and put two entries in the buffer. Each object is 56 bytes on this build — a class with no declared properties measured 40, and the one property adds a 16-byte slot.

That is the whole eligibility rule, and it is why the acyclic loop’s collector counters stayed at zero rather than merely staying cheap. The collector was not idle in that program. It had nothing to be idle about.

The threshold has not been ten thousand since 7.3

The manual states that the root buffer “has a fixed size of 10,000 possible roots (although you can alter this by changing the GC_THRESHOLD_DEFAULT constant in Zend/zend_gc.c in the PHP source code, and re-compiling PHP)”. Two things in that sentence do not survive contact with the file it names. The constant is not the buffer’s size, and nothing involved is fixed. These are the constants that are there:

#define GC_DEFAULT_BUF_SIZE  (16 * 1024)
#define GC_BUF_GROW_STEP     (128 * 1024)

#define GC_THRESHOLD_DEFAULT (10000 + GC_FIRST_ROOT)
#define GC_THRESHOLD_STEP    10000
#define GC_THRESHOLD_MAX     1000000000
#define GC_THRESHOLD_TRIGGER 100

Two separate quantities. The buffer is capacity, starting at 16,384 entries. The threshold is the number of buffered roots that triggers a run, starting at 10,001 — GC_THRESHOLD_DEFAULT, which is 10,000 plus the reserved first slot — and moving at runtime.

gc_adjust_threshold() decides which way after every run. Fewer than GC_THRESHOLD_TRIGGER objects collected — 100 — or a buffer still at or above the threshold when the run ends, and it goes up by GC_THRESHOLD_STEP, growing the buffer first if the new value would not fit. A productive run brings it back down by the same step, never below the default. Holding 100,000 objects alive in a cyclic structure and then generating garbage shows both halves:

event                           roots  threshold   buffer  collected
at startup                          0      10001    16384          0
after 100,000 live roots        40000      40001    65536          0
run #4                              2      50001    65536          0
run #5                              2      40001    65536      50000
run #6                              2      30001    65536      90000
run #7                              2      20001    65536     120000
run #8                              2      10001    65536     140000
run #9                              2      10001    65536     150000

Building the live structure triggered three runs that collected nothing, because everything they scanned was still referenced. Each one raised the threshold, and the buffer grew with it. Run #4 collected nothing either and pushed the threshold to 50,001. From run #5 the garbage was real, and five productive runs walked the threshold back to where it started.

The consequence is that a process holding a large live graph collects less often, not more, and each run it does perform is larger. Which is the wrong direction, and the next measurement is why.

The bill is the graph you keep alive

The threshold decides how often a collection happens. What decides what one costs is not in that constant at all.

This workload holds a doubly-linked structure alive — the shape an identity map, a registry or an in-memory cache has — and then produces exactly 50,000 garbage cycles that each hold a reference into it. The garbage is identical in all four rows. Only the live structure changes.

Live nodes Runs Collected ms per run Loop total
0 5 49,996 0.191 (0.182–0.220) 1.909 ms
10,000 5 49,997 0.356 (0.343–0.383) 2.817 ms
50,000 4 49,997 1.085 (1.044–1.279) 5.422 ms
200,000 1 10,000 6.745 (6.122–8.073) 7.866 ms

The first and last rows collected about 10,000 objects per run each. One took 0.191 ms to do it and the other took 6.745 — thirty-five times the cost for the same result, decided entirely by memory the collector was never going to touch.

The last row also shows the threshold effect from the previous section arriving in production form. Building 200,000 live nodes had already pushed the threshold to 60,001, so only one collection happened inside the measured window and 40,001 roots were still sitting in the buffer when it ended. Rarer runs, bigger runs, and garbage that lives longer.

What a collection walks

The algorithm runs in three passes — gc_mark_roots(), gc_scan_roots() and gc_collect_roots() — and the first one explains the table. Starting from each buffered root, it traverses every reference it can follow, decrementing a copy of the count on everything it reaches. The second pass restores whatever still has a count left, because a survivor is reachable from outside the candidate set. The third frees what neither pass could reach.

Traversal does not stop at the boundary between garbage and live data, because the boundary is what it is computing. A garbage object holding a pointer to your registry means the registry gets walked. Every run. The 200,000 nodes in the last row were visited, marked, and put back exactly as they were, and that work is the 6.745 ms.

This reframes the usual question. “Should I disable the collector” is asking about the garbage. The cost is in what the garbage points at.

What gc_disable() buys, and what it hides

Same cycle-producing loop, collector on and off, everything else held constant:

Loop Peak memory Runs Collector Process total
Collector on 23.354 ms (23.126–23.685) 2,252,368 39 13.553 ms 67.640 ms (65.813–68.813)
Collector off 14.557 ms (14.142–14.918) 70,276,032 0 — 61.398 ms (59.487–62.128)

The loop got 8.797 ms cheaper and peak memory went up by a factor of 31.2. That is the trade in its rawest form, and on a request that has to fit inside a memory_limit it is usually the wrong side of it.

The process-level column is the part a loop timer will not show you. End to end the saving is 6.242 ms, not 8.797, because 6.423 ms of the collector’s 13.553 was free_time — releasing memory. The disabled run does not skip that work. It defers it to shutdown, where the same objects are destroyed outside the loop’s clock.

Nor does disabling the collector stop the buffer from filling. The same loop with the collector off ended with 400,000 roots in a buffer that had grown from 16,384 entries to 524,288 — one pointer each, about 4 MB. It is allocated with perealloc(), which resolves to plain realloc() for a persistent allocation rather than going through the request allocator. So it does not appear in memory_get_usage() and it is not counted against memory_limit. Reading all three memory layers as deltas across the loop:

      used            real            rss           roots    buffer
on     +1,764,032      +2,097,152      +2,080,768    10000     16384
off   +69,787,288     +71,303,168     +75,366,400   400000    524288

With the collector off the kernel counted 4,063,232 bytes more than the engine’s own chunk figure did, and the root buffer at the end of that run held 524,288 pointers — 4,194,304 bytes. It is a fourth quantity that none of the three memory figures reports. Resident memory here varied by about 16 KB across five runs; the other two columns were byte-identical.

The obvious middle path does not rescue it either. Disabling the collector and calling gc_collect_cycles() at a job boundary — every 20,000 iterations, ten explicit collections instead of thirty-nine automatic ones — ran to 27.325 ms (26.260–28.056) against the automatic 24.213 (23.835–24.681) in the same script, with a peak of 7,562,320 bytes. It lost on both axes. Ten collections of 20,000 roots each cost more than thirty-nine of 15,000, because the per-run cost is not linear in the number of roots when they all point at the same place.

One script, two builds

One change in 8.5 is visible from userland, and it is narrower than the summaries of it suggest. php-src PR #19866 marks enum cases and fake closures as not collectable, which keeps them out of the buffer entirely. For closures the exclusion is conditional:

if (Z_TYPE(closure->this_ptr) != IS_OBJECT) {
    GC_ADD_FLAGS(&closure->std, GC_NOT_COLLECTABLE);
}

Nothing bound to $this, no flag to worry about. So the first-class callable syntax on a plain function is excluded and the same syntax on an instance method is not. Two hundred thousand of each, held in an array:

PHP 8.5.10  plain    runs 0  roots      0  threshold  10001  collected 0
PHP 8.5.10  bound    runs 5  roots  50006  threshold  60001  collected 0
PHP 8.4.23  plain    runs 5  roots  50001  threshold  60001  collected 0
PHP 8.4.23  bound    runs 5  roots  50007  threshold  60001  collected 0

Three of those four rows are five collections that freed nothing, cost 4.0 to 5.3 ms, and left the threshold at six times its default for the rest of the process. handle(...) on 8.5.10 is the row that never enters the buffer. $mailer->send(...) on the same build behaves exactly as it did on 8.4, because that closure holds an object and could be part of a cycle.

The enum half of the same commit is real and changes little: 200,000 reads of one enum case put two roots in the buffer on 8.4.23 and zero on 8.5.10, with no collection triggered on either. A case is a singleton, so there was never much there to exclude.

This is a Homebrew 8.5.10 against a Herd 8.4.23 on one machine, so it is a version comparison with a build difference in it. Treat the two build environments as a caveat on the timings rather than on the counters.

Four clocks the engine already keeps

Both of those readings came from the same place, and it is not an extension you have to compile. gc_status() returns twelve fields on 8.3 and later, four of them clocks in seconds:

Array
(
    [running] =>
    [protected] =>
    [full] =>
    [runs] => 0
    [collected] => 0
    [threshold] => 10001
    [buffer_size] => 16384
    [roots] => 0
    [application_time] => 6.7084E-5
    [collector_time] => 0
    [destructor_time] => 0
    [free_time] => 0
)

application_time is wall time since the collector was activated for this request. collector_time is the part of it spent collecting, with destructor_time and free_time naming the two phases inside that number which run userland __destruct() methods and release memory. So one division answers the question this article opened with:

$status = gc_status();
printf("%.1f%% of this request was collection\n", 100 * $status['collector_time'] / $status['application_time']);

On the acyclic loop that prints 0.0%. On the cyclic one it prints 58%.

The ratio between collected and runs is the second number worth logging. A process whose runs each free 15,000 objects is being served by the collector. One whose runs free 200 is paying for a walk over data that was never garbage, and the threshold is on its way up.

When the lever is the right one

Measure before reaching for gc_disable(), because on a workload that produces no cycles it buys nothing — the two acyclic rows above differ by 0.016 ms on an 8-millisecond loop, which is inside the noise. If runs is zero at the end of a representative request, the collector is not in your profile and turning it off changes nothing but your exposure to a future cycle.

Where it does apply, the shape that justifies it is a bounded unit of work whose peak memory you have measured and can afford: a short script, or a job that ends soon enough for the process to reclaim everything at once. The 5.3-era figures still quoted for this — about 7% slower with the collector on, 98% less memory — come from Derick Rethans’ 2010 benchmark, and the direction has held while the magnitude has not. On this loop and this build it is 60% slower and 31 times the memory, which is a much sharper trade than the numbers in circulation suggest.

For a long-running process the more useful lever is usually not the collector at all. A run’s cost is the graph its roots can reach, so cutting the reference from short-lived objects back into the long-lived structure removes the work rather than deferring it. That is what a WeakReference on a parent pointer buys: the same parent-and-child pair that put two roots in the buffer above puts zero when the back-reference is weak, and gc_collect_cycles() finds nothing to do. It is also why an identity map that is cleared between jobs is cheaper than one that is not — the second one is walked by every collection for the life of the worker.

Print gc_status() at the end of one real request before changing anything. Four fields decide it: runs, collected, and collector_time against application_time.

Frequently asked

Will gc_disable() make my worker faster?
Only if it is producing cycles, and then you are trading memory for it. On a loop that produced none, the collector never ran and the two configurations were within run-to-run noise of each other. On a loop that did, disabling the collector cut 8.797 ms off the loop and multiplied peak memory by 31.2.
My GC runs collect almost nothing. What is happening?
The roots being scanned are reachable from something that is still live, so nothing can be freed. The engine reacts by raising the threshold 10,000 at a time, which makes runs rarer but not cheaper — each one still walks everything the roots reach.
How do I tell whether the collector is costing me anything?
gc_status() returns collector_time and application_time as seconds since the collector was activated for this request. The first divided by the second is the fraction of the request spent collecting. On the cycle-producing loop here it was 58%; on the acyclic one it was zero.
Did PHP 8.5 change anything about garbage collection?
Enum cases and fake closures with nothing bound to $this are marked as not collectable, so they never enter the root buffer. Two hundred thousand first-class callables from a plain function ran the collector five times on 8.4.23 and zero times on 8.5.10. The same syntax on an instance method still buffers 50,006 roots on 8.5.10, because that closure holds an object.
Share

Written by

Alden Pike

Alden Pike writes about PHP internals, performance, architecture and production behavior — the layer beneath the frameworks. Measurements over assumptions, trade-offs over universal rules.

More about the author

Related posts

Arrow keys to move, Enter to open.