# Two hundred thousand objects, and the collector never ran

> Two hundred thousand objects created and dropped, and zero collections. What puts an object in the root buffer, and why the collector's cost is set by what you keep alive.

- Published: 2026-09-20
- Tags: memory, zend-vm, profiling
- Source: https://elephantphp.com/blog/the-gc-only-collects-cycles/
- Language: en-US
- Author: Alden Pike

---
Two hundred thousand objects created and destroyed in one loop, and the cycle
collector ran zero times. Not because it was disabled — because nothing in that
loop was ever a candidate.

```php
<?php

declare(strict_types=1);

final class Row
{
    public function __construct(public int $id, public string $title) {}
}

for ($i = 0; $i < 200_000; $i++) {
    $row = new Row($i, 'row');
    unset($row);
}

$status = gc_status();
printf("runs %d  roots %d  collected %d\n", $status['runs'], $status['roots'], $status['collected']);
```

```text
runs 0  roots 0  collected 0
```

Now give the object one reference back to its owner. Same count, same loop
shape, and the collector runs thirty-nine times.

```php
<?php

declare(strict_types=1);

final class Comment
{
    public ?Post $post = null;
}

final class Post
{
    /** @var Comment[] */
    public array $comments = [];

    public function addComment(Comment $comment): void
    {
        $comment->post = $this;
        $this->comments[] = $comment;
    }
}

for ($i = 0; $i < 200_000; $i++) {
    $post = new Post();
    $post->addComment(new Comment());
    unset($post);
}

$status = gc_status();
printf("runs %d  roots %d  collected %d\n", $status['runs'], $status['roots'], $status['collected']);
```

```text
runs 39  roots 10000  collected 585000
```

That loop took 23.354 ms against the first one's 8.040, and the second program
allocates more than the first, so not all of that gap is the collector. The part
that is, is 13.553 ms — 58% of its own runtime, against zero for the loop above
it. One assignment is what moved it there.

## How these numbers were taken

PHP 8.5.10, NTS, arm64, Homebrew build, on an Apple M4 Pro laptop with 24 GB and
12 logical cores, running macOS and not quiesced. OPcache and the JIT are off,
and `memory_limit` is 4G so the disabled-collector runs fit.

Timings are the median of 15 interleaved runs with three warmup rounds
discarded, taken with `hrtime()` inside the process, and the range is given with
each one. Process totals are wall-clock around the whole interpreter, taken
outside it. Counts and memory figures from `gc_status()` and
`memory_get_peak_usage()` were byte-identical across all 15 rounds of every
configuration below, so they are printed as single values. Read the timings
relative to each other rather than as absolutes for your hardware, for
[the reasons the method post sets out](/blog/benchmarking-php-without-lying-to-yourself/).

One measurement compares 8.5.10 against 8.4.23, which is a Homebrew build
against a Herd build on the same machine. That section says so again where the
numbers are.

## What puts an object in the root buffer

Most PHP memory is not collected. It is released the moment
[the reference count on the value reaches zero](/blog/what-a-php-variable-actually-is/),
by the same code that decremented it, with no collector involved. The first loop
above allocated 200,000 objects and freed 200,000 objects that way.

The collector exists for the case that mechanism cannot see. If two objects hold
each other, dropping every outside reference leaves both counts at one, and
neither will ever reach zero. The memory is unreachable and still accounted for.

So the engine watches for the only event that can create such a cycle: a
refcount that is decremented and does not reach zero, on a type that could
contain a reference back. That object is written into the root buffer as a
*possible root*. Not garbage — possible.

```php
<?php

declare(strict_types=1);

final class Node
{
    public ?Node $peer = null;
}

$baseline = memory_get_usage();
$alone = new Node();
unset($alone);
printf("plain object, unset       %+d bytes  roots %d\n", memory_get_usage() - $baseline, gc_status()['roots']);

$baseline = memory_get_usage();
$left = new Node();
$right = new Node();
$left->peer = $right;
$right->peer = $left;
unset($left, $right);
printf("cycle, unset             %+d bytes  roots %d\n", memory_get_usage() - $baseline, gc_status()['roots']);

$collected = gc_collect_cycles();
printf("after gc_collect_cycles()  %+d bytes  collected %d\n", memory_get_usage() - $baseline, $collected);
```

```text
plain object, unset       +0 bytes  roots 0
cycle, unset             +112 bytes  roots 2
after gc_collect_cycles()  +0 bytes  collected 2
```

The first object cost nothing to release and was never buffered. The pair cost
112 bytes that `unset()` did not recover, and put two entries in the buffer. Each
object is 56 bytes on this build — a class with no declared properties measured
40, and the one property adds a 16-byte slot.

That is the whole eligibility rule, and it is why the acyclic loop's collector
counters stayed at zero rather than merely staying cheap. The collector was not
idle in that program. It had nothing to be idle about.

## The threshold has not been ten thousand since 7.3

The [manual](https://www.php.net/manual/en/features.gc.collecting-cycles.php) states that the root buffer "has a fixed size of 10,000 possible roots
(although you can alter this by changing the `GC_THRESHOLD_DEFAULT` constant in
`Zend/zend_gc.c` in the PHP source code, and re-compiling PHP)". Two things in
that sentence do not survive contact with the file it names. The constant is not
the buffer's size, and nothing involved is fixed. These are the constants that
are there:

```c
#define GC_DEFAULT_BUF_SIZE  (16 * 1024)
#define GC_BUF_GROW_STEP     (128 * 1024)

#define GC_THRESHOLD_DEFAULT (10000 + GC_FIRST_ROOT)
#define GC_THRESHOLD_STEP    10000
#define GC_THRESHOLD_MAX     1000000000
#define GC_THRESHOLD_TRIGGER 100
```

Two separate quantities. The buffer is capacity, starting at 16,384 entries. The
threshold is the number of buffered roots that triggers a run, starting at
10,001 — `GC_THRESHOLD_DEFAULT`, which is 10,000 plus the reserved first slot —
and moving at runtime.

`gc_adjust_threshold()` decides which way after every run. Fewer than
`GC_THRESHOLD_TRIGGER` objects collected — 100 — or a buffer still at or above the
threshold when the run ends, and it goes up by `GC_THRESHOLD_STEP`, growing the
buffer first if the new value would not fit. A productive run brings it back down
by the same step, never below the default. Holding 100,000 objects alive in a
cyclic structure and then generating garbage shows both halves:

```text
event                           roots  threshold   buffer  collected
at startup                          0      10001    16384          0
after 100,000 live roots        40000      40001    65536          0
run #4                              2      50001    65536          0
run #5                              2      40001    65536      50000
run #6                              2      30001    65536      90000
run #7                              2      20001    65536     120000
run #8                              2      10001    65536     140000
run #9                              2      10001    65536     150000
```

Building the live structure triggered three runs that collected nothing, because
everything they scanned was still referenced. Each one raised the threshold, and
the buffer grew with it. Run #4 collected nothing either and pushed the threshold
to 50,001. From run #5 the garbage was real, and five productive runs walked the
threshold back to where it started.

The consequence is that a process holding a large live graph collects less
often, not more, and each run it does perform is larger. Which is the wrong
direction, and the next measurement is why.

## The bill is the graph you keep alive

The threshold decides how often a collection happens. What decides what one
costs is not in that constant at all.

This workload holds a doubly-linked structure alive — the shape an identity map,
a registry or an in-memory cache has — and then produces exactly 50,000 garbage
cycles that each hold a reference into it. The garbage is identical in all four
rows. Only the live structure changes.

| Live nodes | Runs | Collected | ms per run | Loop total |
|---|---|---|---|---|
| 0 | 5 | 49,996 | 0.191 (0.182–0.220) | 1.909 ms |
| 10,000 | 5 | 49,997 | 0.356 (0.343–0.383) | 2.817 ms |
| 50,000 | 4 | 49,997 | 1.085 (1.044–1.279) | 5.422 ms |
| 200,000 | 1 | 10,000 | 6.745 (6.122–8.073) | 7.866 ms |

The first and last rows collected about 10,000 objects per run each. One took
0.191 ms to do it and the other took 6.745 — thirty-five times the cost for the
same result, decided entirely by memory the collector was never going to touch.

The last row also shows the threshold effect from the previous section arriving
in production form. Building 200,000 live nodes had already pushed the threshold
to 60,001, so only one collection happened inside the measured window and 40,001
roots were still sitting in the buffer when it ended. Rarer runs, bigger runs,
and garbage that lives longer.

## What a collection walks

The algorithm runs in three passes — `gc_mark_roots()`, `gc_scan_roots()` and
`gc_collect_roots()` — and the first one explains the table. Starting from each
buffered root, it traverses every reference it can follow, decrementing a copy of
the count on everything it reaches. The second pass restores whatever still has a
count left, because a survivor is reachable from outside the candidate set. The
third frees what neither pass could reach.

Traversal does not stop at the boundary between garbage and live data, because
the boundary is what it is computing. A garbage object holding a pointer to your
registry means the registry gets walked. Every run. The 200,000 nodes in the last
row were visited, marked, and put back exactly as they were, and that work is the
6.745 ms.

This reframes the usual question. "Should I disable the collector" is asking
about the garbage. The cost is in what the garbage points at.

## What gc_disable() buys, and what it hides

Same cycle-producing loop, collector on and off, everything else held constant:

| | Loop | Peak memory | Runs | Collector | Process total |
|---|---|---|---|---|---|
| Collector on | 23.354 ms (23.126–23.685) | 2,252,368 | 39 | 13.553 ms | 67.640 ms (65.813–68.813) |
| Collector off | 14.557 ms (14.142–14.918) | 70,276,032 | 0 | — | 61.398 ms (59.487–62.128) |

The loop got 8.797 ms cheaper and peak memory went up by a factor of 31.2. That
is the trade in its rawest form, and on a request that has to fit inside a
`memory_limit` it is usually the wrong side of it.

The process-level column is the part a loop timer will not show you. End to end
the saving is 6.242 ms, not 8.797, because 6.423 ms of the collector's 13.553 was
`free_time` — releasing memory. The disabled run does not skip that work. It
defers it to shutdown, where the same objects are destroyed outside the loop's
clock.

Nor does disabling the collector stop the buffer from filling. The same loop with
the collector off ended with 400,000 roots in a buffer that had grown from 16,384
entries to 524,288 — one pointer each, about 4 MB. It is allocated with
`perealloc()`, which resolves to plain `realloc()` for a persistent allocation
rather than going through the request allocator. So it does not appear in
`memory_get_usage()` and it is not counted against `memory_limit`. Reading all
three memory layers as deltas across the loop:

```text
      used            real            rss           roots    buffer
on     +1,764,032      +2,097,152      +2,080,768    10000     16384
off   +69,787,288     +71,303,168     +75,366,400   400000    524288
```

With the collector off the kernel counted 4,063,232 bytes more than the engine's
own chunk figure did, and the root buffer at the end of that run held 524,288
pointers — 4,194,304 bytes. It is
[a fourth quantity that none of the three memory figures reports](/blog/where-php-memory-actually-goes/).
Resident memory here varied by about 16 KB across five runs; the other two
columns were byte-identical.

The obvious middle path does not rescue it either. Disabling the collector and
calling `gc_collect_cycles()` at a job boundary — every 20,000 iterations, ten
explicit collections instead of thirty-nine automatic ones — ran to 27.325 ms
(26.260–28.056) against the automatic 24.213 (23.835–24.681) in the same script,
with a peak of 7,562,320 bytes. It lost on both axes. Ten collections of 20,000
roots each cost more than thirty-nine of 15,000, because the per-run cost is not
linear in the number of roots when they all point at the same place.

## One script, two builds

One change in 8.5 is visible from userland, and it is narrower than the summaries
of it suggest. [php-src PR #19866](https://github.com/php/php-src/pull/19866)
marks enum cases and fake closures as not collectable, which keeps them out of
the buffer entirely. For closures the exclusion is conditional:

```c
if (Z_TYPE(closure->this_ptr) != IS_OBJECT) {
    GC_ADD_FLAGS(&closure->std, GC_NOT_COLLECTABLE);
}
```

Nothing bound to `$this`, no flag to worry about. So the first-class callable
syntax on a plain function is excluded and the same syntax on an instance method
is not. Two hundred thousand of each, held in an array:

```text
PHP 8.5.10  plain    runs 0  roots      0  threshold  10001  collected 0
PHP 8.5.10  bound    runs 5  roots  50006  threshold  60001  collected 0
PHP 8.4.23  plain    runs 5  roots  50001  threshold  60001  collected 0
PHP 8.4.23  bound    runs 5  roots  50007  threshold  60001  collected 0
```

Three of those four rows are five collections that freed nothing, cost 4.0 to
5.3 ms, and left the threshold at six times its default for the rest of the
process. `handle(...)` on 8.5.10 is the row that never enters the buffer.
`$mailer->send(...)` on the same build behaves exactly as it did on 8.4, because
that closure holds an object and could be part of a cycle.

The enum half of the same commit is real and changes little: 200,000 reads of one
enum case put two roots in the buffer on 8.4.23 and zero on 8.5.10, with no
collection triggered on either. A case is a singleton, so there was never much
there to exclude.

This is a Homebrew 8.5.10 against a Herd 8.4.23 on one machine, so it is a
version comparison with a build difference in it. Treat the two build
environments as a caveat on the timings rather than on the counters.

## Four clocks the engine already keeps

Both of those readings came from the same place, and it is not an extension you
have to compile. [`gc_status()`](https://www.php.net/manual/en/function.gc-status.php) returns twelve
fields on 8.3 and later, four of them clocks in seconds:

```text
Array
(
    [running] =>
    [protected] =>
    [full] =>
    [runs] => 0
    [collected] => 0
    [threshold] => 10001
    [buffer_size] => 16384
    [roots] => 0
    [application_time] => 6.7084E-5
    [collector_time] => 0
    [destructor_time] => 0
    [free_time] => 0
)
```

`application_time` is wall time since the collector was activated for this
request. `collector_time` is the part of it spent collecting, with
`destructor_time` and `free_time` naming the two phases inside that number which
run userland `__destruct()` methods and release memory. So one division answers
the question this article opened with:

```php
$status = gc_status();
printf("%.1f%% of this request was collection\n", 100 * $status['collector_time'] / $status['application_time']);
```

On the acyclic loop that prints 0.0%. On the cyclic one it prints 58%.

The ratio between `collected` and `runs` is the second number worth logging. A
process whose runs each free 15,000 objects is being served by the collector. One
whose runs free 200 is paying for a walk over data that was never garbage, and
the threshold is on its way up.

## When the lever is the right one

Measure before reaching for `gc_disable()`, because on a workload that produces
no cycles it buys nothing — the two acyclic rows above differ by 0.016 ms on an
8-millisecond loop, which is inside the noise. If `runs` is zero at the end of a
representative request, the collector is not in your profile and turning it off
changes nothing but your exposure to a future cycle.

Where it does apply, the shape that justifies it is a bounded unit of work whose
peak memory you have measured and can afford: a short script, or a job that ends
soon enough for the process to reclaim everything at once. The 5.3-era figures
still quoted for this — about 7% slower with the collector on, 98% less memory —
come from Derick Rethans' 2010 benchmark, and the direction has held while the
magnitude has not. On this loop and this build it is 60% slower and 31 times the
memory, which is a much sharper trade than the numbers in circulation suggest.

For a long-running process the more useful lever is usually not the collector at
all. A run's cost is the graph its roots can reach, so cutting the reference from
short-lived objects back into the long-lived structure removes the work rather
than deferring it. That is what a `WeakReference` on a parent pointer buys: the
same parent-and-child pair that put two roots in the buffer above puts zero when
the back-reference is weak, and `gc_collect_cycles()` finds nothing to do. It is
also why an identity map that is cleared between jobs is cheaper than one that is
not — the second one is walked by every collection for the life of the worker.

Print `gc_status()` at the end of one real request before changing anything.
Four fields decide it: `runs`, `collected`, and `collector_time` against
`application_time`.
