How a PHP request actually runs
10 min read
What happens between nginx and your PHP code: PHP-FPM, OPcache and the share-nothing model - and why they explain most PHP performance and scaling problems.

Most PHP developers work inside a framework and rarely see what happens before the first line of their code runs. That is fine until the application slows down under load, returns a 502 on a busy afternoon, or behaves differently on two servers. Then the details matter - and almost all of them come from one unusual design decision: in PHP, every request starts from nothing.
The share-nothing model
In many runtimes, an application starts once and then serves requests for days, keeping objects, caches and connections in memory between them. Classic PHP does the opposite. Each request begins with an empty state: the framework boots, configuration is loaded, services are built, the request is handled - and at the end, everything is thrown away.
This sounds wasteful, and it has a cost, but it also gives PHP some of its most useful properties:
- Nothing accumulates. A memory leak or a corrupted object lives for one request, not for a week.
- Failures are isolated. A fatal error ends one request, not the application.
- Any worker can serve any request. There is no state to keep in sync, so adding servers is straightforward.
Not everything is discarded: OPcache, APCu and persistent database connections survive between requests. What is thrown away is your application's state - and almost everything else in this article follows from that.
From the socket to your code
A typical production setup has two separate processes in front of your code.
The web server - usually nginx - accepts the HTTP connection. It serves static files itself and passes everything else to PHP over FastCGI, through a Unix socket or a TCP port.
PHP-FPM (FastCGI Process Manager, part of PHP since 5.3.3) runs a master process and a pool of worker processes. The master only manages the pool. Each worker handles one request at a time: it receives the request from nginx, runs your front controller (public/index.php), sends the response back, resets its state and waits for the next one.
That last point is the most important sentence in this article. The number of workers is the number of requests your server can process at the same moment. When all of them are busy, new requests wait in a queue. If they wait too long, or the queue itself fills up, the user sees an error.
What OPcache does
PHP is compiled before it runs: the source is parsed into opcodes, and the opcodes are executed. Without caching, this compilation would happen for every file on every request.
OPcache stores the compiled opcodes in shared memory, where all workers in the pool can use them. On a production server it is not optional. A few settings deserve attention:
opcache.memory_consumptionandopcache.max_accelerated_filesmust be large enough for the whole codebase, includingvendor/. When the cache is full, files that no longer fit are compiled again on every request, and nothing in the application tells you.opcache.validate_timestamps=0stops PHP from checking whether files have changed, which saves work on every request. It also means a deploy is not visible until OPcache is reset - usually by reloading PHP-FPM as the last step of the deploy. Callingopcache_reset()from the command line does not help: the CLI has its own cache, separate from the one FPM uses.- Preloading (PHP 7.4+) loads selected classes into memory when FPM starts, removing part of the bootstrap cost for every request.
- JIT (PHP 8+) helps CPU-heavy code such as image processing or calculations. A typical web request spends most of its time waiting for the database and other services, so the JIT rarely changes much there.
Where the time goes
A request spends its time in three places: booting the framework, running your code, and waiting - for the database, a cache, a file, an external API.
While a worker waits, it does nothing else. It is not freed to serve another request. This is why a slow external dependency is dangerous for a PHP application, and the arithmetic is simple:
throughput ≈ number of workers ÷ average request time
With 20 workers and requests that take 100 ms, a server can handle about 200 requests per second. If a payment provider or a reporting API starts responding in 10 seconds, the same 20 workers can handle two requests per second. The CPU sits idle, the pool is full, and every other page on the site - including the ones that never call that API - starts timing out.
More workers help only while memory allows. For a slow dependency, the real fixes are a short timeout on the external call, a circuit breaker that stops calling it while it is failing, a cache in front of it, or moving the call out of the request entirely.
Sizing the pool
The pool size is limited by memory, not by CPU. A useful starting point:
pm.max_children ≈ memory available for PHP ÷ average memory per worker
Measure the second number on the real application, under realistic traffic, rather than guessing. Be careful with ps and top: the resident size (RSS) they show counts the shared OPcache memory in every worker, so it overstates what each additional worker costs. Proportional set size (PSS), shown by tools such as smem, is a better basis. A Laravel or Symfony worker commonly uses tens of megabytes; one that builds large reports can use much more. Setting max_children higher than memory allows does not add capacity: once the server starts swapping, every request gets slower at the same time.
A few other settings matter:
pm:statickeeps a fixed number of workers and is the most predictable choice on a dedicated server;dynamicandondemandsuit shared or low-traffic servers.pm.max_requestsrecycles a worker after a number of requests, which contains slow leaks in extensions.- Database connections: each busy worker can hold a connection. Workers per server multiplied by the number of servers must stay below what the database accepts - a limit that is easy to hit when adding servers "to handle the load".
What this means for your code
Share-nothing shapes how an application has to be written:
- State between requests must live outside PHP. Sessions stored in files only work across several servers if the load balancer keeps each user on the same one; in Redis or a database they work everywhere. File sessions also lock: PHP holds the session file for the whole request, so parallel requests from the same user - several AJAX calls on one page - run one after another. APCu is fast but local to each server. Uploaded files belong in object storage once there is more than one machine.
- Static properties and singletons last one request. That is a guarantee under PHP-FPM - and it disappears with long-running runtimes such as Laravel Octane, RoadRunner, Swoole or FrankenPHP in worker mode. They remove the bootstrap cost by keeping the application in memory, which is a real performance gain, but state left in a static property now leaks into the next request. Code written for one model is not automatically safe in the other.
- Long work does not belong in the request. Sending emails, generating documents and calling slow services should go to a queue, processed by separate workers that are allowed to take their time.
Three timeouts that must agree
A request can be stopped in three different places, and they are often configured by different people:
max_execution_timein PHP. On Linux it counts only the time PHP spends executing code - time spent waiting for the database or a network call is not included. A request stuck on a slow API can run far longer than this value suggests.request_terminate_timeoutin the FPM pool, which kills the worker after a fixed wall-clock time, whatever it is doing.fastcgi_read_timeoutin nginx, which limits how long nginx waits between two reads from PHP - not the total time of the request. A script that sends output from time to time can run well past it.
If nginx gives up after 60 seconds but PHP keeps working for five minutes, the user sees an error while the server keeps spending a worker on a response nobody will receive. Each inner timeout should be shorter than the one outside it: FPM should stop PHP before nginx gives up, and nginx before the load balancer or CDN in front of it does.
Reading the symptoms
Most PHP production incidents announce themselves in the same few ways:
- 502 Bad Gateway - nginx could not get a response from PHP-FPM at all: the pool is down, the socket is wrong, a worker crashed - or the pool's listen queue is full and FPM refuses new connections, which is simply overload.
- 504 Gateway Timeout - PHP took longer than
fastcgi_read_timeoutto send anything. - "server reached pm.max_children setting" in the FPM log - the pool is exhausted. Before raising the limit, find out what the workers are waiting for.
- The slow log (
request_slowlog_timeout) writes a stack trace of any request that runs too long. It is one of the most useful and least used diagnostics in PHP. - The FPM status page shows active and idle workers and the length of the queue - the numbers that tell you whether the server is busy or simply stuck.
How other runtimes handle the same problem
The worker pool is PHP-FPM's answer to concurrency. Other platforms made different trade-offs:
- Node.js runs one long-lived process with an event loop. While a request waits for I/O, the same process serves others, so a slow API does not exhaust capacity the way it does in FPM. The price: CPU-heavy work blocks every request on that process, state persists between requests, and one unhandled crash takes down every request in flight. Using several CPU cores means running several processes.
- Java servlet containers such as Tomcat use a thread per request from a fixed pool - the same arithmetic as FPM, but threads share one heap, so each concurrent request costs less memory. Virtual threads in Java 21 make blocking I/O cheap and largely remove the pool limit for I/O-bound work.
- Go starts a goroutine per request inside one long-running process. Goroutines are cheap enough that waiting costs almost nothing; long-lived shared state is the norm, with the discipline that requires.
- Python has both models: WSGI servers such as Gunicorn with sync workers behave much like PHP-FPM, while ASGI servers such as Uvicorn run an event loop like Node.js.
- PHP itself is no longer limited to FPM. Swoole, RoadRunner and FrankenPHP keep the application in memory, and since PHP 8.1, Fibers let event-loop libraries such as ReactPHP and AMPHP do non-blocking I/O.
What PHP-FPM gives up in raw efficiency it gets back in operational simplicity: no shared state, isolated failures, predictable memory. For most business applications, where the time goes into database queries, that is still a good trade - as long as you know the arithmetic behind it.
Why it's worth knowing
Frameworks, libraries and tooling around PHP have changed completely since the PHP 5 days. This model has not. Understanding it turns most PHP performance questions into arithmetic: how many workers, how long each request takes, and what they are waiting for. I've spent most of my career on both sides of this line - running the servers and writing the code on top of them. Most of the problems neither side can explain alone are solved right here.