It uses a pool of memory. If the pool is full because too many clients become active at the same time, only a part of the clients will be served in a timely manner. The other clients will have to wait for their reads to be processed. For a proxy handling long-running connections, this seems acceptable to have a slight delay in the 1% of cases where too many clients wake up at the same time. Buying 20× the number of servers don't seem a sensible solution.